Max via Cursor, should've mentioned, very possible I'm seeing more Cursor than OAI. All the subagents it picked were also Sol Max... I've seen 7-10 in one turn. Regardless the top reasoning tiers being conflated with subagent count feels double edged
I've found Sol's propensity for delegating to subagents can make it... disastrously expensive, especially with each subagent having some implicit floor on further reasoning/context gathering before action.
The base model is certainly cheaper and more token efficient etc, but on large tasks cost in some way is now n^2
imo there’s a clear greenfield to have doubled down on where cursor was before in proactively keeping devs appraised of the code that they’re generating, and bridging growing gaps between abstracted chat sessions and files/directory structures I might understand less and less
This on the other hand feels like a clear reaction to cc/codex, in a way that even kind of builds an offboarding ramp?
Many are saying codex is more interactive but ironically I think that very interactivity/determinism works best when using codex remotely as a cloud agent and in highly async cases. Conversely I find opus great locally, where I can ram messages into it to try to lever its autonomy best (and interrupt/clean up)
where we draw the line on agent "identity" when the models being orchestrated are generally the same 3 frontier intelligences is an interesting question indeed
I would think this idea of creating a third-party to verify things likely centers more around liability/safety cover for a steroidal increase in velocity (i.e. --dangerously-skip-permissions) rather than anything particularly pragmatic or technical (but still poised to capture a ton of value)
Stumbled on to a similar concept a few years ago, I think there was a paper around some kind of proto code skills for Minecraft around then... awesome to see Anthropic pursue this direction
I myself run five IDE's in parallel with ~ 6 terminal sessions all on Claude Code. There's a time and place for each tool. Also, thanks to my Neuralink, I have Linear and Azure MCP's wired directly into my frontal cortex at all times. This makes it really easy to keep our docs up to date without lifting a finger.
probably a rare area I fully agree with HN on– the IP here seems weak and it's not hard to swap out code editors, nothing like tearing out Salesforce or other sales-driven tooling. and idk if first mover advantage actually means much in the next 10 years given how dynamic the underlying models are.
but undeniably these cos are all a great lesson in just how much cash lies in executing first/near first
the hn cynicism is crazy but unsurprising. turns out can really easily spend your whole life obsessed with a pragmatism that betrays taking any action, however directionally correct
Love this. I’ve worked on a few projects in RPA prior and I’m losing faith in selectors. I think either direct data access like this or AI based CV are the automation arms of the future.