It is surprisingly hard to do in a single prompt. I’ve had good luck though with being very explicit about asking for review, then to think deeply about what actually matters, and then reduce to a very short extremely concise reply. —- the model NEEDS to output those pages of text as part of its thinking process. It’s used to doing it in the assistant reply but can be cajoled into doing it in the thinking tokens.
I’ve found that doing the full requirements capture, planning, writing, reviewing, gardening loop with frontier models has worked quite well since last October, and phenomenally since Fable 5.
The key I’ve found is human peer review. The reviewer jumps on a live call with the developer, pulls up the PR with transcription on, and asks questions. At the end of the call, the transcript passes back into the coding agent and the PR is polished up, becoming more self-documenting, and the humans are left with some degree of common understanding of what’s going on.
I’ve been operating my team of ~15 this way for 9mo to great effect… there is simply no going back to the stone ages.
Hmm I imagine using a server to connect to signal/whatsapp or even email, then using a local model to classify and filter and trim messages and forwarding to SMS, and viceversa. I guess the trouble is I’d need many source numbers :thinking:.
Can you share an example? I've been happily using Fable this afternoon and it just seems like the usual upgrade so far with no interruption to my (fairly standard) SWENG problems.
One capability that I see is missing from opus is this ability to understand an entire system. My hope is that a mythos class model will be able to comprehend even something as complicated as an IOT system with a hardware and firmware layer multiple API’s backend and different kinds of API and web clients.
The main limitation we’ve had to agentic coding is an understanding of this system that spans processes running on different machines and architectures.
Come on, Anthropic ARE the good guys if there are any. Certainly the incentives of trillions will do what money does, but they have assembled an incredibly altruistic and philosophically-minded crew. I’m rooting for them and trying real hard not to get cynical.
(Actually Englewood, NJ)
Interests: AI/ML, Books, Education, Entrepreneurship, Hardware, Hiking, IoT, Investment, Legal Tech, Mentorship, Music, Philosophy, Outdoor Activities, Yoga
---