Agent loops are you give it a task and it tries its best to finish it.
The only results are success and failure. If it's a success you go through the logs to see the actions it took and if it's a failure you do the same thing.
Why would you look at it in realtime when the whole point of agentic work is to get them to run autonomously as long as possible?
a massive number of data engineering agents. We are able to stand up, test, audit, deploy data pipelines much more efficiently with the foundational models.
I don't write code anymore and I doubt I ever will ever again.
On the flipside I review exponentially more code than ever before.
>Just ask Claude to dump out assembly, or a compiled binary, but no, they don't trust the LLM that much
It's not "not trusting" the llm its that the llm has been undergoing reinforcement learning is on coding. Plus generating assembly is extremely token inefficient.
People don't like to hear this but the open models just aren't good for end to end agentic workflows.
There are some very very good small open models that can excel in certain finite bounded tasks, but the foundational models are essential to building out agentic pipelines that actually work.
>And most of the time I see some basic workflows. Summarizing Slack. Answering emails. Doing scheduled scans. Performing research and booking something. Sending emails out of Claude.
This alone will improve the lives of so many people. The real issue with AI commentary is that everyone is guilty of the hedonistic treadmill. We constantly need it to do more and more to get that sense of awe.