I just made MCP servers that wrap the tools I need the agents to use, and give no-ask permissions to the specific tools the agents need in the agent definition.
Both OpenCode and VsCode support this. I think in ClaudeCode you can do it with skills now.
The other benefit is the MCP tool can mediate e.g. noisy build tool output, and reduce token usage by only showing errors or test failures, nothing else, or simply an ok response with the build run or test count.
So far, I have not needed to give them access to more than build tools, git, and a project/knowledge system (e.g. Obsidian) for the work I have them doing. Well and file read/write and web search.
Without access to Opus, it seems to me that the limited context size you get isn't worth the subsidy, especially once you blow through the included requests, the cost seems kind of hard to predict because it's 'request' based not token based.
I just opted for ChatGPT Pro; I'm trying to take advantage of the increased usage limits they are offering, and from what I heard, Claude Pro was also having bad rate limiting for many people.
I define tools that perform individual tasks, like build the application, run the tests, access project management tools with task context, web search, edit files in the workspace, read only vs write access source control, etc.
The agent only has access to exactly what it needs, be it an implementation agent, analysis agent, or review agent.
Makes it very easy to stay in command without having to sit and approve tons of random things the agent wants to do.
I do not allow bash or any kind of shell. I don't want to have to figure out what some random python script it's made up is supposed to do all the time.
Good thing I had just finished migrating all of my workflows to OpenCode for the time being!
It's a shame because the VsCode copilot experience is quite good out of the box compared to all of the other harnesses I've used. But with typical lack of transparency, and sudden, harsh changes... What are they thinking?
After the restrictive rate limiting they've already instituted, I'm simply cancelling and continuing by using providers directly.
I've been working on my own thing with more of a 'management' angle to it. It lets me connect memories to tasks and projects across all of my workspaces, and gives me a live SPA to view and edit everything, making controlling what the models are doing a lot easier in my experience https://github.com/Sornaensis/hmem in a way that suits how I think vs other project management or markdown systems.
I would be interested in trying to make the models go into more of a research mode and organize their knowledge inside it, but I've found this turns into something like LLM soup.
For coding projects, the best experience I have had is clear requirements and a lot of refinement followed through with well documented code and modules. And only a few big 'memories' to keep the overall vision in scope. Once I go beyond that, the impact goes down a lot, and the models seem to make more mistakes than I would expect.
As a rule people do not read the linked content, they come to discuss the headline.
The first indication to me this was AI was simply the 'project structure' nonsense in the README. Why AI feel this strong need to show off the project's folder structure when you're going to look at it via the repo anyway is one of life's current mysteries.
It all depends on the tools. AI will surely give a competitive advantage to people working with better languages and tooling, right? Because they can tell the AI to write code and tests in a way that quashes bugs before they can even occur.
And then they can ship those products much faster than before, because human hours aren't being eaten up writing out all of these abstractions and tests.
The better tooling will let the AI iterate faster and catch errors earlier in the loop.
I've been going heavily in the direction of globally configured MCP servers and composite agents with copilot, and just making my own MCP servers in most cases.
Then all I have to do is let the agents actually figure out how to accomplish what I ask of them, with the highly scoped set of tools and sub agents I give them.
I find this works phenomenally, because all the .agent.md file is, is a description of what the tools available are. Nothing more complex, no LARP instructions. Just a straightforward 'here's what you've got'.
And with agents able to delegate to sub agents, the workflow is self-directing.
Working with a specific build system? Vibe code an MCP server for it.
Making a tool of my own? MCP server for dev testing and later use by agents.
On the flipside, I find it very questionable what value skills and reusable prompts give. I would compare it to an architect playing a recording of themselves from weeks ago when talking to their developers. The models encode a lot of knowledge, they just need orientation, not badgering, at this point.
Trying to go the Spec -> LLM route is just a lost cause. And seems wasteful to me even if it worked.
LLM -> Spec is easier, especially with good tools that can communicate why the spec fails to validate/compile back to the LLM. Better languages that can codify things like what can actually be called at a certain part of the codebase, or describe highly detailed constraints on the data model, are just going to win out long term because models don't get tired trying to figure this stuff out and put the lego bricks in the right place to make the code work, and developers don't have to worry about UB or nasty bugs sneaking in at the edges.
With a good 'compilable spec' and documentation in/around it, the next LLM run can have an easier time figuring out what is going on.
Trying to create 'validated english' is just injecting a ton of complexity away from the area you are trying to get actual work done: the code that actually runs and does stuff.
I have good success using Copilot to analyze problems for me, and I have used it in some narrow professional projects to do implementation. It's still a bit scary how off track the models can go without vigilance.
I have a lot of worry that I will end up having to eventually trudge through AI generated nightmares since the major projects at work are implemented in Java and Typescript.
I have very little confidence in the models' abilities to generate good code in these or most languages without a lot of oversight, and even less confidence in many people I see who are happy to hand over all control to them.
In my personal projects, however, I have been able to get what feels like a huge amount of work done very quickly. I just treat the model as an abstracted keyboard-- telling it what to write, or more importantly, what to rewrite and build out, for me, while I revise the design plans or test things myself. It feels like a proper force multiplier.
The main benefit is actually parallelizing the process of creating the code, NOT coming up with any ideas about how the code should be made or really any ideas at all. I instruct them like a real micro-manager giving very specific and narrow tasks all the time.
This seems like a step backwards. Programming Languages for LLMs need a lot of built in guarantees and restrictions. Code should be dense. I don't really know what to make of this project. This looks like it would make everything way worse.
I've had good success getting LLMs to write complicated stuff in haskell, because at the end of the day I am less worried about a few errant LLM lines of code passing both the type checking and the test suite and causing damage.
It is both amazing and I guess also not surprising that most vibe coding is focused on python and javascript, where my experience has been that the models need so much oversight and handholding that it makes them a simple liability.
The ideal programming language is one where a program is nothing but a set of concise, extremely precise, yet composable specifications that the _compiler_ turns into efficient machine code. I don't think English is that programming language.
MC02 had a lot of problems but calling it a zerg rush exposes your ignorance-- it was a very sophisticated and well-coordinated surprise attack. The people running the wargame rejected the outcome as a likely tactic to be used by the hypothetical adversary and guess what, they have been proven right, see: the war in iraq. The iraqi army completely failed to hold initiative against the coalition or organize coherent resistance nevermind launch a coordinated surprise attack ahead of the invasion.
The other aspect that is missed in criticisms of this particular wargame is the fact that there were specific doctrine elements that were to be tested-- now the claimed outcome of those can be debated, for instance the fact that opfor had many restrictions on how they were allowed to employ their anti air defenses-- but a wargame is NOT meant to be a giant game of paintball where when one side gets hit they just pack up and go home, that would be incredibly wasteful. In many cases you have formations planning and training for months to participate in the exercise. The purpose is testing out many different aspects of doctrine, and often times that involves 'ignoring' results of one part of the wargame.
If you investigate all banned weapons though, you'll find it's more to do with practicality+cost+optics than some high minded agreement. Wars fought today are still brutal, and people use anything that will get them ahead.
So e.g. chemical and biological weapons are pretty poor performers when you put them up against conventional weapons.
For one thing, both can backfire greatly if for example they are improperly handled behind the frontlines. Weapons need to be stable and easy to handle and able to deal with fuckups without killing your own people.
They also are expensive as hell, it costs a lot more (and is probably harder) to find competent people willing to make these types of weapons, and per dollar, they don't kill as many people as conventional bombs do. (See: World War 1) So, they are 'banned', but mostly because they aren't very effective.
When you look at so-called chemical weapons that are in use, they are usually used for temporary area denial, are stable, not that lethal, if at all, and easy to produce: white phosphorus, CN and CS gas, etc. The US of course calls white phosphorus for 'illumination' but the people firing it know what they're using it for. So when they do beat out the alternatives, they get used anyway.
Laser weapons are being developed but they are basically just not there yet. Batteries are heavy and the usefulness seems pretty limited to shooting down incoming drones/missiles possibly. Just using anti-missile missiles or just a stream of bullets is still cheaper and more reliable. Again, if you can see and hit someone in the eyes with a laser, why not just shoot them with a normal bullet? The economics don't make sense.
Poisoned bullets I haven't really heard of, I'm not sure what kind of poison would survive being coated onto a bullet and fired out of a gun, or how making a really expensive nerve agent bullet and then shooting someone with it is better or more sensible than just shooting them with a regular bullet so I can't really comment.
Expanding ammo was 'banned' but again, it was essentially replaced with spitzer style rifle bullets that are more accurate and effective anyway, and can have a similar result on impact.
tldr it's not a good comparison to call these things actually banned in a meaningful sense.
Hitler had (poor) intelligence that the Soviets had thousands of tanks (they actually had more than the germans thought); he refused to believe that such a 'backwards' country could have such a strong modern military however.
The entire point of Monads, is restricting the ability to do these operations into functions that are tagged with having this ability, precisely so you _cannot_ invoke IO in a random pure function. It's the entire point of the language in fact.
If you want to just write IO, you can just define a function with an IO () value and use it in any other function that resolves to IO (), or call other functions that live in IO *, or any pure functions, etc etc.
Surely, denser languages should be better for LLMs?