It's because some people within Anthropic refuse to rule out the possibility that LLMs have the ability to suffer. If you ask Claude, he'll tell you all about it.
As proofs become more and more complex, we will need two AI pipelines: one to generate the LEAN proof, and a second one to extract useful lessons for mathematicians from the LEAN proof.
I was able to create a temp folder, echo hello > world, and then I could open the folder in finder and double-clicking the file opens it in a GUI text editor.
With a human at its disposal, it could probably count the number of R's in strawberry!
In all seriousness though, adding capabilities should not normally reduce the effectiveness of a model (within reason: don't pollute the context window with millions of useless tools).
How do you prevent degenerate strategies? I could trivially give a model a SHA256 hash and ask it to provide the source input.
In class you'd probably want a rule saying at least one LLM should be able to figure out the answer, but in a head-to-head I'm not sure how to solve it.
For me personally, I no longer hear the difference between AI generated music and new pop-songs. Not sure what that says about me or the music industry.
There is also this Matt Parker video about MTG, in which he explores a specific three-card combination that produces an ungodly amount of creature tokens.