Opus 5 ARC-AGI-3 likely benchmaxxed(xcancel.com)
xcancel.com
Opus 5 ARC-AGI-3 likely benchmaxxed
https://xcancel.com/quietnning/status/2080786711861407883
2 comments
I switched from 4.8 to 5 yesterday and I'd have to say that I've noticed an improvement, though not hugely perceptible, the model seems to want to do more work per turn, possibly a tiny bit faster, and little less verbose when it isnt necessary to be so. Anthropic says it is more token-efficient, I hope that's true because I intend to use it over 4.8 from now on. I haven't even tried Fable yet. Opus has always performed well enough that I don't really yearn for more.. unless it was the same price.. which it isn't so.. no need.
Benchmarks mean nothing anymore. I don't even look at them, especially the ones the companies release themselves. Independent benchmarking may still provide some value but even then, I doubt it.