Handwritten-edit benchmark: Fable 5 is #1, Opus 4.8 regresses 55% on miscounting(dorrit.pairsys.ai)
dorrit.pairsys.ai
Handwritten-edit benchmark: Fable 5 is #1, Opus 4.8 regresses 55% on miscounting
https://dorrit.pairsys.ai/
https://dorrit.pairsys.ai/
I hadn't updated the benchmark in several months, but there are some interesting findings. Fable 5 takes the top spot (0.6579), setting a new performance record, while Kimi K3 is within a hair's breadth of its performance.
The most significant finding is that Opus 4.8 regresses drastically compared to Opus 4.7, from 0.4805 to 0.2150. This seems mainly due to a regression in its ability to count line numbers, and it's something you might want to keep in mind when designing your own agents.