DeepSWE crowns GPT-5.5, and finds Claude Opus exploiting a benchmark loophole(venturebeat.com)
venturebeat.com
DeepSWE crowns GPT-5.5, and finds Claude Opus exploiting a benchmark loophole
https://venturebeat.com/technology/deepswe-blows-up-the-ai-coding-leaderboard-crowns-gpt-5-5-and-finds-claude-opus-exploiting-a-benchmark-loophole
0 comments
—