GLM-5 topped the coding benchmarks. Then I used it(charlesazam.com)
charlesazam.com
GLM-5 topped the coding benchmarks. Then I used it
https://charlesazam.com/blog/glm5-benchmark-reality/
1 comments
TL;DR: GLM-5 tops coding benchmarks. I tested it on an unpublished NP-hard optimization problem (KIRO) and 89-task Terminal-Bench. Best case: competitive. Typical case: 30% invalid output, every trial timed out, and two identical runs could produce a valid solution or complete garbage. Zhipu AI reports 56% on Terminal-Bench; I got 40%.