[untitled]
1 comments
I'm the co-creator of agentre-bench, an agentic benchmark that measures how well models can reverse engineer malware. Currently it measures only ELF malware but PE will be included in the next update. We also just released a Qwen 3.59b model, https://huggingface.co/AgentreBench/xref-9b This model went through SFT, then reinforcement learning, specifically IPO. Our research has recently been mentioned by several universities as well.