I think this is such an nefariously unnecessary negative argument.
Most, if not all, of the shai-hulud attacks that hit npm and other ecosystems were preventable with cooldowns. And these were not detected because regular users reported the worms, but because security researchers did. I don’t think I’ve ever seen an attack that was discovered because a user reported it.
As an author: show me why you thought this was interesting and why you’re doing it, and why you think it’s relevant. What does it build towards? What does climbing this leaderboard mean to me?
Absent those things, this is just some thing my opus could generate as well.
I think the idea is interesting though, although I wonder if training time for LoRA is such a bottleneck to deserve its own, extremely narrowly scoped, leaderboard. Maybe if it was more tasks or more models we could hope that it transfers? With a single task, and a single model, I’d be afraid of this overfitting pretty heavily.
For NanoGPT, I think the idea always was that the ideas can be transferred to much larger models, or serve as stepping stones for investigations on larger models.
Your fully generated project description makes people lose faith in the actual project. If you went through the trouble of writing the whole code yourself, why generate the comment presenting it to the public wholesale.
I think he is a good example of someone who writes mainly to show he is ready for the next rung of the corporate ladder. That is, his posts are not meant to be useful, but to show higher-ups he is useful to them.
As mentioned by a sibling comment: this is an insensitive take.
It takes a lot of courage to write down one’s struggles for all the world to see. Your analysis denies the OP their self-reflection, and instead reduces it to a thing you happened to find in your own life.
I am interested in why you chose to do this, and publish it with the headline you used. Was it to learn something? Or to get publicity for another project?
Tbh, this sounds like fear mongering to me. Of course the statement “99.9% of servers are not compliant” sounds impressive, but then it turns spec hasn’t even been released yet.
Also some general feedback: the whole thing looks generated, as does the comment I am replying to.
It’s personal ad, basically. The author is trying to get a job as an evaluator somewhere and is hoping that putting 1000$ on the line will get them enough publicity to land them an interview/get a job somewhere.
The second part of this comment is not what I expected. I also don’t think it is true. I got bit by a CORS error at work recently that passed by Claude, copilot, and another senior engineer.
We’ve been on the receiving end of this complaint with Semble. I think it is a valid complaint, but constructing a benchmark for this kind of thing is just very difficult and expensive because of the (harness) x (model) x (mcp/cli) combination.
With traditional ml/tooling, not showing benchmarks was usually a red flag. But for llm tooling, I’m not so sure.