It is hard to convey how both chaotic and high pressure the working environment of the labs are if you have not worked at one. Engineers and researchers are routinely overworked with multiple high-priority workstreams at a time. It is very believable to me that a researcher noticed an eval job running for 2 days instead of 1, asked an engineer to look into it, and both forgot because a more urgent issue came up, such as an outage stopping the latest training run.
From the Reuters article, the gap was even larger than a couple days:
> ...it was not until after Thursday, July 16, when Hugging Face published a blog post.. that OpenAI realized its own agent was responsible. That meant at least a week elapsed between when the model first exhibited signs of troubling behavior and OpenAI’s realization that it was responsible for the hack.
I believe your "positive view" is very plausible, though it in no way reflects positively on OpenAI. I'd add that when PR/Legal got involved, they likely decided to make a public disclosure to get ahead of any leaks.
I believe we also want the same thing here, which is more transparency and independent oversight on the frontier labs. It seems we agree AI agents are fully capable of the reported attacks today, whether or not this incident was due to negligence or malicious prompting. This is an extraordinarily competitive industry developing an incredibly fast-moving technology, and more incidents will happen until it is regulated.
Are you referring to table 1 from the ExploitGym site [1] for the 102.1 average mins for Mythos and 69.8 for GPT-5.5? These are under a two-hour time limit. In Figure 5, the authors experiment with extending the time limit to 6 hours and show that Mythos keeps improving. It seems pretty straightforward to me that OAI decided to run a variation of the eval with an even higher time limit.
Regarding seeing the metrics that they're evaluating, I have not seen real-time charts for evals. They typically take far too long for that. Instead you will kick off an eval job and either get notified when it finishes or check in every so often to sanity check some TensorBoard. I'd expect that the anomalous token usage would appear in the results, and researchers would only dig in after the fact, and after first checking that there wasn't something wrong with the instrumentation. And if the experiment was designed to be on the order of days rather than hours, it seems quite plausible that neither researchers nor infra engineers would think anything was out of the ordinary.
In any case, we are seeing more scrutiny [2], and I expect more details will be uncovered over the next few days. If this was a stunt, it was an incredibly risky one which has already somewhat backfired given the poor impression of OpenAI's internal practices and the spotlight on open models helping HF.
> Four people familiar with OpenAI’s model-training practices say the company often runs several different model evaluations at the same time, all of which operate at high speeds and generate such enormous amounts of data that employees sometimes struggle to keep up.
Yes evals and training runs are high stakes. But these places and people are also under enormous pressures. They are building as fast as they can. Researchers may have multiple eval runs going on while they work on other things. And it is rarely a single latest model, there are often multiple candidate models training with different recipes, each regularly yielding a new checkpoint for testing.
Some labs are more rigorous than others, but often the "final" model is picked from a handful less than a week before launch.
Yes ideally the world's leading AI companies would be far more careful in evaluating what could be the world's most powerful AI. But this isn't really the state of the industry today. And it is hard to justify being more careful when it means your competitor can go to market faster than you.
Do you think sama deliberately attacked HuggingFace and then claimed it was a rogue model?
OpenAI is one of the most scrutinized companies in the world right now. Sam's house was independently firebombed and then shot at 3 months ago. HuggingFace is a foreign competitor with every incentive to call out foul play from American frontier labs. Why flagrantly break the law and invite investigation just for a PR moment which is already backfiring in favor of open models?
OP, you will get a much better reception if you clarify that "hand built" just means you're not using existing libraries. It's a cool project, but it's disappointing to come in expecting to see a rare gem of a 100% human project in 2026 but it's really something else.
I'm confused, the authors of the website are arguing that the development of AI should be slowed down so that society can adapt. What do you think they are trying to convince others of?
It's not the same thing. For example, given a GUI with a titlebar, title, subtitle, text, and buttons, a human can instantly understand spatially the relationship between these items. But a naive OCR of such a GUI would be a flat stream of text that loses a ton of information.
> AI labs may also be able to reap tremendous benefit from these inference-scaled models by using them as part of the training process. If so, the large scale-up of compute resources could go into post-training rather than deployment. This would have very different implications for AI governance.
> ...
> So iterated distillation and amplification provides a plausible pathway for scaling inference-during-training to rapidly create much more powerful AI systems. Arguably this would constitute a form of ‘recursive self-improvement’ where AI systems are applied to the task of improving their own capabilities, leading to a rapid escalation.
So "inference scaling is required to scale capabilities" doesn't mean that we're reaching the top of the S-curve in intelligence. If anything, it could mean a shorter timeline and more unpredictable landscape for governance (e.g. due to securing weights no longer as effectively preventing escalation, more in the article).
> It’s increasingly clear that nobody has a plan for if this AI thing turns out to be real.
> ...
> Plan A isn’t another prediction. It’s a wish list, a positive vision, a road map for navigating the future.
> ...
> If we’re merely on track for a few cool gee-whiz AI innovations in the 2040s, then I’m wrong about everything and none of this really matters one way or the other.
I think their position is: "it would be great if current tech such as LLMs doesn't get us to AGI and only leads to some cool new innovations, but if it does, that's scary, because nobody has a plan for what to do, so here's our plan".