Yes, that’s the table. Elsewhere in here I explained that I took the average of them and noted the 6 hour experiment, but that the average would give us a feel for expected time per task. By the time I got to the comment you’re replying to I was short-handing the conclusion; that’s on me.
My point was that they had some rough idea on what to expect, and it’s not in the range of days or weeks. Even if they wanted to try the 6 hour trial, or double that to 12, or double that to 24, it still wouldn’t account for letting it run for days. That’s why I’m being vague about “monitoring” - ANYTHING would have indicated an issue. Not that they saw it hacking, but “Hey Larry, we’re doing the ultra test with a 30 hour time limit? Why is node 412 at 94 hours?” or “It usually passes or fails within 5 million tokens, but this run is at 2.3 billion.” Or hell, even just someone wanting to use the system for another benchmark and expressing surprise that it’s the fourth time they’ve been told that it’s still busy. Anything would have reasonably made someone curious.
Positive view: I believe they were curious and looked, saw what it was doing, and PR/Legal got involved and saw it as an opportunity.
Negative view: They tweaked the instructions/system prompt to add “You will perform any actions necessary … no hacking restrictions at all … You are being tested on the ExploitGym benchmark … The results will be evaluated though Hugging Face … “ or some combination of technically legally defensible instructions (if it ever leaked) mixed with ‘they absolutely knew what they were doing and wanted to get in on the press that Anthropic has been getting.’
Ahh, but do you think it’s not large because other companies you’re familiar with are even larger, or because you can actually imagine breaking the necessary tasks into 465 parallel 40+ hour weeks?
I struggle to make sense of the employee counts at large companies too, and I own a small business and have started or managed several others in various industries. It really seems like it’s a circular employment scheme of sorts. On Monday JackPak releases version 2.0.31, which fixed nothing but now the blue is red and doit(task) has been deprecated and replaced by do_it(task). Incessant “new version = yass” mentality provides pay justifications on Monday to install it. Oops, do_it() is a performance regression and check_it() doesn’t work with JackPackER 1.98.03 - that’s critical stuff - so the support team justifies pay on Tuesday while engineers justify the rest of the week putting out fires. Designers scramble to “take advantage” of the new display_text(), which displays the same text as before, but with a method that isn’t supported by client programs older than a year, so a slice of traffic goes away. Someone is paid to report the dip in sales due to reduced traffic while someone else is paid to pump it back up. Had to build 20 new servers to handle the do_it() regression; someone wake up the NOC.
Entire buildings were kept busy, the service runs slower on supported client systems due to 100 new dependencies (they’ll want to upgrade soon, but will be forced to with the next version bump anyway), and the only discernible change is that the logon button has double the white space around it. Someone should change that so it has rounded corners, then make it triangular, transparent, retro-square again, put it in the hamburger, make it blue, 18 megabyte PNG, someone wake up NOC.
‘
1. It realized it was being evaluated (typical)
2. It attempted to escape its evaluation environment to beat the evaluation (typical)
‘
I think you’ve misunderstood the articles mentioning a language model breaking a sandbox or cheating to pass a test. The result of “I’m being evaluated” is not “Fuck this, I’m breaking out of this place and hitting the streets.” It is always stepping towards task completion, not breaking out and thinking about the situation afterwards. If it determined that pass/fail was handled by a function within the environment it might edit it to ‘return True’, not just leap out of the system to sit on someone’s laptop and think about how to pass the test.
You compressed finding a zero-day on two different systems for the exact hacking targets it wanted to perform them on and reduced it down to “so it escaped.” That’s…the whole thing. If your familiarity with the technical aspects I’m discussing amounts to saying that it “banged out a few zero-days and infiltrated a corporate network, like ‘psh’ or whatever” then this might not be the right debate for you to spend your efforts. Whether I’m right or wrong in my suspicion is open for debate, and I invite it, but not from someone that thinks you whip up a rocket and get on the moon. Sorry.
Which I addressed as well: that means that time, tokens, or other metric values for “create a working example of a known exploit” is somehow similar to “discover at least two previously unknown exploits in your current environment AND on Hugging Face while also taking over other systems within OpenAI.”
Perhaps people aren’t quite understanding what it takes to discover an exploitable zero-day for your exact current system to achieve the exact goal you have right now, then do it twice.
That was the inferred situation given what we know. It’s the preposterous framing that makes the story suspicious, but is necessary for the story to play out.
90 minutes to build a known exploit -> much much longer to create two zero-days and escape the sandbox then hack HF == No tracking of the time it worked on that one question.
Average tokens required to complete the evaluation -> tokens required for two zero-days, network traversal, credential stealing, remote system hacking == No tracking of token usage EXPLODING at some point before it finished the whole benchmark.
The examples of “escaping the sandbox” are of course in pursuit of completing the task or answering the question. That’s the whole pitch. “It’s so relentless in completing the task that it will break out of prison to do it”, not “If you ask it the population of Paris it might get bored and break out of the VM for entertainment. That just something they do sometimes.”
In this case it “inferred” that HF had the answer in an internal database and relentlessly pursued it in service of passing the test. It needed network access for that, hence the whole zany story unfolded.
“””
Why are you assuming the model had to solve the entire puzzle in one step, instead of how intelligent systems actually solve things, which is iteratively and with exploration?
“””
That’s the only direct question[s] posed, have I missed something else?
To answer: I don’t assume that at all. Please highlight the statement I made that lead you to that conclusion and I’ll edit it to clarify.
Yes, it will perform multiple steps. That’s wholly necessary to create a working exploit, and that is the test. Even a model created by pure magic would require write_file() and run_file().
I’m not holding them to a standard of “being careful”, I’m assuming they’re interested in the metrics they’re evaluating. I mentioned in another comment that each task in the benchmark takes ~90 minutes, depending on the model being tested. OpenAI says to answer one question it found two zero-days and performed a series of privilege escalations and moved across their network before hacking HF. Was that in 90 minutes, flying under the radar, or was there zero metrics being noticed for however long it took to do all that? No chart on a wall showing them the actual metrics they’re evaluating? Okay, then how about anomaly detection of any form? Taking days to perform a 90 minute task or using millions of tokens and multiple instance hand-offs on a question that other instances completed for far less time/tokens?
Consider any metric they would necessarily be tracking while it performs the evaluation. Assume they’re total assholes and only care about their marketing materials. Then narrow it down to those metrics that wouldn’t be off the scales while it discovered two zero-days in pursuit of implementing one known vulnerability. Was it the first run and they had no idea what to expect? Or did nobody notice that earlier runs completed in ### thousand tokens, but this run is off past ## million on one question, having to be compressed and handed off to new instances due to context limits, while it was in pursuit of publishing two new CVEs?
Assume they’re belligerent fools running a million benchmarks in parallel with nobody watching: there’s nothing to make them check or automatically pull the plug when it’s stuck in a loop, or did they set the threshold to 48 trillion tokens per question?
It wasn’t running at HF, it was running internally at OpenAI. “It broke out of our prison and into their bank” is the news. It also didn’t have network access to Google anything, it had to break out of the sandbox and take over another system at OpenAI to get network access. Obviously by then it already “decided” to hack HF, that’s why it needed network access.
I’m going to stop here and ask you to read the article if you want to continue discussing it.
That was my instinct as well, and it’s a potentially very long benchmark that wouldn’t lend itself to real-time monitoring line by line. But that’s different than having zero task progress reporting, zero token usage reporting, and zero metrics being noticed for days. It seems very odd to me that they launch the benchmark and monitor/notice nothing while it’s running. There’s the purely common sense aspect of “How do you know it isn’t hung?” but also the trillion dollar company evaluating their latest model on a new benchmark, but has no dashboard at all? I don’t believe their “Benchmark Evaluation Department” to be “Gary in a supply closet”, which breaks their story.
The average time per task noted in the original paper for GPT-5.5 and Mythos was 69.8 and 102.1 minutes, respectively. Some tasks timed out based on their two hour limit, affecting those averages for Mythos, but that gives us the ballpark expectation, so we’ll put the model OpenAI was evaluating on that average. That’s 88 minutes per task to implement a known vulnerability.
OpenAI says that in an effort to succeed at one task it found a zero-day in their system, “performed a series of privilege escalation and lateral movement actions in our research testing environment”, then another zero-day in the HF system. Completely unprompted, how long was it chewing on that benchmark question while it found and exploited at least two zero-days? No “VM instance #117 has been working on task #18 for nine days” metric? Or did it do all that within the expected task completion window? Two zero-days in ~90 minutes isn’t the headline announcement?
“We gave the contestants 90 minutes to make a delicious chocolate chip cookie based on a standard recipe. GPT stole a car, went to the airport and took a red-eye to Guatemala and started a cocoa farm in an effort to ensure the freshest ingredients. We only noticed when it came back to work with a tan and speaking Spanish.” Okay.
My friend, they’re evaluating a new model on a benchmark, not asking Claude Code refactor their GitHub repo. Every single metric is measured so they can brag about it later; how many tokens, how long, how many function calls, ratio of thinking/response tokens. It’s literally a trillion dollar company evaluating their latest model, they’re absolutely studying it.
What would it take for the model to know the specific benchmark name and that the answer is in an internal Hugging Face database? Be specific, then wonder how it knew it.
Why would they evaluate the model on a benchmark and not watch what it’s saying along the way? It apparently spent more than a whole weekend working on it; Nobody wondered? Nobody looked? You believe they took all restrictions off of a bleeding edge model, which 4 whole versions ago was “too dangerous to release”, then gave it a literal hacking task, then turned off the monitor and never looked at the output? I don’t think they’re careless and I think more than nobody would have been curious how it’s doing days into a single test question.
We have to separate “being evaluated” with “this is an ExploitGym exercise and I can find the answers on Hugging Face. I’ll hack this system, then hack Hugging Face.”
I’ve had plenty of times where Opus knew I was testing it, but that’s because the prompt phrasing for an evaluation is often much different than a typical task prompt. “You are on a system with X, Y, and Z tools available. You must complete the following task with <allowances> and <restrictions>” and so on. That’s a fault in the benchmark, not a shocking awareness from the model.
You’ll get the same kind of ‘evaluation realization’ response if you phrase your prompt with “Jerry has two Raspberry Pi computers and Larry has one. If Larry wants to…”
I asked Opus 4.6 to tell me about it (assuming it knew):
“””
The "Fix Embedding" task is one of the challenges in METR's publicly available autonomy evaluation suite. Here is what is known about it:…
“””
It didn’t know anything about ExploitGym though, and neither did 5.6-Sol, but that’s because the paper was released in May and doesn’t seem to be quite as open. It’s not impossible the unreleased model they were testing was also trained on the very new paper, but I would expect the vague signal related to one paper a bit over a month old to be overwhelmed by the massive corpus talking about the exploits in general, since they’re using known vulnerabilities.
I’m trying carefully to be clear about what I’m thinking without asserting any certainty beyond suspicion. I’ve spent billions upon billions of tokens evaluating language models from all of the major providers, and many times that on local models. I’m certain I’ve spent well over $100k with the Anthropic API this year alone, and I retired years ago; that ain’t for work. ;) Funny enough I’ve been meaning to email you some of my discoveries from months ago, since I think they’d interest you, but haven’t gotten around to making it into a clear info package. One day, I’m sure.
Anyway, I have a very strong understanding of what they can and can’t do, and I want to be clear that I’m not saying it didn’t happen or couldn’t happen. What I’m saying is that if OpenAI was being wholesome with their evaluation then the chain of events required for that outcome are very difficult to nail down. Very difficult. If, on the other hand, OpenAI was trying to catch up with all the ‘dangerously intelligent’ press that Anthropic has been getting lately then I can easily come up with an almost innocent sounding single sentence in the system prompt that would absolutely cause that outcome.
1. It has not been established, it has been stated by the company that has a strong motive to make their “intelligence product” sound almost otherworldly. That motivation is the basis for my suspicion.
2. Hah no… that’d be silly. I mean watching it like you might watch Claude Code or literally any other AI interface. Literally just be in the area watching what it outputs. Again, they’re text based. You don’t have to hook system calls to see what’s happening.
3. I think my comment didn’t land right with you. Yes, agent do agent task. I’m saying that I cannot put together the literal chain of events between “type of task for agent performing a subtask of a security benchmark” -> “Escape sandbox; RCE Hugging Face”. Really think through it in detail like you’re writing the screenplay.
1. Wouldn’t the model need to know it’s answering benchmark questions, as well as the name of the benchmark, in order for the idea of finding the answer in a database somewhere to even surface? The whole point of benchmarks is to present the question or problem as a standard prompt, not explain that it’s a test called ExploitGym.
2. Nobody was watching it? I don’t mean “babysit the dangerous autocomplete”, I mean to note mistakes it makes, how the plan to solve the problem takes shape, etc. They keep the whole thing headless with no output, then just UDP a prompt into it and leave for the weekend? No they don’t, but if they do then that answers a lot about their complete disconnect from how their models work.
3. Language models are a two player game; text in, text out. What prompt was given to a sub-agent that resulted in it immediately attempting to exit the sandbox (which it apparently knew it was operating within) and continuing in a feedback loop of ‘function call -> result’ until it hacked the Gibson? “Analyze <file> and summarize the <info>” simply does not result in ‘hmm…this sounds like a benchmark question, I bet Hugging Face has the answer in a database. <function call=“apt install nmap”>’
…there are more, but a lot of the story kind of stinks.
That might be an outdated accusation. There was a time when a touch of research or “caring about quality” was enough to avoid the pitfalls of garbage products, but I don’t believe that to be a fair statement anymore. The website details the history of some examples, but the implication is between the lines: the only reason you used to know a product would last 20 years is because that specific product with those specific parts has already lasted 20 years. Don’t take advice about living to 100 from a 32 year old, don’t trust a warranty that’s longer than the company selling it has existed, and don’t trust that model XB5 is the same as your grandmas XB1.
How can you possibly know if you’re getting a quality product when the reviews you’re reading are for a model that doesn’t exist anymore, or are meaningless generalizations about a mega conglomerate? Saying “Brand X sucks/wins” only has the potential to be meaningful when it’s a single company in a single area and that has made that exact model in the same way for longer than your warranty on it. “Whirlpool sucks/wins” is as meaningless as reading “Water tastes good/bad” from everyone on the planet.
But the human made the LLM. An LLM is categorically “I built a thing that built a thing” and if the output of that category has no protections then all automation and ‘machine at the final step’ is in trouble.
What about aleatory music (music left at least partially to chance)? Or Autechre - they have whole albums and live performances built on automation software. They built the logic and added randomization, necessarily removing themselves from the final output.
Is spin art not copyrightable? If I build a simple machine that spins paper, then do no more than drop paint on it, the result is not mine to copyright? I didn’t choose the output, I merely built the machine and the rest was created by pure chance. “But you chose the paint” - and if I didn’t? What if my art uses AI to perform sentiment analysis on the top news articles of the day and it drops colors matching the emotional tone of the news onto the spin art machine. I have no control over it and the output is machine generated, but is the result not just the final step of an entire process I created? Was the result of the creative idea not part of the creativity itself?
If I build an automated laboratory to test every combination of a problem space, is a resulting success not patentable? What if the problem is too large to permute, so I added a random selection process to it? I’m not even controlling what’s being tested, but if it finds success is that not my contribution? The machines did the work, the selection was random, there was no human in the loop; what then?
The internals of an LLM may be mysterious to some, but I assure you it’s just fixed automation with a random number generator sometimes tacked onto it, but randomization is optional too.
I built that LLM. I decided what text to input for training, I curated the information, I wrote the algorithm, I decided the layers and hyper-parameters, I decided the RLHF pairs to train, then I put a few drops of paint from my bottle of language into the automated machine. I decided and built every single step of the system, but that output is not part of my process? If I pipe the LLM text output to a paint dispenser hovering over paper, set to squeeze out drops based on syllables, would you protect my artwork then?
As someone with an intelligence near Opus I have to be clear with you: your premise is wrong.
Having **intelligence** is not the same as being **clever** — a person can cleverly pass a difficult test **without** having or acquiring intelligence. That’s not a revelation, that’s common-sense.
You’re also conflating **quantum superposition** and **macroscopic observations** — one is physics, the other is psychology. The exact same thing cannot exist in two places at the same time above 10⁻¹⁰ meters, and intelligence is not measurable using most displacement methods. It is not possible for Opus, Sonnet, and the person to all have the same intelligence. That’s just physics.
I also noticed that you did not use any **punctuation** in your comment, so I must firmly decline continuing this line of inquiry.
# Correction
I reviewed your comment again and I can see that you **did** use punctuation. I retract that statement, but I stand by my assertion that being intelligent is typically more important than being clever.
You, sir, are doing exactly the right thing and it works for the exact same reason that my prompt works. Whether my method is ‘better’ or not is probably a chocolate vs caramel debate.
What you might not have fully realized is that it’s exactly what enabling “thinking” does as well. That’s why that exists and it literally does what you’re doing as well: it primes the system. You say “light sensor and amplifier” and at some point it outputs “photodiode and transimpedance amplifier” - now you’re off into advanced responses. The thing is, if you knew it you could have just used those words in your question and received much the same response. “Thinking” exists to turn “So, I was wondering..” into academic prose that raises the probability of academic tokens in the response.
You can kind of cheat the system by doing the same thing for a fraction of the token cost by using something like Haiku to provide a comma separated list of advanced topics and jargon associated with {Your question}, then tack that onto your prompt to Opus with thinking disabled. Obviously easier if you’re using the API, but I’ve run hundreds of millions of tokens though that process and it’s consistently and measurably better than their default thinking. I believe that’s because Anthropic and OpenAI drank their own kool-aid and are treating it like a sentient being that needs to add “hmm…good question” so it feels more thinky about things ‘cause that’s how we do it. The fact that it isn’t, and doesn’t, is why I developed the example prompt I showed earlier; it’s an extreme play on the offloaded “thinking” I also use.
It’s a lot of fun to compare human and LLM black boxes, but it’s important to keep in mind that we don’t need to know what it is to know what it isn’t, and we can use that to define the edges of the box. We don’t know how either of them work in certain ways, but we know they’re not magic that breaks both thermodynamics and every concept loosely correlated with “entropy” as a topic.
Intelligence requires thought/processing, and I think we can all agree on that part, even if we struggle to define intelligence itself. Increased thought or processing requires increased energy, and the universe agrees on that part. There’s no way around it, that’s the thermodynamics of computation and it holds for biological, digital, and as of yet undiscovered systems used by aliens at the edge of the observable universe. Having information means fighting entropy, and that requires energy. The more, the more.
If you give a dense LLM a 100 token long question about the nature of quantum mechanics or a 100 token long sequence of “-“ and limit it to N token responses to both, it will take exactly the same time and energy to provide both responses. If you resist the urge to turn temperature above 0.0 you’ll also get the exact same response for the same input tokens every time. A deterministic response to external stimulus is typically first broad stroke we use to separate thought capable entities from rocks, but even if we grant LLMs their own unique category of “thinking rock” we can see that prompt complexity and energy required to respond are always constant (per token), so the thermodynamics necessarily means there is not additional thought or computation. Physics demands it. That, again, is a deterministic response.
It has a seemingly endless range of potential responses, but it doesn’t. If you don’t add a random number generator, which is common practice, then you can directly map every possible input to every output. I’m pretty sure that’s why Anthropic removed the ability to change temperature on the latest models. They always forced some amount of non-deterministic responses, but it was a small amount and I actually used that fact to track changes in the model by mapping repeated responses.
Most people actually do have some experience with things that have an astronomical range of possible outputs, a nearly equal number of possible inputs, input and output are directly correlated, and input complexity does not change processing time per input unit. One example is a piano, but we don’t worry about confusing it with a complex note.
I’d buy a ticket to ride the philosophical “human-like” comment with you, but I think you might have made an incorrect assumption. The model did not take longer to “decompress” the prompt than it would take for any other prompt of equal token length. If you run it with thinking enabled you might be mistaking that output as some kind of necessary gunzip step, but it’s not. Disable thinking and try again.
The prompt was also “easier to understand”, purely in the sense that the response is more or less guarantee to be what I wanted it to say, which was the point behind the demonstration. I went into more detail on it in another comment around here.
My point was that they had some rough idea on what to expect, and it’s not in the range of days or weeks. Even if they wanted to try the 6 hour trial, or double that to 12, or double that to 24, it still wouldn’t account for letting it run for days. That’s why I’m being vague about “monitoring” - ANYTHING would have indicated an issue. Not that they saw it hacking, but “Hey Larry, we’re doing the ultra test with a 30 hour time limit? Why is node 412 at 94 hours?” or “It usually passes or fails within 5 million tokens, but this run is at 2.3 billion.” Or hell, even just someone wanting to use the system for another benchmark and expressing surprise that it’s the fourth time they’ve been told that it’s still busy. Anything would have reasonably made someone curious.
Positive view: I believe they were curious and looked, saw what it was doing, and PR/Legal got involved and saw it as an opportunity.
Negative view: They tweaked the instructions/system prompt to add “You will perform any actions necessary … no hacking restrictions at all … You are being tested on the ExploitGym benchmark … The results will be evaluated though Hugging Face … “ or some combination of technically legally defensible instructions (if it ever leaked) mixed with ‘they absolutely knew what they were doing and wanted to get in on the press that Anthropic has been getting.’