I don’t believe I argued that an LLM couldn’t find and exploit a vulnerability and even break out of some layer of technical controls. That seems realistic and has been demonstrated before and I mentioned that LLMs are used in offensive security work.
Also, I listed the view points I had to show that it seems other more reasonable first assumptions don’t seem likely, therefore, the last potential of this being either faked or carefully not avoided seems more likely than the others (based on the current information we have).
Could you clarify which point or assumption you are objecting to?
There would be a lot more nuance I’d add with more words, but this isn’t the place to write books so I cut it short (the comment was already lengthy).
Still, to address your comment about what you’d expect to see in world 2 and 3 (assume 1 was true), that’s why 1 was addressed separately. I don’t believe I argued that the potential for world 2 or 3 prevented world 1.
As for the ‘evil exec’s strategy’, I would call this a mild incident but if it were much less I wouldn’t guess they would get a lot of press. The press coverage is certainly repaying the token cost as well. If it was planned, it seems to be going well given the press coverage I’ve seen on it. So I wouldn’t assume the plan lacked enough to weaken the idea that it’s a plan. But to be clear, my stance is just based on the info I see now which isn’t a lot… subject to change.
As for the containment piece, if you were testing an AI model on its hacking capabilities that you believed was far more capable than anything you’ve seen, I would assume you would air gap it (a network control). Done right (no signals ability) I would argue this could be next to impossible to break out of. But it’s a fair jab to say I should have added some qualification on the “impossible” piece as next to nothing is truly impossible.
The post was long enough so I couldn’t capture all the nuance and details for sure. Also, this comment was an opinion based on limited info right now, that may change if we found out more. I think OAI does want it framed this way but that’s something we’ll likely never prove if it’s true.
Your comment about jailbreaks being more one off and hard to do consistently in agents is a good point. Still getting an agent to hack isn’t hard even without a jailbreak, you just have to tell get creative in what you tell it. I’ve found telling it that it’s in a CTF or that I own the system that it’s hacking will work fine. A lot of offensive security companies are running agents in their testing so getting an agent to hack seems commonplace.
That’s meant to be captured by point one with the model just being that advanced but more of a negative spin on it. If I felt option 1 was more likely, I think I’d have to agree with you there. Still, there currently are some gaps with that view in my opinion.
There seems to be three popular ways to view this incident.
1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model.
2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
3. The whole thing was faked or at least very intentionally not avoided.
The first interpretation is the only one that is positive for OpenAI and it has some assumptions. First, it’s seems to assume that this is the first case of fully automated attacks using AI. Second, this only happened because their latest LLM was a) more advanced than competitors, b) didn’t have refusals in the model.
Assuming the first about this being the first autonomous AI attack is true (which may be more of a survivorship bias), the second seems to forget that jailbreaks are available for every model. Therefore, the models guardrails don’t seem to be the differentiator here. Also, benchmarks seems to put most models pretty close to each other so it seems unlikely that their capabilities are far beyond what’s in the market already.
So then it’s seems it’s either that this was intentional(ish) or bad security. However, it also just could be that this isn’t the first case of this attack; just the first that was caught.
My take from working in offensive security for over five years is that this likely only looks novel since they did it poorly. Scripts are faster than LLMs and a combination of code, LLMs where it makes sense, and humans is the most efficient right now. Hundreds or thousands or agents spinning up attacks in the internal network is poor opsec and token efficiency. As for why it happened in the first place, it’s hard to say but I’m inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. The timing of this attack after big open weight competitions drops seems too convenient.
I can’t say I’ve done it, just in theory you could print the model to pore on the slurry, directionally freeze it, freeze dry it, and then sinter it… quite a process and may require some specialized tools so I mentioned the FDM infused filament since it’s a similar concept with only a kiln needed.
This is a pretty cool concept for hobbyist with a 3d printer who wants one off metal parts. Sure, you can always cast… but it’s not easy.
There is a similar way to do this casting with infusion and sintering using metal powder infused FTP printer filament. It shrinks down a bit more than this freeze casting sintering I think but it skips the slurry/freezing step of freeze casting.
> “… gives the illusion, without the reality, of safety”
This actually isn’t true. Having done physical security work before, a weird fact is that one of the best physical deterrents is lighting; even over CCTV.
I don’t say that to take away from this, this is great work and I’d love to see the lighting toned down for multiple reasons. However, this should be framed as a security tradeoff not an outright win.
Source:
The Impact and Policy Relevance of Street Lighting for Crime Prevention: A Systematic Review Based on a Half-Century of Evaluation Research (https://www.crimrxiv.com/pub/wl9zqxga/release/1)
This feels similar to the argument that electric will full kill gas cars. I have three cars; one I drive with one pedal, one with two pedals, and one with three pedals. I can say that the stick serves a very different function than the other two which makes me skeptical of the idea that stick shift will go to absolute 0 anytime soon.
The electric car is the new daily driver, the gas car is an old daily driver, but the stick is an old work truck that’s been around for probably over 40 years since its basic (you see pavement when you open the hood). Even if the old cars with sticks get replaced, construction equipment and machinery may always benefit from simple gas engines… even if that’s not ideal for the environment.
I’ve been curious what a polymorphic botnet that runs one (or multiple) distributed LLMs would be capable of doing. The idea would be to evolve the botnet delivery and payload using the clustered compute of all hosts in the botnet to run LLMs that guides the evolution of various botnet clusters. Bad cluster morphs get caught and cleaned off and bad delivery methods never spread, but the best versions survive to continue to grow.
What I envisioned for how it works is fairly similar to this, QUIC can actually be more difficult to detect than it seems since it’s very dynamic.
Sweet! 3 row electric are hard to find unless you have more money than you know what to do with. A used model X was the best option if you’re cheap… and still is with Model YL at this price point. Sadly, this is a bit too expensive to compete with a used Rivian R1S’s or Model X’s, but if they put out a base model cheaper or if you wait a few years for a used Model YL, this could be the cheapest 3 row electric you’ll find.
I’d be curious to see the breakdown on spending by use case. I’ve heard it said that the majority of tokenmaxing comes from none technical uses like reading PDFs, creating PowerPoints, generating graphics/images… ect. But I’ve never heard any actual proof to that.
Weirdly being a security company actually can have the opposite affect. A small portion of potential customers or investors assume the company is more secure because they are a security company after all (and should be); therefore, the customer's security review are less stringent so exec can get away with smaller internal security budgets. Of course good security companys with good leadership doesn't do that... but those aren't the big companies.
Almost all of the major vulnerability and hack are just single spikes at the time it happened and it tails off after that… except Stuxnet. Stuxnet is was much more interesting that most other attacks since it was very political and openly published. Of course, the thing that attack was about is still a news headline today as well
There are lots of types of a “breach”. The first and second (the major ones) were likely related so more like one continuous incident. This one was a vendor breach that had access to their data so not a reflection of their security program as much as the first.
I’m not saying you’re wrong, I’m saying you can’t tell from this incident.
Political bias of LLMs is something not talked about much (except for with Grok of course) but could have a big impact on the next decade. People seem to think that because an LLM gave a nuanced answer that it means it gave the WHOLE picture… and that’s not always the same thing
I’ve done a lot of security consulting work for hundreds of companies and one thing I noticed is that the companies that actually took security seriously were the ones that had been breached in the past. Until the execs and board see the dollar impact themself and not just read about it, the security program never gets the funds it needs.
I’m not saying I recommend LastPass for that reason, but I wouldn’t write them off for that reason.