For all we know, the prompt provided compelling evidence that the requestor had authorization to pentest the target server. Or there may have been nuance in the network configuration that made it seem like such access was authorized.
In the absence of details about the prompts used, the environment, or the network configuration, we do not have enough information to know for certain. So any claims that this is an issue of alignment are based on pure speculation and generous "reading in between the lines" with regard to what has been said publicly by OpenAI and Hugging Face
Also, I object to your anthropomorphizing. It's not clear that any crime occurred. My lay understanding is that intent is required to prosecute under CFAA, and as much as frontier labs would have us believe otherwise, they have no more ability to intend than the text field into which I type this message.
The author is trying to provide a counterweight to the volume of articles that simply repeat OpenAI's account of the events and their interpretation without much pushback.
They are not suggesting that OpenAI or HF have lied about what happened, but rather that OpenAI is advancing a narrative framing their models as supremely dangerous and capable, while positioning themselves as the only ones qualified to manage that danger.
At the same time they are not being particularly transparent about what actually happened (e.g. was this one-shotted or if not how many trials did they run and what were the outcomes of those, was it emergent as a part of routine cyber-capabilities tests, how much prompting was involved, what prompts were used)
Note that this is at a time when they are lobbying for a regulatory approach that would give frontier labs special treatment.
I would guess the editorial team at The Guardian may not like articles that get too in the weeds of technical details and questions like these that the vast majority of their readers wouldn't understand. I don't know. But I empathize with your disappointment. I don't think it's fair to say that they are contributing "nothing" especially given what most reporting on this has looked like.
I should clarify why this annoying to me, beyond what I wrote above.
Nothing is transparent about the nature of the experiment, but at the very least it should have been qualified that this is not a standard deployment.
It is difficult to assess how "real-world" the setting was because we don't know any particularities. We don't know about the prompting, the environment setup and configuration, or any other details that would allow anyone to differentiate this from a benchmarking setting.
> UK AISI’s evaluation shows that models such as GPT‑5.6 Sol are increasingly able to sustain complex, multi-step cyber operations over long time horizons. This incident implies these theoretical capabilities do apply in real-world settings.
I should clarify a bit more why this is annoying beyond what I wrote above. The main issue is that this was not a standard deployment, and the lack of particularities make the size of the gap between "real-world" and "benchmarking"/"lab" difficult to assess.
We don't know about the prompting, the context, the environment + configuration, or any other details that would allow anyone to differentiate this from a benchmarking setting.
I agree with you that the article was disappointingly light, but there is not in general a singular correct answer when it comes to interpreting narrative and I don't think that is a useful way to evaluate what the author has written, nor real-world writing in general. The author does not implore you to have a singular 'correct' point of view on the matter but rather to hesitate before repeating the claim that "the AI broke out of the lab and went rogue" because we don't know enough to know that this is what happened.
As you say, it seems likely that the broad strokes of OpenAI's narrative is true. This is never in question in this article, so the author doesn't "stop short" of accusing them to have lied, he never moves in that direction and that has little to do with the thesis. The author is saying that OpenAI's framing of what happened is part and parcel with their longstanding PR strategy, to position models as supremely dangerous and themselves as the uniquely qualified stewards of those.
Since many critical details were not provided, everyone has to read in between the lines, and there are particular common readings I'm seeing both in discussions and in published articles that are problematic in the sense that they are effectively hallucinations -- i.e. we don't have enough information to make those interpretations. This is where the critical thinking comes in.
For instance, I see many are assuming that this event was emergent, arising as a part of routine cyber-capabilities testing, rather than induced or suggested by specific and careful prompting. Either are possible, but we don't even know so much as how lengthy the prompt they used was, let alone how suggestive it was with regard to the approaches the models used. Many are assuming that this was one-shotted, but again, we don't know how many times this particular evaluation was run and what the outcomes were of all the other runs. It could be that this was completely emergent and that it was one-shotted. But I suspect if it were, OpenAI would have said as much.
You don't really need anyone to be lying here. It is likely that the broad strokes of the narrative are true and that no collusion or conspiracy took place here.
The issue is that a lot of important details in that narrative are missing, and the devil is really in the details here. I suspect that those details would make the result seem less exciting and that this event would move the needle far less for them if they were more forthcoming.
A decisive detail would be the prompt used. OpenAI gives virtually nothing here, not a sanitized prompt and not even so much as a description of how long the prompt was and what sorts of instructions it contained. Many are inferring the model behavior to have been fully emergent and unprompted, arising naturally from routine cyber-capabilities testing. But we can't know this because we don't know anything about the prompt or the context the model had access to.
Another detail: how many times did they perform this particular experiment before they obtained this result? What were the outcomes of all the other runs? Many are assuming this was a one-shot result, which I suspect is what OpenAI intends for us to infer. But we can't know that to be true.
One annoying claim from the OpenAI side is that long-horizon goals in real world settings are now effectively settled. Previously there were some bounded and tempered benchmark results, but now OpenAI can point to this event and announce "AI independently went rogue and escaped the lab, what more do you want?". This bypasses the need for anything quantifiable or wading through multiple detailed case studies to get a more sober view of model capabilities. It relies instead on the emotional weight of the spectacle.
I hate to be so obnoxious but I think you might be overestimating the average user of this website with regard to media literacy, and if that's the case, then the basics of it seem very relevant for the homepage
None of what was disclosed shows that this is what happened, by the way, since we know absolutely nothing about what the specific prompts were that led to the incident.
You aren't going to locate incontrovertible evidence that what happened as or wasn't engineered. Anything like that is going to be private and that is unlikely to change. And that's not really an interesting question anyway.
As widely as they shouted from the rafters the news of the so-called breach was, what OpenAI provided was sorely lacking in crucial details.
We are missing, for instance, prompts that were involved, agent architecture + system/tool permissions + scaffold architecture, whether this was a one-shot occurrence and if not, the number + durations + outcomes of other runs involved + how each of those matched whatever scoring criteria were used, and the extent to which the exploits themselves were truly novel or just assembled from easily accessible clues.
In lieu of these items, the author here suggests that we use some media literacy and critical thinking to read in between the lines instead.
In doing so, one sees that instead of specifics, OpenAI gave a breathless narrative rife with superlatives ("unprecedented") that reads as promotional material moreso than a security disclosure, naming specific OpenAI models and alluding to an even more capable pre-release model.
They go on to claim the events imply long-horizon goals work decisively in real world conditions, so that now instead of merely citing boring benchmarks they can point to this and say "AI broke out of the laboratory and went rogue". Naturally, they situate themselves as the uniquely qualified steward for these supremely powerful and dangerous models.
Nevermind the fact that this was no ordinary deployment and the assessment here depends on the gimmick and emotional weight of the spectacle rather than something quantifiable (i.e. a boring benchmark).
Note there's no real requirement of conspiracy or collusion between OpenAI and HuggingFace here BTW. But my sense is that if they provided any of the specifics I suggested earlier that this outcome would not be as exciting or frightening
I mentioned this in a sibling comment but even for mathematicians, the intimidating notation and the more formal language might give the wrong impression about how we think about math. Actual thinking and even discussions with other mathematicians tend to be looser and more concrete and tactile, but the notation and language are there in part to act as a sort of lingua franca to help everyone stay on the same page, since everyone thinks at least a little bit differently. It also helps to keep you honest and catch situations where your thinking was muddied, since this language is so specific and writing things down has a funny way of catching things. And good notation goes a long way towards making the simplicity of an idea clear, or completely muddy in the case of bad notation.
Yeah it is a lot of simple ideas stacked one on top of the other, but the edifice is so large from some vantages that the building blocks aren't visible, or tractable to think about independently. And sometimes the ideas are very subtle, so you can only develop fluency partly by spending lots of time playing with those blocks by building your own little structures. You also develop fluency by talking to other mathematicians
I like to emphasize that the ideas are usually very simple at their core. Sometimes they map to kinds of objects or reasoning that non-mathematicians use implicitly all the time in their daily lives, mathematicians just have words for them and so are able to use them explicitly.
And I suspect the density of the language/terminology may give the wrong impression about how mathematicians think about the math they are working on. I mean, different people think / experience / practice math differently of course but IME the underlying thought about a particular problem tends to be much looser and concrete than formal math writing would imply.
That more formal language is needed of course because at the end of the day, it is how we communicate our thoughts in the way that other mathematicians can understand them, not to mention how we can check our own thinking
I think don't be too hard on yourself. If you haven't already, it may be worth considering whether ADHD or something else might be playing a role. We aren't machines, we're messy bundles of biology and understanding our own particular messy bundle can go a long way towards getting it rolling in the right
direction.
> 2. BI dashboards(we removed spend on tableau, retool and appsmith)
Can you explain this in more detail? You saying that DeepSQL can create BI dashboards that makes Tableau and friends redundant -- are you talking about analytics dashboards consumed by business teams or more like db ops dashboards used by backend teams? Does the user prompt the agent for them, or does it build dashboards that it infers are needed?
The title is ambiguous but not wrong. It's a command-line game in the sense that the controls are primarily text commands issued over a diegetic console within the GUI. It is not a purely text-based TUI and does not run inside your factor terminal emulator.
Ooh, slick rhetorical moves you just executed. You couched your comment in the jargon of formal logic, and in doing so situated the parent under a convenient ad-hoc pseudo-formalization where you could paint them a raving illogical, thereby dismissing most of what they wrote through low-effort counterexamples. Mathematical certainty on your side, you make a chivalrous concession to the parent, reinforcing your noble-stoic character, while sealing their dismissal in the mind of any observer.
I claim that it is trivial, albeit obnoxious, to dismiss most informal utterances in the way you have just done. This is because informal language is complex and enthymeme in the extreme -- we don't tend to articulate every e.g. assumption and premise and contextual relation up front in every utterance. Language is messy because the concepts under discussion are complex and messy, and vibes are powerful tools for wrangling that complexity and getting at meaning. And the majority of paradoxes and contradictions one can identify carry no signal whatsoever. Perfectly reasonable statements can be rife with them.
For example: my interpretation of what the parent wrote was something like "the 'making stupid mistakes' line together with the rest of the blog post signal towards neurodivergence, and someone who is harshly self-critical. so with urgency, I warn the author to be weary of comments on HN, as there is a strong likelihood of being led down an unhelpful path that might aggravate this, particularly given the harsh self-talk and what we know about the subcultures that frequent this place"
Speaking of unarticulated contextual relations, "I often make careless mistakes" is literally in the DSM-V as a diagnostic criterion for ADHD.
I agree with Gwern in that I think for the vast majority of people, short-term melatonin supplementation is useful and can cause little harm, and it is extremely safe as far as supplements go.
But I don't think it does anyone any favors to oversell the idea that it has "few" or "no" side effects -- it has mild side effects, most commonly reported in the literature are daytime fatigue, headaches and GI symptoms, and also nightmares. Mild doesn't mean it isn't a nonstarter for some people.
It's also important to remember that there are major gaps in what we know about melatonin; notably the effects of chronic supplementation are not well-studied, but earlier final awakening has been documented and this is quite commonly reported in anecdata -- I can contribute a datapoint there, as can most people in my circles who have used it.
To be clear, melatonin is great and useful, but as someone with a rare lifelong chronic sleep disorder who is intimately familiar with this substance, I think it's most useful when we're clear on what we know, what we don't know, and what actually are the limitations on a substance.
Just because downing a bottle of it probably won't cause systemic organ failure or otherwise any kind of medical emergency in most people doesn't mean there aren't tradeoffs to consider when using it, especially if you are sleep-challenged
I wonder how intelligible classical Maya is with modern Maya languages/points on the Maya continuum. For instance, does the classical word for fox share any resemblance to any Maya word for it today?
I can imagine it going either way really but would probably guess there was vastly more drift in the case of Maya. I would naively guess that the printing press would have a dampening effect on language drift, and that the kind of repression of both the language and culture under colonialism would encourage it.
I am sorry to be harsh but I find it amateurish that they would use an AI generated hero image for this and presumably fabricated LLM output -- fabricated by an AI image generator no less
Whenever I create an image like this for the purpose of a demo, I make certain that it demonstrates either real input/output or at least is exemplary of real input/output because the whole point is to instill confidence in the tool. Sure, if the raw outputs aren't clean/comprehensible enough for presenting to stakeholders or others, fine, clean them up to make them comprehensible or add explainers, but there shouldn't be any need to fabricate the inputs.
I feel obligated to respond to the hypothetical "But they don't want to tie it to a particular restaurant or brand" -- you don't have to! Doordash has taken generic food photos for this exact purpose.
hm [at / arroba] listen dot systems