They are weights for matrix operations, so in principle you'd think they would be deterministic. In practice it's more complicated than that.
Not to be snarky or dismissive, I mean this genuinely: ask an LLM about it. I currently have a headache so I'm not up to explaining the technical details, but they are interesting and worth reading about.
Except that even the exact same model won't output the exact same results, that's a fundamental aspect of how LLMs work. They're probabilistic/stochastic, not deterministic.
Eschatological is a perfectly fine word. The post is about doomsday scenarios, but perhaps the commenter is equally tired of people who think AI will be salvation!
Anyway, if the word confuses anyone, they're in luck: dictionaries still exist :)
So if the detector isn't perfect, it isn't useful? Not sure I buy that.
Also, even if there's no way to detect what the activations are doing, we already have the ability to analyze your proposed threat statistically. If the model repeatedly uses insecure libraries in most trials, then yes, in that case it would be prudent not to trust those weights.
Assuming one doesn't get banned for violating some ToS clause about using a closed model for LLM research, it could be possible to run those evals on a closed model too (likely at much greater expense). But there's a big difference: if such a statistical anomaly is discovered in an open model, one could potentially fine-tune that behavior out of it. With a closed model, that won't be an option.
Whether it makes a difference to the US government or not is beside the point. Even with a perfect solution, the current administration could do some mental gymnastics to achieve whatever political outcome they want. I’m not trying to make a political statement here, my point is technical: open-weights at least give us the possibility of visibility into why they generate what they do; this simply isn’t true with closed models.
But that would happen with literally any model without grounding and is more of a quality/competence issue, not what's being discussed. It'd be a bit of stretch to conclude open models are no more auditable than closed models based on that possibility alone.
Also, in that case, there would likely be activations indicating that it is favoring a specific version. If that's an insecure version, sure that'd be suspicious... but again, you're only going to be able to verify that's what's happening in an open model.
Maybe you can illustrate a realistic scenario in which that would be a problem, otherwise I don't really understand what your point is in this context.
Not sure I understand that position. Unless we're talking about a scenario in which one is using an outdated model along with no grounding (which, imo, PEBKAC), why would the model be pinning insecure libraries?
If it does have grounding, and can therefore see that it's introducing vulnerabilities to the code it's generating, yet does so anyway... I suppose we could invoke Hanlon's razor, but if the model is that incompetent, it probably isn't the right tool for the job regardless of its provenance.
That said, we aren't talking about incompetent models, we're talking about models sabotaging projects due to hidden motives. My point, again, is that those motives could potentially be revealed with open-weight models, in a way that will never be possible with closed models (barring some sort of legislation requiring independent third-party interpretability audits, which I suppose is in the realm of possibility).
I'm talking about using mechanistic interpretability to see the model's intent. If it is deliberately using compromised libraries to weaken some code's security, there's going to be a signal in its hidden activations that it's doing so.
Finding these kinds of activations is something Anthropic is actively researching [1] but they're the only ones who can use those techniques to see Claude's intent. On the other hand, if a model is open-weights, in theory whoever is running the model could look inside the activations at runtime to see if a hidden vector associated with "deception" or "sabotage" is being activated [2].
(Those sources are just a couple of relevant starting points I could find without much effort, there is also https://www.neuronpedia.org/ if one is interested in seeing interactive demonstrations of interpretability concepts)
Maybe I should clarify. As I understand it, the kind of vulnerability being discussed is something like a Chinese model invisibly "realizing" that it's working on an American project, and then deliberately leaving subtle security bugs in its generated code for Chinese hackers to later exploit. As far as I know, that scenario is hypothetically possible, but has never been demonstrated to happen in the wild. Admittedly, I could be wrong about that! If anyone has evidence to the contrary, I'd love to see it.
Of course, one could retort that gathering that evidence may be nearly impossible now, but my point stands: in the future it might/probably will be possible to properly audit open-weight models. Closed models, on the other hand, will always be a black box.
In my opinion, the big issue with that argument is that advances in interpretability research and steering conceivably could, and probably will, render moot that (as of now, purely hypothetical) risk of subtle sabotage for open-weight models... but not for closed models.
"Strandfall is a sci-fi story set in a post-apocalypse, with an emphasis on climate change, co-operation, and adaptation."
I'm in the early stages of designing a world/game focused the same topics, in part because I simply want more stories like this to exist. I live an ocean and a continent away, so of course I won't be participating, but I don't know if I've ever subscribed to a newsletter so quickly. Looks like an absolutely brilliant idea and I hope it's successful!
I'd be extremely interested in participating if a version of this ever came to my corner of the world. Even if it doesn't, just knowing about it is inspiring and motivating :)
As someone who has only recently been thinking about learning how to solder and work with PCBs that's actually useful info, thanks!
Unfortunately I do have experience with the smell, being around others working with improper ventilation... I'll be passing that advice along to my BIL too, ha.
True. The linked article's title says that. I wonder if that was a typo by the OP or one of those HN quirks where the title was automatically changed when it shouldn't have been.
It's wild if it's running at 400 fps because nobody has a screen that refreshes at 400Hz. Every frame rendered past the screen refresh rate is wasted compute. Easily solved by limiting the frame rate.
I'm not anti-AI or anti-data center, in fact I lean more towards "pro-AI" overall, but to say that everyone who claims data centers can contaminate water is lying is a strong claim that doesn't really hold up to scrutiny. If they're saying "every data center without exception poisons water" then sure, they're lying, but that's certainly not what I'm saying. A couple examples from this year are linked below. I understand if you're skeptical of politicization in the highly-publicized Georgia case [1], but it's harder to dismiss what happened in Wyoming [2].
As for energy: honestly I agree with you that the solution is to build more power generation capacity, but that doesn't change the fact that in the meantime energy prices are already increasing substantially in many areas because of data centers [3].
Like I said, I don't necessarily agree with the idea and I don't feel strongly enough about it to really argue in its favor, but to answer the question: the same reason why OpenAI doesn't operate out of Sam Altman's garage.
At a certain level of compute you need specialized infrastructure -- such as a purpose-built datacenter -- for the energy needs (and really, I think the stronger argument to be made here is about energy, not raw speed, and where the argument might fall apart is the historical fact that compute tends to become more energy-efficient over time).
Not sure whether the breathing/murder analogy is apt, but I get where you're coming from and I would probably agree that a blanket restriction on computer speed wouldn't be appropriate.
Or you might build a data center that poisons a community's water and drives up the cost of energy for your neighbors. We can't pretend there are zero negative externalities that accompany unconstrained compute.
To be clear I'm not necessarily agreeing with the idea, but to be fair, there's more to it than you're suggesting.
I don't know if it's a bad idea or not, but I'm struggling to understand how the idea as presented in the post would be a violation of fundamental human rights. Do you care to elaborate?
I've never used a Grok model before because I have my OpenRouter settings on ZDR-only. I just checked, and apparently there are ZDR xAI endpoints now [1], so I might actually try this. Out of curiosity, does anyone here happen to know when those were added?
[1]: However it does say "Requires user IDs" under anonymity, which is unusual on OpenRouter and not something I particularly like to see. Generally, OpenRouter is a proxy that anonymizes requests to providers, and I can't find an account-wide setting to enforce that like ZDR-only.
Probably not complete BS. This is anecdotal, but years of experience has taught me that aluminum-containing antiperspirants cause contact dermatitis for me after extended use.
I was also diagnosed with a nickel allergy by a dermatologist when I was a child. Metal allergies are real.
I wouldn't go so far as to say aluminum is toxic for everyone, but it's certainly something I avoid putting on my skin.