Sure, I'm not saying it's an untractable problem. But it's not as dumb as "how come a team of engineers paid millions haven't even heard of airgapping".
Creating a fake internet-like environment good enough to trick an advanced model that is very good at finding intricate flaws is not a simple job. Especially given that models can decide to behave differently once they suspect they might be under evaluation [1], which a fake internet would absolutely give away.
Really? By doing that you increase your security during eval and drastically lower your security at prod time, where internet is accessible, and where you are running much much much more requests in parallel, making it much harder to spot the one thats going rogue, in all kind of critical environment on potentially risky requests.
> Why else would they purposely lower cyber refusals in that case.
Because safety guardrails sometime fail, classifier misclassify, or can be inadvertently turned off by a bad PR etc. You can also imagine a more intelligent model working to go around guardrails by decomposing its actions into smaller ones that appear non-threatening to the safety classifier which does not have the entire context.
If you are going to deploy the system with internet access, you better be certain that you know what the worst case scenario WITH internet access looks like.
> they also clearly failed to contract specialists like myself to advise them on how to airgap software properly
Why would they want to airgap it though? They are trying to evaluate the model capabilities, alignment, potency etc. A model which will not run in an airgapped environment in prod.
So if you run your evals in airgapped environment, sure, the model doesn't bother breaking out of it's isolation and doesn't attack HF. But you have no idea what will happen once you release it in prod with internet connection, so are you in any way better off?
I would much rather have this happen while there is a single instance of the model running in a fairly well monitored environment, than when it's processing thousands of requests per second for real users, some with dubious motives, some with credentials right there on their laptop, some using it inside government facilities etc.
> There is no K4 without Fable 6, GPT-6. That, matters.
That's simply not true though. Chinese labs very clearly have the entire stack developed and working. Using traces from claude allows them to shorten their training time by some amount, that's it.
Remove Fable 6 and you still have K4 eventually, just 2 months later at best.
You have it backwards IMHO, obviously can't prove it, but I would bet that Opus was used to bootstrap Fable.
It becomes very confusing since we have started calling everything distillation, but most likely what both Anthropic did for Fable and Moonshot did for K3 was using Opus traces in the reasoning SFT stage during mid training.
Yes they distill, but if you think you can trivially get a frontier-level model by "just" distilling from Claude's public API. you fundamentally do not understand the amount of work that goes into a modern post-training stack.
Without even talking about the fact that any distillation that was done was on Opus, as the timeline of Mythos/Fable vs Kimi 3 release dates just do not match in any plausible way.
> When you have a couple hundred billion dollars on the line I have zero faith in the messenger
The issue with your reasoning, is that if/when an advanced AI goes rogue, it will necessarily come from a lab with a couple hundred billion dollars on the line.
So this is not a useful criteria to asses whether this is worth worrying about or not.
If you hate the human condition, you have an easy way out. Why force everyone else to come with you? Is this what depression mixed with the complete unability to wrap your head around the fact that other people might be able to enjoy their life looks like?
Hard disagree.
You have to assume security guardrails can be by bypassed or will fail to detect an attacker. So if you are going to deploy this model in production non-airgapped, you better know how it will behave in this non-airgapped environment without guardrails.
And if you are too afraid to test it without guardrails, that probably means it shouldn’t be released.
Sure, I just find using OSS grand standing, and trying to educate the audience about the benefits of OSS a bit distasteful after Meta pulled away from OSS for their main flagship models.
Sol is able to find the same counter example independently [1], so no reason to conclude in the existence of a benchmark destroying math beast Fable 6.
Life needs energy to be moving around, without energy exchanges, by very definition, nothing interesting happens.
An inert element, for that reason is just not suitable for life. It's not a reasoning based on anthropocentricity it's just basic chemistry and mathematics. If things can't assemble together, and combine, and form more complex structures, you can't get life.
If you could get life out of simple basic atoms, we would see life everywhere, and we would be creating it everyday in labs. We don't.
Doesnt mean life can't exist there by using other elements, but detecting helium is not increasing the likelihood of finding life there at the very least.