Why the world needs Ukrainian victory
snyder.substack.com3 pointsby credit_guy0 comments
> If you could run a nuclear reactor with U-235 as fuel or Pu-241 (both mixed with 95% U-238), which one would you choose and why?
For a human this would not be tricky at all. For an LLM it could be, because this question certainly does not exist in any sort of training, because Pu-241 does not exist in pure form, it only exist as a minor component of reactor-grade plutonium, where Pu-239 would dominate, with Pu-240 coming second and Pu-241 coming third. > Some uses are also simply outside the bounds of what today’s technology can safely and reliably do. Two such use cases [...]: Mass domestic surveillance. Fully autonomous weapons.
If the US Government can't be trusted with such uses, then how can you trust millions or billions of users with arbitrary usage?
Fable 5 -> 0.42
Opus 4.8 -> 0.45
Sonnet 5 -> 0.45
Opus 4.7 -> 0.46
Grok 4.3 -> 0.52
There's an obvious jump at Grok 4.3, and it would not surprise me that the similarity there is because Grok used Anthropic models for training too (you can get the similarity list for Grok and it does look like the top most similar models for Grok are either Anthropic models or some Chinese models).
The damning evidence that K3 used Anthropic models for training is that K3 is more similar to those models than it is to K2.6. If you look at the Anthropic, OpenAI or Google models, they are most similar with their own other models. Not so with K3, where K2.6 is less similar than 15 other models.
Now, why is Fable 5 the most similar to K3 and not Opus 4.8. I think it's quite likely that K3 did some fine tuning at the end, when Fable 5 became available. They probably had all the infrastructure in place, and Mythos had been announced for months, so they were probably waiting for the second the newest Anthropic model was released to start using it for synthetic data generation.