I also suspect the questions asked matter a lot, and the system prompts matter a lot, because "the map is built from nothing but the words they choose" - so this is more a measure of linguistic style than anything.
If you use the Claude Code harness on two models you will probably get very similarly styled output. I would not be surprised if K3 "stole" a lot of the harness that was leaked.
Notice how the only rows that even come close to being gray is GPT-5.4+Mini, diverging even from other GPT models. Is this because it has a wholly different training set, or (more likely IMO), did it just have a system prompt that leads to a different style?
My takeaway is that closed model providers are dangerous. OpenAI and Anthropic are more motivated than anyone to prove that models can be dangerous, and so they will make dangerous models. "Look how dangerous our models are!" No bro YOU are the danger.
Not to mention the "distilled" models almost certainly have other training inputs as well, again watering down the meaning of the word.
And if just partial output is all that it takes to declare a model distilled, then every model trained on Internet content since 2023 is now technically a "distilled" ChatGPT and Claude model.
1) Your own Wikipedia link goes on to describe using logits. Yes, language evolves to mean multiple things, and that is my point: Anthropic is pushing for a watered down definition. Furthermore, Anthropic hides thinking, so you do not even really get model outputs, you get some downstream partials. Further-furthermore, Anthropic does not describe their own models as distilled when they produce training data. Why? Because generating training data != distillation.
2) This is a stretch: it allows Anthropic to arbitrarily define "attack" via TOS, and ignores the fact that the generated training data is literally paid for by the "attackers".
"distillation attack" is such a loaded term that really pisses me off.
Distillation is a technical term with real meaning, and historically requires logits which Anthropic does not provide.
"Generated training data" is the correct term. It's not an "attack". And Anthropic undoubtedly also generates training data for each new generation of models, yet you never see them claim Fable is a distilled Opus.
Am I dumb or does this chart make no sense? Or why does the line only go up even with compaction? Or maybe "overall trajectory size" is hiding some meaning I don't understand?
"threatens" - Dawg there aint a lead anymore. I've been testing K3 and it is outperforming Sol and Fable on a project of mine, fixing stuff I couldn't get either to.
Hey I love the idea, don't let the haters get you down just because the site was vibe coded, that's trivial to improve upon.
I do think more information is needed to understand exactly where these claims are coming from. Not sure the "not legal advice" tiny text at the bottom is enough, people are going to see the claims and most won't double check.
Yeah I sometimes see people on here getting defensive when you call out AI slop, saying maybe it's just a human who writes like Claude, and I really don't care- slop is slop.
That battle tested bit is so true, always was. It’s hard for engineers to admit their hard work is worthless until they have let other people use and shape it.
I think it was always hyperbolic to call rewrites the single greatest mistake (I can think of worse) but I think most of the wisdom still holds. And AI generated rewrites are arguably riskier because now NOBODY is familiar with the massive codebase.
If OpenAI is this shady culturally then all kinds of dark secrets could come out.
They just released 5.6 to great fanfare with benchmarks showing even their weakest model (Luna) supposedly smoking Fable. Having used it a couple hours I'm certain this is a lie, and I'm now curious what else is.
If you use the Claude Code harness on two models you will probably get very similarly styled output. I would not be surprised if K3 "stole" a lot of the harness that was leaked.
Notice how the only rows that even come close to being gray is GPT-5.4+Mini, diverging even from other GPT models. Is this because it has a wholly different training set, or (more likely IMO), did it just have a system prompt that leads to a different style?