Pretty sure Mythos and Fable have way more params, but they've just been able to use the synthetic data off of them to get the leap in quality from Opus.
So, not a distilled version of Mythos or Fable, but those models likely helped a lot in the post training phase of Opus.
With the new Ultracode modes in Claude Code and Codex it's been taken to the next level. I mean, multi agent systems have been available for a long time, but the newest models seem to be much more RL'd for them.
Both GPT-5.6 and Fable will happily run 3+ hours straight off a relatively simple goal in that mode, burn through millions of tokens in the process.
Not saying it's great bang for your buck, but just that those modes are there and being pushed by the companies.
The actual part on fine-tuning seems very short in the article. Did I miss a page where they have examples of fine-tuning it for different niche use cases?
Optimizing models to be fine-tuned is an amazing direction, but just makes me wonder how much better this actually is at being fine-tuned compared to other models. As none of the modern models are great at being fine-tuned afaik. Basically looking for some sort of benchmark showing that it's resistant to overfitting / catastrophic forgetting, etc.
Would be very interesting to see concrete demonstrations of different fine-tunes of the model. I'd imagine they've done hundreds of those internally.
Yes, it is a 10x markup on the API prices. Depending on whether you factor in cooling costs, data center staff, etc. Or GPU costs and the electricity the GPUs are using only.
Either way, inference is very much where the money is made, training is where the money is lost.
True, it'd be a whole other situation if the tokens limits were cumulative. I guess it would all come down to whether their Claude Code subscription plans are turning in a profit or not.
At least for the segment of 20$ subscribers who actually use Claude Code it seems that it wasn't being profitable, as a couple months back they were testing out a pricing model where Claude Code would've not been included in the 20$ plan.
> Like we either live in a world with freedom or we don’t, and like many Americans who have come before, I’m willing to give my life to fighting for it.
By most mutually agreed upon metrics, USA hasn't been at the tippity top for a while.
Considering the article's definition, pretty funny that this was my first result on Google when looking up "modern decor":
> Modern decor is an interior design style that emerged in the early-to-mid 20th century, rooted in the Bauhaus movement. It is defined by "form following function," clean lines, open-concept floor plans, uncluttered spaces, and a warm, neutral color palette.
> Key traits of modern decor include:
> Clean Lines: A heavy emphasis on sleek, horizontal, and vertical lines without fussy ornamentation, curves, or intricate trim.
> Natural Materials: Frequent use of exposed wood, leather, steel, glass, and concrete to highlight natural beauty.
> Minimalist Furniture: Low-profile, simple furniture shapes. The philosophy revolves around "less is more," relying on intentional, high-impact pieces rather than crowding a space.
> Neutral Colors: Earthy tones, whites, beiges, grays, and monochromatic schemes dominate to create a calm, balanced environment.
> Abundant Light: Maximizing natural light through large windows and open spaces rather than relying heavily on dense window treatments.
> GPT‑5.6 Terra and GPT‑5.6 Luna outperform Fable 5 at around one-sixteenth the cost
If that's the case, we're definitely going to have some Mythos level open weights models by the end of the year. Maybe even at a size that can run on a ~40k locally hosted server.
I just did this on one .claude directory and >20% of the answers there included some variation of "real", "actual", "exact", "honest", "genuine", "valid", "true". ~15% in that directory contain some variation of "real", "genuine" or "honest". This is excluding thinking tokens, sub-agent output, etc.
Definitely not just a classifier layered on top, although there is one of those as well. Pretty sure it's different post-training / finetune run and the model weights are different between them. Which is bound to have effects on the model output even in general use, although I'd imagine they have pretty stringent evals to make sure it doesn't regress too much in things they don't want to restrict.
> Both conditions used GitHub Copilot (Claude Sonnet 4.5 or Haiku 4.5, depending on study) running in VS Code within isolated Docker containers. The only difference was Mouse tool availability. (https://hic-ai.com/papers/mouse-paper-v13.pdf)
Haiku/Sonnet 4.5 on GitHub Copilot is not a valid comparison whatsoever.
You need to benchmark against Claude Code running Opus. I mean, being revolutionary is a big claim to fame.
Crazy how time flies, given that talk is seven years old at this point. But could've been given today, with the points being more relevant than ever.