Yeah, I used to get stuff directly from China, then after the postage rate changes I was getting packages that had gone China -> Thailand -> Azerbaijan -> USA. Nowadays they seem to batch packages for the good shipping, from Aliexpress in China to (I assume) some US subsidiary, and then from there it gets parceled out to a shipping company seemingly at random (Maersk, Amazon, USPS).
> Claude Code 2.1.207 and OpenCode 1.17.18, both pinned to claude-sonnet-4-5
So not only is this article AI-written, but the testing was entirely done by AI, too? I can't see any other reason to use such an old model.
> Our traffic passes through a local LLM gateway that wraps requests in its own envelope, a constant we measured at roughly 6,200 tokens with bare calibration requests
Why do you need to do calibration requests to figure out how your own gateway is affecting requests?
> Its subagent lane did not complete cleanly through our gateway
> We attempted to toggle extended thinking in both harnesses and are declining to publish numbers. Our gateway applies its own thinking policy, neither harness's toggle demonstrably survived the path, and anything we quoted would be noise.
Why is your own gateway screwing with your testing?
> For one thing, CF is probably the largest entity serving pirated content internationally while hiding the identities of actual perpetrators for privacy.
Some of them will publicly break it for whatever reason, an NVIDIA engineer (@blelbach) recently made some posts on X about his results using 5.6 Sol Ultra to optimize stuff and then deleted them shortly after (presumably got yelled at).
No, depending on the complexity of the issue models can be into loops, where they go "this is definitely an issue and must be fixed", and then the resulting fixed code gets "this is definitely an issue and must be fixed", and then the resulting fixed code has the original 'issue'.
>Our safety assessments found that Sonnet 5 shows an overall lower rate of undesirable behaviors than Sonnet 4.6, and is generally safer to use in agentic contexts.
which is obviously painting that as a good thing. So reading the next sentence as "in other good news" is reasonable.
You think that if someone can get a model to write a beginner's guide to exploiting code that requires writing your own purposefully vulnerable program, then the creators of that model should be arrested?
Anthropic refused to allow the US Government to conduct mass surveillance, which made the US Government mad. OpenAI was fine with it as long as it was 'legal' mass surveillance. OpenAI is not going to get banned, even if their next model is both better and more dangerous than Mythos.
It's only rewarding hype if the ban gets dropped. If "foreign Anthropic employees that live in the US can't use Fable/Mythos" stays it harms them, if they don't drop the ban and Fable/Mythos stay limited to "every single person who uses the model must individually provide their ID to prove American-ness" it harms them.
Everything other than full self-hosting has this exact power