Interesting how everyone's favorite language seems to be even better in LLM era, almost like passion, skill level and having LLMs matters more than the language.
>That decomposition (perfect on templates, regressed on novelty) is the signature of “scaffold-then-internalize” training on genre-specific data, not a general gain in interactive abstract reasoning.
They're smuggling a claim that benchmarks like ARC-AGI measure "interactive abstract reasoning" here, which is what is claimed by the people that make these benchmarks, and also not proven.
>The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind.
DeepMind hasn't been on the frontier for a while, their current best model is behind Anthropic, OpenAI, Moonshot (Kimi k3), xAI (Grok 4.5), Z.AI (GLM 5.2), and even Meta (muse spark). Gemini 3.6 is behind GLM 5.2, released a month earlier, open weights and cheaper.
You can paint the OpenAI story as a way to try to appear as dangerous as Anthropic with all the Mythos stuff.
I think the way you interpret the null result is downstream of considering creatine as "a supplement". You can make the null result say anything by changing your prior about creatine or supplements. That's the issue with priors.
You also seem to reject the possibility that some things help just a bit. The author has another article in the same vein about "things that can help maybe a bit but the evidence we have doesn't really help detecting small effects" https://dynomight.net/vitamin-d/
The mix of humanizing the LLM, calling it "clanker" and being very aggressive towards it is really weird. I don't think it's a good habit to take, it feels like it could bleed into how you interact with people. Many interactions are through text interfaces these days.
There's also an alternative which is that token prices are already high, downtimes do exist (semi frequent on Claude) or are managed by serving degraded versions/quants (speculation that has never really be proven afaik) or reducing thinking time ("juice" values from open AI), and that free tiers are not used much at all or are a loss leader.
Interesting, good find! Yeah I may be wrong and this may be an error in the leaderboard. Weirdly it shows no reasoning cost and no reasoning tokens used, but for example here https://huggingface.co/datasets/arcprize/arc_agi_v1_public_e... the answer is super short but it says "4945" completion tokens.
I'm not referring to this paper, I'm referring to this leaderboard: https://arcprize.org/leaderboard. Set it to "arc agi 1", "base LLM" and you'll see deepseek at 57%. Submitted 2025-12-01, $0.120 per task. The paper you linked was later than that, and also says "We do not report an official ARC Prize leaderboard score".
So this paper doubled the price to get the same exact result at base Deepseek 3.2 at launch, and wasn't even tested on the verified set.
DeepSeek V3.2 was tried without reasoning and it got 57% on ARC AGI 1. It's a 7 month model, so I'm pretty confident that base LLMs would be able to solve ARC AGI 1 without reasoning/CoT.
The thing about not much difference between models and the harness making them deterministic and useful is wrong. Also models have different strengths and weaknesses and some are better at almost everything by a large margin compared to others.
As for your speculation, I think it's hinging on some companies releasing models for free or no big differences between models. In a world with hyperscalers and companies training models you can quickly recreate Anthropic or OpenAI by having an hyperscaler ally with a model training company, train a good/a better model, and not release it.
This is true, but there's a big difference between saying "15-20 hours battery life, which is 11000 5-seconds activations, which last you a few years with 10 5-seconds activation a day" and "years of battery (btw in small text the real number is given). Especially since they mention that this project is hackable/you can do other things with it, knowing in advance you have something like ~100k button presses means some projects feel perfectly and some others won't really work.
I don't really care about the environmental consciousness, my issue is that presenting a product with a battery that lasts for years when it actually lasts 15 to 20 hours makes me feel like I'm being lied to.
"Battery that lasts for years" being actually 12-15 hours of recording is a huge turn off honestly.
>How long does the battery last?
>Roughly 12 to 15 hours of recording. On average, I use it 10-20 times per day to record 3-6 second thoughts. That's up to 2 years of usage.
They then say:
>Wait, it's single use?
>Yes. We know this sounds a bit odd, but in this particular circumstance we believe it's the best solution to the given set of constraints. Other smart rings like Oura cost $250+ and need to be charged every few days. We didn't want to build a device like that. Before the battery runs out, the Pebble app notifies and asks if you'd like to order another ring.
My oura has lasted ~3 years, I recharge it twice a week usually, and I think it has spent way more than 15-20 hours turned on.
Considering progress in the rest of computing stuff (RAM, CPUs, storage) is kind of "linear"/exponential it sure looks like it's an engineering gap and we're on the right track. GPT 3 was 175B parameters and is today crushed by models that are 32B parameters, that's a lot of progress in 6 years.