I think the way you interpret the null result is downstream of considering creatine as "a supplement". You can make the null result say anything by changing your prior about creatine or supplements. That's the issue with priors.
You also seem to reject the possibility that some things help just a bit. The author has another article in the same vein about "things that can help maybe a bit but the evidence we have doesn't really help detecting small effects" https://dynomight.net/vitamin-d/
The mix of humanizing the LLM, calling it "clanker" and being very aggressive towards it is really weird. I don't think it's a good habit to take, it feels like it could bleed into how you interact with people. Many interactions are through text interfaces these days.
There's also an alternative which is that token prices are already high, downtimes do exist (semi frequent on Claude) or are managed by serving degraded versions/quants (speculation that has never really be proven afaik) or reducing thinking time ("juice" values from open AI), and that free tiers are not used much at all or are a loss leader.
Interesting, good find! Yeah I may be wrong and this may be an error in the leaderboard. Weirdly it shows no reasoning cost and no reasoning tokens used, but for example here https://huggingface.co/datasets/arcprize/arc_agi_v1_public_e... the answer is super short but it says "4945" completion tokens.
I'm not referring to this paper, I'm referring to this leaderboard: https://arcprize.org/leaderboard. Set it to "arc agi 1", "base LLM" and you'll see deepseek at 57%. Submitted 2025-12-01, $0.120 per task. The paper you linked was later than that, and also says "We do not report an official ARC Prize leaderboard score".
So this paper doubled the price to get the same exact result at base Deepseek 3.2 at launch, and wasn't even tested on the verified set.
DeepSeek V3.2 was tried without reasoning and it got 57% on ARC AGI 1. It's a 7 month model, so I'm pretty confident that base LLMs would be able to solve ARC AGI 1 without reasoning/CoT.
The thing about not much difference between models and the harness making them deterministic and useful is wrong. Also models have different strengths and weaknesses and some are better at almost everything by a large margin compared to others.
As for your speculation, I think it's hinging on some companies releasing models for free or no big differences between models. In a world with hyperscalers and companies training models you can quickly recreate Anthropic or OpenAI by having an hyperscaler ally with a model training company, train a good/a better model, and not release it.
This is true, but there's a big difference between saying "15-20 hours battery life, which is 11000 5-seconds activations, which last you a few years with 10 5-seconds activation a day" and "years of battery (btw in small text the real number is given). Especially since they mention that this project is hackable/you can do other things with it, knowing in advance you have something like ~100k button presses means some projects feel perfectly and some others won't really work.
I don't really care about the environmental consciousness, my issue is that presenting a product with a battery that lasts for years when it actually lasts 15 to 20 hours makes me feel like I'm being lied to.
"Battery that lasts for years" being actually 12-15 hours of recording is a huge turn off honestly.
>How long does the battery last?
>Roughly 12 to 15 hours of recording. On average, I use it 10-20 times per day to record 3-6 second thoughts. That's up to 2 years of usage.
They then say:
>Wait, it's single use?
>Yes. We know this sounds a bit odd, but in this particular circumstance we believe it's the best solution to the given set of constraints. Other smart rings like Oura cost $250+ and need to be charged every few days. We didn't want to build a device like that. Before the battery runs out, the Pebble app notifies and asks if you'd like to order another ring.
My oura has lasted ~3 years, I recharge it twice a week usually, and I think it has spent way more than 15-20 hours turned on.
Considering progress in the rest of computing stuff (RAM, CPUs, storage) is kind of "linear"/exponential it sure looks like it's an engineering gap and we're on the right track. GPT 3 was 175B parameters and is today crushed by models that are 32B parameters, that's a lot of progress in 6 years.
Yeah, in an ideal world where I'm the ideal me I wouldn't use my phone in my bed, but I haven't found a way to stop doing that which I can stick with, so I try to limit the damage.
Part of what I wanted to say is, there is conventional wisdom, then there is how you actually put that wisdom in practice in a way you stick with. I've struggled a lot with the implementation, but sometimes by throwing lots of stuff at the wall I find something that brings me halfway there. It's not the "golden way" but it leaves me in a better place than before, with a bit better sleep, a bit more self knowledge, and a small victory.
Hard to answer precisely without knowing what conventional wisdom didn't stick.
The common levers I know and that worked at least a bit for me:
- start by having a fixed waking time, and get sunlight or bright light quickly after waking up. Normally relatively fixed sleep time is supposed to follow. For me waking up is the easy part, transforming that into getting up and going outside is harder. Another option here is a strong (like, really strong) lamp on a timer, or letting the morning light in your bedroom (this one is usually not recommended I think, most people seem to be blackout curtains style, but for me it gave me a nice 6am waking time with good sleep last summer).
- melatonin. Two main ways: using it as a kind of hypnotic, so ~30 minutes before sleep, experimenting with 0.3mg to ~2mg doses ; then using it as a circadian regulator, this is a good resource https://lorienpsych.com/2020/12/20/melatonin/, search for "TO TREAT" in the page.
- app timers, for me it was mostly no twitter and no youtube, or a very low time for each.
- light, ie reduc light before sleeping. Not just blue light and not just screens, if I'm on my phone in bed I'll reduce the luminosity a lot, same with computer, same with e-reader. I also try to avoid using too much the lights in my room. More light tend to make me feel more "wired" and less ready to sleep.
- "meditation" to cut rumination, by which I mean "lay down in my bed, gently try to find sensations in the body and to stay focused on them, by gently I mean it's a very low stakes game where the goal is to find sensations in the body and give them attention, but losing focus for a while is not a big deal".
- shower in the evening, as I don't like feeling dirty when I am in my bed, but also not just before bed as sometimes I don't really want to go take a shower and this delays my bedtime
- clean bedsheets, bedroom, stuff in/on your bed
- AC in the summer, I wouldn't be able to sleep properly without it
- sleeping mask. It helps going to sleep, but it falls of my head every night so it doesn't prevent waking up with light too.
- making getting good sleep the priority of the evening. This is easy/possible for me due to my circumstances (ie low responsibilities in the evening). The way I do it is that unless something is actually important, what I'm trying to accomplish in the evening is prepare myself for sleep and get good sleep. This can look like not starting a movie at 11pm, not booting up games, not eating a super heavy meal, not drinking too much water after 6pm to avoid waking up to pee, if I have things I want to do try to do them early so they're done earlier, move some stuff I want to do every day like spaced repetition in the morning.
I've tried to like Go with HTMX, but the big issue was always Go templates. I feel like if there was something like JSX/TSX but for Go, it would be a way better dev experience, but right now it's mostly a pain. Templ tries to go in that direction, but a year or so ago editor integration and tooling weren't great.
There is value in splitting things but there is also a cost. You have to train the specialized model, for that you have to know your use case, you have to hope the use case is going to be stable over time, you then have to see if you can remove english -> slovakian or coding from a model without affecting the useful parts.
You also seem to reject the possibility that some things help just a bit. The author has another article in the same vein about "things that can help maybe a bit but the evidence we have doesn't really help detecting small effects" https://dynomight.net/vitamin-d/