"Stripping extended thinking: Extended thinking blocks (shown in dark gray) are generated during each turn's output phase, but are not carried forward as input tokens for subsequent turns. You do not need to strip the thinking blocks yourself. The Claude API automatically does this for you if you pass them back."
It's more nuanced in the various modes, but i haven't seen it boil down towards Thinking Tokens surviving more than two turns.
I've thought about the high-jacking of reasoning-chains as a potential vector, but never saw a proven implementation in american models since, from my understanding, all major vendors throw out the reasoning tokens between turns.
As part of my consulting, i've stumbled upon this issue in a commercial context.
A SaaS company who has the mobile apps of their platform open source approached me with the following concern.
One of their engineers was able to recreate their platform by letting Claude Code reverse engineer their Apps and the Web-Frontend, creating an API-compatible backend that is functionally identical.
Took him a week after work. It's not as stable, the unit-tests need more work, the code has some unnecessary duplication, hosting isn't fully figured out, but the end-to-end test-harness is even more stable than their own.
"How do we protect ourselves against a competitor doing this?"
I hope there's an uncensored version of the Internet Archive somewhere, I wish I could look at my website ca. 2001, but I think it got removed because of some fraudulent DMCA claim somewhere in the early 2010s.
I hate using voice for anything. I hate getting voice messages, I hate creating them. I get cold sweats just thinking about having to direct 10 AI Agents via voice. Just give me a keyboard and a bunch of screens, thanks.
I grew up with a mentality of "you can't do that, there's a rule against that" and had to slowly break out of it as much as I could. Just knowing that there's people like you out there makes me happy. I applaud your freedom.
What Kagi or anyone could work on, is an actually working version of YouTube Kids.
I literally Pi-Hole Blocked all of YouTube after my son started reading the Bible after a Minecraft Influencer started preaching throughout most of his videos to the point my son became a bit too much interested in the topic.
Not that I'm a rabid atheist or would deny my child such a thing, but if THAT can enter my 8yr olds brain via his short allowed time where he can browse by himself, i'm worried what else is coming his way through it.
I'd love to give him access to valuable videos between rules I describe by natural language and can test myself, but nothing like this exists.
Good callout, will try! I haven't considered switching tools, it's mostly convenience of just continuing, instead of stopping mid-way through and switch out the tools. But also I only code intermittently, a couple of days a week at most these days, because it's only part of what I do, so I can get to experiment with new tooling much less than i'd like.
This is only surface-level deep. Cursor already has Quotas for their paid plans and Usage-based Pricing for their larger models, which I run into and fall over to their usage based model every month.
Imo most of their incentive on context-pruning comes not just from reducing the token amount, but from the perception that you only have to find "the right way"tm to build that context window automatically, to get to coding panacea. They just aren't there yet.
For me the most interesting case is HeidiSQL. I find it easily the most useful SQL GUI client, but it crashes pretty frequently, but not frequently enough for me to stop using it over the alternatives.
I often wondered how to strike the balance right on these things, since apparently all options can lead to success.