The 15% number is based on two runs from the ./experiment folder, on an instance from the SWE Bench Pro benchmark. I picked two trajectory logs, from the pro_pilot_openlibrary_haiku_v2 folder, that I consider to be representative of the token savings one can expect. There's huge variance just between individual runs of the same model on the same instance, though, so it's somewhat difficult (and expensive, at API rates) to cut through the noise. But I ran over 80 trials, and I believe that number is accurate.
And I'm publishing the "token science" research in the repo itself, so anyone who's curious about it can review the logs and inspect the python scripts that I used. And if someone out there has the time, inclination, and money, to run some of those experiments themselves, it'd be great to get consensus on some of those results.
And if you'd like to share feedback directly, rather than publicly, my email is on my profile page.
And it reads to me like they have some other reason to move on from SWE Bench Pro, but they don't want to say what it is. They say right up top, "~30% of the tasks are broken." But that leaves ~70% un-broken, which seems pretty good to me. It would be nice if they would also say: "Here's the list of instances that are broken: <CSV>". Or, "Here's the subset of SWE Bench Pro we will use going forward." They're letting the perfect be the enemy of the good.
Still plugging away on SourceMinder (https://github.com/ebcode/SourceMinder). I know, I know, everyone and their brother is working on token-saving schemes for LLMs. But I’ve found it to be useful even without the LLM. I’m working on a proper website for it now where you can try it out in the browser (wasm port) — try before you don’t buy (it’s GPL). Feedback welcome.
I’m working on a tool that is a more token-efficient code search than grep. I don’t have hard numbers yet, but it’s been working for me to get longer sessions. https://github.com/ebcode/SourceMinder
SourceMinder: A “code index” tool that finds symbols in a codebase and creates a single table sqlite database for the index. It uses tree-sitter to parse the AST and add the symbols and what they are (function, class, argument, etc) to the db. I currently have it working with TypeScript, C, Go and PHP. I’m working on adding Perl next, after someone requested it here on HN.
Yep! Java, Perl, and Ruby are all on the to-do list. I’ll start with Perl since you asked. And it takes me about two weeks to implement a new language. Check back in January.
> A 10% improvement every month gets to be a 10x improvement in (math...)
1.1^24=9.85, so yeah, if you could reliably get a 10% speed-up each month, you’d get to 10x in roughly 2 years. (But I’d expect the speed-up per month to be non-linear.)
I picked these up at a used bookstore ages ago, since they had the three-volume set. My recommendation would be to familiarize yourself with just the table of contents that’s printed on the binding, and when you come across something adjacent in your day-to-day work (e.g. Search), review the papers in that section. Those books are an excellent snapshot of the field at the time.