Glad you liked it! Would love to chat more about what kind of production workflows you are looking to improve. You could grab whichever time is convenient for you: https://cal.com/team/almanac/demo
LoCoMo is not quite our use case. It tests conversational memory, while Almanac is more for coding agents trying to find their way around a codebase.
These results show output quality at the same token budget, not token savings yet. We’re working on coding agent evals that should be more representative.