After visiting Anthropic, OpenAI and Cursor offices two weeks ago, it became clear the frontier teams are working differently, and it's all about the way they validate their changes.
We ran the exact same Amazon shopping task with 9 leading AI models in the browser. Same site, same steps, same environment. Only the model changed. A few things stood out:
1. Fastest model: 70 seconds
2. Slowest model: 340 seconds
3. Cost range: $0.03 to $1.04
4. Only 2 of 9 models picked the right product!
Regarding the icons, I was trying to build the tables with some sort of visual cues, if it's not helpful, it'll be removed, @cimmanom what do you mean by the second part?
@partycoder thanks for the guidance, we'll work hard to get all those data points, i'd even like to add an ability to for advanced user to request a customized summary what'd you think about that? Still looking into figuring out tests though, for now, i've been thinking about importing the tags from the readme
Guys as the author of this project one of the guidelines i had was to create it as modular and familiar as i can, so whenever I could I tried using a popular module(mongoose, passport etc.)
I deeply encourage anyone who wants to break, replace, remove, add any part of this project to go ahead and do so! Like any recipe this is just our serving suggestion.