AA isn't the best way to measure relative cost in real world use because some of those benchmark questions are extremely hard for the models. Some models give up quickly on hard questions, other models spin their wheels for a long time before declaring defeat (or getting the answer on token 200k!).
A useful measure of real world cost (complementary with total cost like they already report, of course) would be "cost for correct answers". You could look at the ratio between the two costs to get a measure of laziness which many would find quite useful.
It seems unlikely that anyone will offer 3.0 for less money than Moonshot does, though you will likely be able to find a host (W&B?) that serves it at quite a high token/s.
You need to compare cost per task buddy boy. Cost per token doesn't tell you much when you don't know how many tokens a model will use to accomplish a task
This is huge. I've always wanted to use Git in the terminal but never been happy with the underlying language that other TUIs were written in. Now that I know that I'm using developer-managed memory I can be much more comfortable and confident changing between branches, pushing, and pulling, and even merging code. Thanks Simone!
Agreed 100%. This guy thinks there's a limit on the demand for intelligence. You think that Fable 7 which can run a billion dollar corporation on its own has no consumer demand just because we have fable 5 at 9k tok/s? Who do you think will be the biggest customer of such a model? Fable 7, obviously.
The assumption that LLMs can't answer this question. You can pay approximately $0.000001 to correctly answer this question in approximately 50 milliseconds using Google's cheapest model.
Are emotions actually salient parts of text or is that just because you have a human brain which is tuned to recognize emotions?
You're using "I find it easy to recognize emotions in text therefore it is a simple task", but we know for a fact that some tasks which are easy for humans are hard for LLMs, like counting objects in an image, while other tasks are easy for humans and easy for LLMs, like adding single-digit numbers.
It's not readily apparent to me that precisely modelling emotional state of the characters in a piece of text is the second and not the first, which you seem to assume. In fact, the work as presented seems to indicate that it's a class much closer to the first.