bigger change here might not be model quality, but debuggability.
once you hide the reasoning, remove the knobs, and let the model choose its own effort, it gets much harder to tell whether the model got worse or just got harder to inspect.
subprime mortgages sprinkled on top of prime ones, treated as prime ones. because they were printing money. subprime code sprinkled on the backbone of software we use everyday. because they are printing code. reckoning
Important to note that this model excels in reasoning capabilities.
But it was on purpose not trained on the big “web crawled” datasets to not learn how to build bombs etc, or be naughty.
So it is the “smartest thinking” model in weight class or even comparable to higher param models, but it is not knowledgeable about the world and trivia as much.
This might change in the future but it is the current state.
This "vibe" check that it's even better than GPT-4 Turbo is not what its Elo rating shows on the Chatbot Arena based on not 1 but thousands of user votes.
GPT-4 (Turbo) is in a league of its own still.
This is based on users choosing the better from 2 models at a time, and calculating an ELO rating from who-beats-who.
BYOT - bring your own tests style.
Gives a better picture of real-world performance and more robust against contamination.
They collected over 6000 and 1500 votes for Mixtral-8x7B and Gemini Pro.
While ELO ratings are widely used to rank performance in Chess or among sports teams, here's a disclaimer by the makers of the leaderboard:
---
> Please note Arena is a "live eval" and pretty much a sampling process to estimate models capability.
> That's why we show the confidence intervals through bootstrapping. Statistically, these models (e.g., GPT-3.5, Mixtral, Gemini Pro) are very close and only looking at their ranking can be misleading.