You won't like it, but the answer is Apple. The reason is the unified memory. The GPU can access all 32gb, 64gb, 128gb, 256gb, etc. of RAM.
An easy way (napkin math) to know if you can run a model based on it's parameter size is to consider the parameter size as GB that need to fit in GPU RAM. 35B model needs atleast 35gb of GPU RAM. This is a very simplified way of looking at it and YES, someone is going to say you can offload to CPU, but no one wants to wait 5 seconds for 1 token.
"The constraint system offered by Guidance is extremely powerful. It can ensure that the output conforms to any context free grammar (so long as the backend LLM has full support for Guidance). More on this below." --from https://github.com/guidance-ai/guidance/
I didn't find any more on that comment below. Is there a list of supported LLMs?
*Please note, I'm not in favor of censorship, it's just that this analogy is inaccurate
Olive Garden isn't given access to something it requires to operate at the pleasure of the government. Broadcast TV on the other hand...
All of broadcast TV is allowed because the government says it is. ABC/CBS/NBC/FOX don't own the radio spectrum they are operating on, the government does and they grant the right to use it to those companies. There's a long list of things that the government requires them to do in order to keep this pleasure. One of them used to be the Saturday morning cartoons. I miss those.
I've used Waymo countless times in SF. It's typically 15% cheaper than an Uber/Lyft and trip time/wait are generally the same. I much prefer the Waymo.
This is the model that was code named "Sonic" in Cursor last week. It received tons of praise. Then Cursor revealed it was a model from xAI. Then everyone hated it. :/ I miss the days where we just liked technology for advancement's sake.
*edit Case in point, downvotes in less than 30 seconds