AI means it has npu, Max+ is marking the memory channel, PRO is a normal label for chips that have extra security baked in, it has been this way since forever.
update: the knowledge cut off date is "unknown" now.
funny because some people downvoted me believed that there is no relation between knowledge cut off date and real world events. that's not how it works!
tested the models on aistudio. despite that the knowledge cut off is march 2026 it still knows nothing about 2025!
you can check by asking "list notable world events in 2025, only list unplanned" on aistudio. or you can ask for Charlie Kirk, it also does not know. I tried it multiple time to ensure that I didn't not get routed to older models!
> but google has search
irrelevant, without deeper knowledge about cutting edge technologies or latest libraries, all of it suggestions are crap. even you ask it to search it will still use outdated keyword thus only getting outdated information.
what a horrible article. full of misinformation and dishonesty.
1. training new base models are expensive for sure, but fine-tuning them are relatively inexpensive enough the labs can continue to do so forever. the main reason why frontier models are so good is because the massive input they generated from user usage. they are using that information to strategically build better training data. and this is why no other models can catch up, til now that is.
but if chinese models are good enough, and free to host, and cheaper to use, then the consequence is the frontier labs will lost valuable user inputs and the chinese labs will gain more. as time goes by this will be a domino effect.
2. nvidia is not only the player in the hardware scene. amd mi350p is getting popular, and huawei is pumping SuperPoDs. what does this mean for us? chinese models will surely use chinese hardware, and optimize for them. the other people will pick amd because compare to nvidia they are cheaper. with open weight models and open source inference stacks, they are freely to experiment and improve the stack, thus further lower the inference cost and nvidia dependency.
and they even plan to build their own inference hardware, too.
and nvidia loses market share meaning all the fund it gives to openai or anthropic will be cut, too.
Did they cry about it? No right? Don't apply your own standard then judge them about it, petty people.
Stackoverflow aimed to be a knowledge base. And knowledge base has a ceiling limit. They simply reached the point that almost all questions (regarding the knowledge) were asked for them. You can argue that newer or niche libraries or languages knowledge is still lacking there, but I have never seen them getting closed, just not answered.
Not worth it. I have just tried a single prompt in the web interface and it is still not finish reasoning. It thinks too much and often repeats the same stuff over and over.
Combine with the price it will surely more costly than gpt 5.6.
it is funny because nobody ever bother points out that they overcharge you for text input token price.
sure it was pretty resource intensity a few years before, but with turbo quant, sparse attention and various techniques, plus the advancing of hardware (dedicated prefill machine, memory pool for kv caching) the cost should be drastically reduced, and yet they still keep the same cost formula.
I can't help but laugh whenever someone proudly share how many billion input tokens they spent in their code sections and how much they saved with the subscription, meanwhile it is pretty much just electricity cost for the providers.
That's for the long term. Anthropic only needs short term solutions for the sake of IPO. They will do whatever they can to sabotage other companies (specially the Chinese ones) to reach the same parity with best claude models.
I doubt you can do that. MTP magic happens because for texts, we have a lot of low value fixed tokens that almost always get generated in the sequence (like punctuation, function words, language keywords etc). for most important ones (the entities, the content words, variables) you still need the full model.
so there is alwasy a maximum limit for how well MTP can do.
edit: now I read the article fully, seems like they utilize some very effective MTP algorithm. and somehow the quality is still decent enough.
though, I doubt that the quality really only drip a bit like they claimed. maybe for the benchmarks, but for general uses the heavily quantized models very often so worse result.
Is this somehow satire? This is just the dgx spark with keyboard and monitor in a convenient format. Since it has more stuff, I'm sure that the price mark up will increase too.
Up to $5000 because why not?
With that money you can build a real PC with rtx 5090!
from what I understand, it's because unlike the other models, MAI models haven't yet fine-tuned against the synthetic datasets specifically designed to boost the benchmark scores.
I personally do not like Microsoft, but congrats them to release this model.
While the scores are not good compare to other open weight model, the important thing to note is their training data (as they claimed) is very clean, without any synthetic datasets.
I bought one AMD MI50 32GB back then when they were sold rather cheap (around $150-$170). it can easily generate over 70 tokens per second for gemma 4 26B moe model (q4).
I have no doubt that we will have another wave of cheap retired server gpus just like before. And that is the time when everyone will have their own models at their home.
Or we can just buy the newest medusa halo mini pc. they will be pretty decent, too, albeit pricey.
we all know it is impossible goal to make. surely AI will be even more useful in the future, but as long as china exists and continue to undercut the price, the goal will be never meet.
> We're talking about a world where you need 5% of every knowledge workers salary to go into tokens. 20% if you're a developer.
with that much money, the companies can easily buy their own hardware and hosting free public models, no need for those expensive subscriptions.
Finally some good news. There are a lot of niche products (like handheld emulators or pocket devices) are on verge of collapse right now due to the ram price.
I was waiting for a new GPD win max with amd hx 385 or newer CPU. But they are holding the production plans right now, it sucks.
because for most people they don't need what deno promises.
me for example only use nodejs or bun to run a basic sveltekit server, so it can render the html for the first time. all core functionalities are delegated to backend services written in crystal or rust. I don't need some bloated js runtime that hoard 500MB of ram for that purpose (crystal services only take 20+ MB each).
bun promised a lean runtime, every essential functionality is written in zig to increase the speed and memory footprint. and javascriptcore also uses less memory compare to v8. the only thing we expect is for bun to stabilize and can run 24/7 without memory leaking or crashing.
I don't see anything wrong with the name.