Zen3 isn't the best performing CPU today, though of course it'll never approach GPU level (which is pretty much ASIC level for matrix muls with special units). CPUs are getting dedicated "AI accelerators" too so it'd be interesting to compare per watt. The real limit is almost certainly memory bandwidth, not flops.
It would also be very interesting to see someone like Fabien Giesen / ryg do a maxed out AVX512 version for Zen5. His code's so fast it makes Intel 13900k's self destruct.
Yes, yes, yes! I'm absolutely ready and waiting with dual Strix Halo machines here and really want something approaching Opus at home. Speed is secondary concern for now, that would absolutely change the world.
Qwen 3.6 27b 8b quant 16b kv cache is already pretty good on the Strix.
> Most of us would struggle to imagine a world in which U.S. soldiers, police officers, Border Patrol agents and elected leaders could be dragged before an international court, tried by judges from random countries across the globe, found guilty under international laws we neither consent to nor control, and then imprisoned thousands of miles from America.
Strange how they don't have any problems with abducting people from other nations (plenty of recent examples) to stand trial in the US under US laws, complete with meme'ing about being sent to alligator Alcatraz etc.
I met him briefly at the Internet Archive during my first trip to America, and asked for an autograph. He asked for $50 and so I sheepishly slinked away because at that time I was making 20k eur. Oh well.
My man :) Alex Evans has been my gfx coding hero since the 90s with his demoscene work, and I had the privilege of working briefly with him at Lionhead.
I have a Framework Desktop as primary PC (great cooling, beautiful case with handle) and the Bosgame M5 dedicated for AI use.
I was also a bit wary about Bosgame but TBH they've been great and the machine is rock solid, if a little noisier than and not as pretty as the FD. You can just buy from them directly and be fine, best computer deal out there by a mile.
The Strix Halo mini PCs use the exact same chip, and have a much smaller footprint than any laptop. Have you seen the size of these machines? I can and have easily popped my daily driver computer into my very small backpack to attend a demoparty for example.
With the laptop you probably won't get silent operation at the peak 100-140w, i.e. you've now massively overpaid for lower performance.
I think people buying laptops for AI use are, sorry, just plain crazy. You overpay for the screen and keyboard and battery and whatever, plus you get much worse thermal performance because of basic physics (area vs volume). My Framework Desktop has a Noctua cooler which works really well.
[Tangent: all my life I've been downvoted into a smoking hole in the ground, particularly on reddit r/hardware, for questioning the wisdom of laptops for high performance computing, including gaming. Everyone insists they need the mobility, and then just leave it plugged in the whole time, absolutely refusing to admit it's about aesthetic preference.]
IIRC llama.cpp doesn't implement DSv4's compressed attention mechanism, and while it does use (credited) parts of llama.cpp, it's focused on this great model for now. Much of this is covered better in the repo's readme.
I have two 128gb Strix Halos and have been extremely excited about Antirez's (Redis author) work on DS4, especially with 4bit quant using two machines: https://github.com/antirez/ds4
Right now the speed isn't good for GLM 5.2, Deepseek V4 Flash speed is okay for me (actually reading the output) and quite usable. See kyuz0's great recent video here: https://www.youtube.com/watch?v=PkKXm_mKCCM
With a bit more speed and model improvements, local AI becomes a reasonable practical thing! The biggest problem is all the tech companies making consumer hardware completely unaffordable, and I don't think this is accidental. Look at Micron's profits and share price lately...
I got my Strix machines for ~2k eur each, best computers this 90s kid has ever owned, but those days are gone :(
Then all they do is drive the usage of open models underground (copyright infringement is illegal too, and still common), stifle US companies operating legally, and accelerate the rest of the world decoupling from the US.
I hope they do it! It will have a positive long-term effect just like the Iran war footgun accelerates renewable energy transition.
Are these models still relevant for people outside the US? I get the impression we're stuck on GPT 5.5 and Opus 4.8 pretty much permanently now, and relying on Chinese models in future.