Qwen 3.6 27B is quite good for agentic coding, and practical to run on consumer hardware. You need a system with either 32+ GB VRAM, or a unified memory system with 48+ GB VRAM and a decent integrated GPU. While not cheap, such a setup is still attainable for much of the world, and will eventually get cheaper over time. Open models hosted on non-American clouds also remain an option with a much lower barrier to entry, for cases where privacy is less critical.
I run Qwen 3.5 122B-A10B on my MacBook Pro, and in my experience its capability level for programming and code comprehension tasks is roughly that of Claude Sonnet 3.7. Honestly I find that pretty amazing, having something with capability roughly equivalent to frontier models of an year ago running locally on my laptop for free. I’m eager to try Qwen 3.6 122B-A10B when it’s released.
It’s worth also comparing Qwen 3.5, it’s a very strong model. Different benchmarks give different results, but in general Qwen 3.5, GLM 5, and Kimi K2.5 are all excellent models, and not too far from current SOTA models in capability/intelligence. In my own non-coding tests, they were better than Gemini 3.1 flash. They’re comparable to the best American models from 6 months ago.
I disagree about the image quality at typical sizes - I find JPEG-XL is generally similar or better than AVIF at any reasonable compression ratios for web images. See this for example: https://tonisagrista.com/blog/2023/jpegxl-vs-avif/
AVIF only comes out as superior at extreme compression ratios at much lower bit rates than are typically used for web images, and the images generally look like smothered messes at those extreme ratios.
JPEG-XL provides the best migration path for image conversion from JPEG, with lossless recompression. It also supports arbitrary HDR bit depths (up to 32 bits per channel) unlike AVIF, and generally its HDR support is much better than AVIF. Other operating systems and applications were making strides towards adopting this format, but Google was up till now stubbornly holding the web back in their refusal to support JPEG-XL in favour of AVIF which they were pushing. I’m glad to hear they’re finally reconsidering. Let’s hope this leads to resources being dedicated to help build and maintain a performant and memory safe decoder (in Rust?).
Those ratios seem way off if you're referring to the M1 Max and not the base M1. If we use Geekbench CPU performance, the Ryzen 9 7945HX (which is from 2023) is around 12% faster single core and 32% faster multicore than the M1 Max (which is from 2021). If you look at the 2024 M4 Max, it's substantially faster than the Ryzen and Intel you mentioned.
153 GB/s is not bad at all for a base model; the Nvidia DGX Spark has only 273 GB/s memory bandwidth despite being billed as a desktop "AI supercomputer".
Models like Qwen 3 30B-A3B and GPT-OSS 20B, both quite decent, should be able to run at 30+ tokens/sec at typical (4-bit) quantizations.
It's colleges they they have been clamping down on, as they were bringing in absolutely massive numbers of mostly Indian students who were coming mainly to work in low-end jobs and get out of India rather than to legitimately study.
The number of graduate students being allowed in hasn't changed significantly, and undergraduate university students are also continuing to be brought in at rates similar to pre-pandemic times.
Mistral models are largely along the likes of what you were asking for. However, Grok (any version) absolutely is not a “don’t say gay” model, it talks about sexuality of all forms quite openly and fairly and is happy to produce creative content of any level of explicitness about these topics. It’s the least censored unmodified model I’ve encountered on any topic. People dismiss Grok as a Nazi model based on Musk’s politics without using it themselves.
That’s the thing - old German compact luxury sedans from the 80s had the control feel, balance, and light weight you get from a Porsche, while also being practical family cars. There’s nothing like that made today. They were also decently safe and comfortable and reliable and generally just good.
Also the bigger ones like the W126, while not as light and agile as a Porsche or Lotus, still had similar control feel, very comfortable and spacious interiors, and could glide over the worst most broken and potholed roads better than any modern car I’ve driven. They’re also much simpler than any modern luxury cars, much less to break, and they just keep going and going as long as you take basic care. From personal experience, a much younger used W220 or W221 S class needs far more maintenance and repair than an old W126.
The more powerful but still reliable engines and nicer transmissions of the late W140 or W220 would be nice to have in a W126 though. My problem with the newer S classes is the complexity and fragility of the rest of the car.
Of course, these are 40 year old cars and need more care and maintenance than a new car, but they’re not too bad either as long as you get a good example of the car. They’re pretty reliable once sorted, and can last a very long time and very high mileage as long as they’re at least somewhat cared for.
Aside from the Mazda MX-5 (which isn’t the most practical car), almost all small, simple, and light cars made today are econoboxes. They’re not designed to have the rich control feel, balanced and satisfying handling near the limits, responsiveness, material quality, suspension sophistication, etc. compared to say German luxury compact cars of the 1980s (BMW E30 or M-B W201). Even cars like 90s Hondas, while front wheel drive and built to a much lower price point, had rich control feel, liveliness, and agility that modern cars don’t give.
Modern luxury cars from essentially all brands around the world have become huge, heavy, numb, and over-complicated. They’re much faster and quieter than the say the old Benzes and BMWs of the 80s, but they don’t have the fun raw feel, small size, light weight, tossability, and simplicity of the old cars.
A BMW E30 or M-B W201 have a weight somewhere between a Mazda MX-5 and Subaru BRZ, but are far more practical than either for passengers and cargo despite being around the same width and only slightly longer.
The only modern cars with similar size and weight are some European market compact cars and econoboxes like the Mitsubishi Mirage, Nissan Micra, and Chevy Spark (which are also disappearing from North America). For steering feel, handling, general raw and connected driving feel, powertrain responsiveness, and interior quality, these modern economy cars can’t compete. Some of the European market specific B-segment cars come closest to those older compact luxury cars, but they still don’t match them for the qualities I described.
Kit cars generally suck from a practical perspective compared to well engineered 80s/90s cars and aren’t a very practical option either.
In general, I agree. However, many older cars were small, light, simple, and raw - characteristics that have largely disappeared from modern cars. Automatic transmissions from the mid-90s and earlier generally sucked, though good old manual transmissions are not much different from good modern ones.
As an example, I owned a W126 S class from the late 80s, and it was fun in its unique way and no modern cars replicate its experience. It had somewhat heavy and very feedback rich steering feel, and Porsche-like firm and tactile pedal feel, while having a super supple ride over the most awful roads with SUV-like ground clearance and tremendous suspension travel. The car was also super simple and reliable; my 300SE had nearly 400k km with all original powertrain when I sold it, it never let me down, and it weighed less than a modern A class or CLA. While not as safe as modern cars, it was exceptionally safe for its era and comparable to normal cars of the early 2000s for crash structure safety.
The W140 (I used to own one too) had a much better powertrain, but it lost the raw tactile scrappy nature of its predecessor, and nor could it handle super awful potholed roads as well as the W126. There are no modern cars that combine the rich raw tactile control feel and super supple ride the W126 had.
Look at cars like the BMW E30, or Mercedes-Benz 190E (W201), or the superbly engineered workhorses that the W123 and W124 were. There are no modern cars that replicate the genuinely delightful driving experience of those.
While cloud models are of course faster and smarter, I've been pretty happy running Qwen 3 Coder 30B-A3B on my M4 Max MacBook Pro. It has been a pretty good coding assistant for me with Aider, and it's also great for throwing code at and asking questions. For coding specifically, it feels roughly on par with SOTA models from mid-late 2024.
At small contexts with llama.cpp on my M4 Max, I get 90+ tokens/sec generation and 800+ tokens/sec prompt processing. Even at large contexts like 50k tokens, I still get fairly usable speeds (22 tok/s generation).
Privacy, both personal and for corporate data protection is a major reason. Unlimited usage, allowing offline use, supporting open source, not worrying about a good model being taken down/discontinued or changed, and the freedom to use uncensored models or model fine tunes are other benefits (though this OpenAI model is super-censored - “safe”).
I don’t have much experience with local vision models, but for text questions the latest local models are quite good. I’ve been using Qwen 3 Coder 30B-A3B a lot to analyze code locally and it has been great. While not as good as the latest big cloud models, it’s roughly on par with SOTA cloud models from late last year in my usage. I also run Qwen 3 235B-A22B 2507 Instruct on my home server, and it’s great, roughly on par with Claude 4 Sonnet in my usage (but slow of course running on my DDR4-equipped server with no GPU).
On my M4 Max MacBook Pro, with MLX, I get around 70-100 tokens/sec for Qwen 3 30B-A3B (depending on context size), and around 40-50 tokens/sec for Qwen 3 14B. Of course they’re not as good as the latest big models (open or closed), but they’re still pretty decent for STEM tasks, and reasonably fast for me.
I have 128 GB RAM on my laptop, and regularly run multiple multiple VMs and several heavy applications and many browser tabs alongside LLMs like Qwen 3 30B-A3B.
Of course there’s room for hardware to get better, but the Apple M4 Max is a pretty good platform running local LLMs performantly on a laptop.
You should use flash attention with KV cache quantization. I routinely use Qwen 3 14B with the full 128k context and it fits in under 24 GB VRAM. On my Pixel 8, I've successfully used Qwen 3 4B with 8K context (again with flash attention and KV cache quantization).
They just recently released the r1-0528 model which was a massive upgrade over the original R1 and is roughly on par with the current best proprietary western models. Let them take their time on R2.