- You value simplicity more than performance or price-to-performance.
- You accept that the hardware will depreciate rapidly.
- You’re prepared to buy two or four of them.
OR: - You want to run frontier models right now as cheaply as possible
- You want to run high-parameter models on a 15a breaker/line
Otherwise, get a normal, high-bandwidth GPU. - One Spark: More memory, but mostly enough for poor-quality, extremely low-bit quants of larger models.
- Two Sparks: Enough for mid-tier parameter models at reasonable quants, such as DeepSeek V4 Flash and HY3.
- Four Sparks: Enough for GLM 5.2 at a reasonable quant. You'll need a $1000+ switch too.
The problem is that Sparks are slow compared with almost everything else in their price range. Many factors affect inference speed, but memory bandwidth is one of the biggest. A $4,000-plus DGX Spark provides only 273 GB/s. | GPU | Memory bandwidth | VRAM | Approx. price |
| ------------------- | ---------------: | -------------: | ------------: |
| DGX Spark | 273 GB/s | ~115 GB usable | $4,000+ |
| RTX 5060 | 448 GB/s | 16 GB | $600 |
| Radeon AI Pro R9700 | 640 GB/s | 32 GB | $1,200 |
| RTX 4000 Pro | 672 GB/s | 24 GB | $2,300 |
| RTX 4500 Pro | 896 GB/s | 32 GB | $3,500 |
| RTX 3090 | 936 GB/s | 24 GB | $1,200 |
| RTX 5000 Pro | 1,344 GB/s | 48 GB | $6,000 |
| RTX 5090 | 1,792 GB/s | 32 GB | $4,000 |
| RTX 6000 Pro | 1,792 GB/s | 96 GB | $12,000 |
Yes, the Spark has substantially more memory. But going from roughly 24 GB to 115 GB does not necessarily unlock substantially better model quality. In many cases, it only lets you load heavily compressed 2-bit versions of larger models, such as DeepSeek V4 Flash, with serious quality degradation. - Perform far above what their parameter counts suggest.
- Fit comfortably within 24–32 GB of VRAM at reasonable quantization levels.
A 4-bit quant of Qwen 3.6 27b (18 GB) will out-perform a 2-bit quant of DeepSeek v4 Flash (90gb). - One RTX 5090, RTX 5000 Pro, or RTX 4500 Pro.
- Two RTX 3090s, RTX 4000 Pros, or R9700s, provided the motherboard can bifurcate two physical x16 slots into x8/x8.
If I were building a system from scratch: - A DDR4- or PCIe 4.0-era consumer CPU and motherboard that supports x8/x8 bifurcation.
- Two RTX 3090s, RTX 4000 Pros, or R9700s.
If I were already planning to buy a new Mac: - A MacBook Pro M5 with 64 GB or 128 GB of unified memory.
For context, these are the systems I currently run: - EPYC Turin with four RTX 6000 Pro Max-Qs.
- EPYC Milan with four RTX 3090s.
- AM4 with two RTX 3090s.
- AM4 with two RTX 3090s.
- Intel Raptor Lake with two RTX 5060 Ti.
- MacBook Pro M3 128GB Unified
I should have prefaced my post - I almost bought four Sparks a couple months ago, but ultimately opted to buy two more RTX 6000 Pro Max-Q's.
It was a painful choice because the two 6000's were more expensive than four Sparks, and ultimately gave me only 384 GB VRAM.
It was even more painful when GLM 5.2 was released, and a 4x Spark setup could run it at a decent quant, but 4x 6000's cannot with any headroom.
But the 6k's absolutely destroy the Sparks on prefill and inference speed. Model intelligence is compressing. The smaller VRAM pool will matter less over time than slower prefill/inference speed.
That is to say, I'm sure we'll end up with <500B parameter models that are Fable-level in the next 8 months or so. Performant quants of those will fit comfortably in 384 GB.