NVLink, NVSwitch, and All That
blog.doubleword.ai5 points·by somnial··2 comments
The Anatomy of an Instruction Pipeline Hazard
hiraditya.github.io9 points·by somnial··0 comments
Width vs. Depth: Speculating on the Margin
blog.doubleword.ai17 points·by somnial··1 comments
Pushing memory bound CUDA kernels past the speed of light with data compression
fergusfinn.com2 points·by somnial··0 comments
Speculative KV coding: ~4× losslessly compressed KV cache using a small model
fergusfinn.com2 points·by somnial··0 comments
[untitled]
1 points·by somnial··0 comments
70x faster cold(ish) starts for SGLang
fergusfinn.com1 points·by somnial··0 comments
LLM powered data structures: A lock-free binary search tree
fergusfinn.com1 points·by somnial··0 comments
Parallel Primitives for Multi-Agent Workflows
fergusfinn.com1 points·by somnial··0 comments
Scheduling in LLM Inference
fergusfinn.com1 points·by somnial··0 comments
How fast can an LLM go?
fergusfinn.com2 points·by somnial··0 comments