Techniques for more efficient LLM serving (up to 10x)(octo.ai)
octo.ai
Techniques for more efficient LLM serving (up to 10x)
https://octo.ai/blog/acceleration-is-all-you-need-techniques-powering-octostacks-10x-performance-boost/
1 comments
Interesting benchmarks. I'm surprised vLLM is performing so poorly given all the corporate investment/attention it's getting. Is it because the contributions are more distributed than the average LLM team?