hi!
open sourced a serving config for Kimi K2.6 from 90 tok/s 508 tok/s on 8xMI300X. Same weights / 0 quality loss.
Scaling is linear @15.8 tok/s per slot latency is constant. REpo has command launcher, Dockerfile, benchmark tool. Known limitations: BF16 KV only (FP8 crashes due to an AITER 384-expert constraint)
YARN is still in its infancy and doesn't run in production for user-facing applications at any large site. Also, there are big limitations when it comes to scaling YARN. YARN also has a lot of overhead on each machine the worker processes run on.
Twitter runs mesos in a multi 10k node cluster for user facing, production applications.
Scaling is linear @15.8 tok/s per slot latency is constant. REpo has command launcher, Dockerfile, benchmark tool. Known limitations: BF16 KV only (FP8 crashes due to an AITER 384-expert constraint)