We do some on-the-fly optimizations as well (like compiling into CUDA graphs or fusing together calls) which ends up resulting (for some inference engines) faster token throughput too.
Hey! I'd love to learn more about the position - the hook @ jpl email on your website is bouncing though. Do you mind sending me contact details so we can chat?
I've followed the same path, out of curiosity do you mind if I PM you asking a few questions? I'm a couple months out of undergrad and I'd like to pick your brain a bit.