I'm on the team. The number I'd pull out of the post: of the ~1.2s response latency, only ~100ms is us. The rest is STT, LLM, TTS, buffering between them, and whatever the user's connection adds on top. At this point the avatar renders faster than the words it's waiting for.