Llama 2 chat with vLLM and tensor parallel guide(docs.mystic.ai)
docs.mystic.ai
Llama 2 chat with vLLM and tensor parallel guide
https://docs.mystic.ai/docs/llama-2-with-vllm-7b-13b-multi-gpu-70b
https://docs.mystic.ai/docs/llama-2-with-vllm-7b-13b-multi-gpu-70b
- 7B, 1x A100, 25GB VRAM, 49 tok/s, $0.0113 /1k tok - 13B, 1x A100, 37GB VRAM, 32 tok/s, $0.0174 /1k tok - 70B, 2x A100, 150GB VRAM, 13 tok/s, $0.128 /1k tok