If you're in SF, you don't want to miss this.
The Qwen team is making their first public appearance in the United States, with the VP of Qwen Lab speaking at the meetup below during SF teach week.
https://partiful.com/e/P7E418jd6Ti6hA40H6Qm
Rare opportunity to directly engage with the Qwen team members.
Honestly depends on when they got in. Seed investors? They're probably fine with their preferences. Series B and beyond? That's where it gets messy. What round you thinking?
- Up to 1350+ FP8 TFLOPS on Hopper GPUs
- No heavy dependency, as clean as a tutorial
- Fully Just-In-Time compiled
- Core logic at ~300 lines - yet outperforms expert-tuned kernels across most matrix sizes
- Supports dense layout and two MoE layouts
Wan2.1, a comprehensive and open suite of video foundation models that pushes the boundaries of video generation. Wan2.1 offers these key features:
- SOTA Performance: Wan2.1 consistently outperforms existing open-source models and state-of-the-art commercial solutions across multiple benchmarks.
- Supports Consumer-grade GPUs: The T2V-1.3B model requires only 8.19 GB VRAM, making it compatible with almost all consumer-grade GPUs. It can generate a 5-second 480P video on an RTX 4090 in about 4 minutes (without optimization techniques like quantization). Its performance is even comparable to some closed-source models.
- Multiple Tasks: Wan2.1 excels in Text-to-Video, Image-to-Video, Video Editing, Text-to-Image, and Video-to-Audio, advancing the field of video generation.
- Visual Text Generation: Wan2.1 is the first video model capable of generating both Chinese and English text, featuring robust text generation that enhances its practical applications.
- Powerful Video VAE: Wan-VAE delivers exceptional efficiency and performance, encoding and decoding 1080P videos of any length while preserving temporal information, making it an ideal foundation for video and image generation.
- Efficient and optimized all-to-all communication
- Both intranode and internode support with NVLink and RDMA
- High-throughput kernels for training and inference prefilling
- Low-latency kernels for inference decoding
- Native FP8 dispatch support
- Flexible GPU resource control for computation-communication overlapping
X: https://x.com/deepseek_ai/status/1894211757604049133
Don't think the decision is based on infra, or any technical reasons.
It's more on the service support side.
How a 200-person company supports 44M iPhone users in China?