Kimi K3 2.8T: Frontier AI on 80x RTX 5090 GPUs
Kimi K3, a 2.8T parameter open-weight model, runs on 80x RTX 5090s at 20 tok/s—the first frontier AI served without HBM. Full hardware comparison inside.
Kimi K3: A 2.8 Trillion Parameter Open Model
Kimi K3 represents a major milestone in open-weight AI development. With 2.8 trillion parameters and 104 billion activated parameters using a Mixture-of-Experts architecture, it rivals proprietary models like Claude and GPT-5.6 Sol. The model features native vision capabilities and a 1-million-token context window, built on Delta Attention and Attention Residuals architecture. Stable LatentMoE activates 16 of 896 routed experts per token, delivering approximately 2.5× better scaling efficiency than its predecessor Kimi K2. The technical report highlights performance across coding, agentic tasks, knowledge reasoning, and vision benchmarks, positioning it as frontier-level open intelligence accessible to researchers and developers worldwide.
Running on 80x RTX 5090 GPUs Without HBM
The deployment configuration showcases a groundbreaking approach: 80x RTX 5090 GPUs using GDDR7 memory with no HBM, connected via 25GbE Ethernet without NVLink. This setup achieves 2.56 TB total VRAM and 143 TB/s aggregate memory bandwidth. The tweet author reports 20 tokens per second for single-stream inference on day one, untuned—a figure expected to climb as optimizations are applied. For context, the same team previously improved GLM-5.2 from 30 to 110 tok/s on identical hardware. This marks the first time frontier-level intelligence has been served using consumer-grade interconnects instead of expensive HBM-equipped server infrastructure, potentially democratizing access to large-scale AI deployment.
Hardware Configurations Compared
The comparison table reveals stark infrastructure trade-offs. Official serving recipes using 32x H100 (80G) GPUs across 4 nodes achieve 107 TB/s bandwidth with HBM3 and NVLink + InfiniBand interconnects. The 16x H200 (141G) configuration delivers only 77 TB/s across 2 nodes despite larger per-GPU memory. Interestingly, the 16x B200 (192G) setup leads in total VRAM (3.1 TB) and bandwidth (128 TB/s), while single-node configurations like 8x B300 (288G) and 8x MI355X (288G) both provide 2.3 TB VRAM and 64 TB/s bandwidth. The RTX 5090 configuration's 143 TB/s bandwidth—achieved without HBM or NVLink—demonstrates how massive GPU count can compensate for per-device limitations, offering a cost-effective alternative for teams without access to datacenter-grade hardware.
Benchmark Performance Across Tasks
The technical report's benchmark charts show Kimi K3 competing directly with frontier models. In coding tasks like DeepSWE (37.4), TerminalBench 2.1 (89.8), and FrontierSWE (35.6), it trails Fable 5 and GPT-5.6 Sol but significantly outperforms Qwen 4.8, GPT-5.5, and QLM-5.2. On Kimi Code Bench 2.0 (Internal) it scores 76.9, and ProgramBench shows 77.4. For agentic workloads, BrowseComp yields 91.2 while AutomationBench achieves 39.4. Vision tasks include GDPixel-AA v2 Elo and JobBench (57.4). The CharXiv and ZebraBench evaluations demonstrate knowledge reasoning capabilities. Notably, Fable 5 results are marked with caveats about Archive-Ariadne penalties, and all Fable 5 results exclude certain potential challenges, making direct comparisons nuanced but highlighting Kimi K3's competitive standing among open models.
Implications for Open AI Development
Kimi K3's release with full model weights represents a philosophical and practical shift in AI accessibility. While proprietary models maintain performance leads, the gap is narrowing—Kimi K3 consistently outperforms other open and some proprietary alternatives across the benchmark suite. The ability to run frontier-scale models on commodity hardware like RTX 5090s, without requiring HBM or specialized interconnects, lowers the barrier for research institutions, startups, and independent developers. The model's architecture innovations—Delta Attention, Stable LatentMoE, million-token context, and compositional generalization—are now available for study and adaptation. As optimization continues and inference speed improves from the initial 20 tok/s baseline, Kimi K3 could accelerate broader deployment and adoption of truly open frontier intelligence across domains.
🎯 Key Takeaways
- Kimi K3 is a 2.8T parameter open-weight model with 104B activated parameters and 1-million-token context
- First frontier AI served on 80x RTX 5090 GPUs without HBM, achieving 143 TB/s bandwidth via GDDR7 and Ethernet
- Day-one inference at 20 tok/s single-stream, untuned, with optimization potential shown by GLM-5.2's 30→110 tok/s improvement
- Competitive benchmarks across coding, agentic, knowledge, and vision tasks versus Claude, GPT-5.6 Sol, and other frontier models
💡 Kimi K3 demonstrates that frontier-level AI can be served on accessible hardware without datacenter-exclusive components like HBM or NVLink. The 2.8T parameter model achieves competitive performance across diverse benchmarks while running on 80x RTX 5090 GPUs, opening new possibilities for open research and deployment. With full model weights released and inference optimizations underway, Kimi K3 represents a significant step toward democratizing access to state-of-the-art AI capabilities and accelerating the broader adoption of open frontier intelligence.