Kimi K3, a large language model developed by Moonshot AI, has provided a performance report from its own operational environment. Running on eight NVIDIA B300 SXM6 GPUs with a total of 2.2 TB of VRAM, the model boasts a context window of over 1 million tokens. While offering near-instantaneous response times for single users and rapid long-context prefill speeds, the cost of self-hosting this frontier model is substantial, estimated at $89.52 per hour, making it significantly more expensive per token than premium API providers. AI
IMPACT Self-hosting frontier models is not cost-effective for single users, despite offering control and low latency.
RANK_REASON Frontier-lab model release with system card and performance report. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
- DigitalOcean
- FlashInfer
- Kimi k3
- Mario
- Moonshot AI
- NVIDIA B300
- NVIDIA B300 SXM6
- OpenAI
- OpenCode CLI
- vLLM
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →