A new C-based inference engine called WASTE has been developed to run large language models, specifically the 2.78 trillion parameter Kimi k3 model, by streaming activated weights directly from NVMe storage. This approach allows the model to operate beyond available RAM by utilizing disk as an extended memory. The engine is designed to be dependency-free and embeddable, keeping the core model in memory while efficiently managing expert weights from disk. AI
IMPACT Enables running extremely large models on hardware with limited RAM, potentially democratizing access to advanced AI capabilities.
RANK_REASON The cluster describes a new software tool (inference engine) for running existing models, not a new model release or research breakthrough.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →