A new inference engine called Colibri allows users to run extremely large Mixture-of-Experts (MoE) models, such as those with 744 billion parameters, on standard desktop hardware. Instead of compressing the model to fit into VRAM, Colibri employs a 'placement' strategy, similar to a Just-In-Time (JIT) compiler for weights. The dense components of the model reside in RAM, while the numerous experts are streamed from fast NVMe disk storage on demand. This approach enables the execution of massive models at a rate of one to two tokens per second, making it suitable for batch analysis or privacy-sensitive tasks, though not for interactive chat. AI
IMPACT Enables running very large LLMs on consumer hardware, potentially democratizing access for specific use cases like batch processing.
RANK_REASON The item describes a new inference engine that enables running large models on consumer hardware, which is a tool-related development rather than a frontier model release or research paper.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →