A developer has created a C-based inference engine called Picchio that allows large Mixture-of-Experts (MoE) models to run on systems with less RAM than the model size by streaming experts from disk. The project recently added support for the MiniMax-M2 model, which, when converted to INT4, is approximately 122 GB. This setup achieved about 0.48 tokens per second on a 12-core laptop with 32 GB of RAM and an NVMe SSD, demonstrating that cache size is a critical factor for performance. AI
IMPACT Enables running large AI models on consumer-grade hardware, potentially lowering the barrier to entry for AI experimentation.
RANK_REASON The cluster describes a new inference engine that enables existing models to run on less hardware, rather than a new model release or research breakthrough.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →