AirLLM has released updates that significantly reduce the memory requirements for running large language models, enabling powerful models to operate on consumer-grade hardware. Recent additions include support for Qwen3.8-27B, which requires only 3.33GB of VRAM, and Kimi K3 (2.8T), the largest open-source model to date, runnable on under 4GB of VRAM. These advancements are achieved through techniques like per-expert streaming, allowing models such as DeepSeek-V3 (671B) to run on approximately 12GB of VRAM. AI
IMPACT Enables running large language models on lower-spec hardware, democratizing access to advanced AI capabilities.
RANK_REASON This is an update to a tool that improves inference efficiency, not a new frontier model release or significant industry event.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →