A new open-source project called TurboFieldfare allows the Gemma 4 26B-A4B model to run on Macs with as little as 2GB of RAM. Developed in Swift and Metal, the runtime achieves this by keeping only the core model components and KV cache in memory, streaming necessary experts from SSD as needed. This innovation enables larger language models to be accessible on consumer hardware with limited memory. AI
IMPACT Enables running larger LLMs on consumer hardware with limited memory, potentially increasing accessibility.
RANK_REASON This is a third-party tool/runtime for an existing model, not a release from a frontier lab.
Read on Hacker News — AI stories ≥50 points →
- Apple Silicon
- Gemma 4 26B-A4B
- llama.cpp
- M2 MacBook Air
- Mac
- macOS
- Metal
- Mlx
- OpenAI
- Swift
- TurboFieldfare
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →