Researchers from UC Berkeley and UT Austin have developed FreeToken, an open-source serving engine designed to run large Mixture-of-Experts (MoE) models on single workstation GPUs, significantly reducing the hardware barrier for advanced AI. This system addresses limitations in existing engines by employing bandwidth-adaptive execution and semantic-aware caching to efficiently manage computation between the CPU and GPU. FreeToken enables models like the 753B GLM-5.2 to run at interactive speeds on consumer hardware, making powerful AI capabilities more accessible to individual developers and smaller teams. AI
IMPACT Lowers the hardware requirements for running large frontier models, potentially democratizing access for individual developers and small teams.
RANK_REASON The cluster describes a new serving engine for large language models developed by academic researchers, detailing its technical mechanisms and performance.
- Anthropic
- DeepSeek-V4 Flash
- FreeToken
- GLM-5.2
- Kimi k3
- Linux
- llama.cpp
- Microsoft Windows
- NVIDIA
- Ollama
- OpenAI
- University of California, Berkeley
- University of Texas at Austin
- Claude Code
- KTransformers
- MoE-Infinity
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →