PulseAugur
EN
LIVE 12:01:55

FreeToken enables large MoE models on single workstation GPUs

Researchers from UC Berkeley and UT Austin have developed FreeToken, an edge-native Mixture-of-Experts (MoE) serving engine designed to run large language models on single workstation GPUs. This system addresses the challenge of deploying powerful models like GLM-5.2 (753B parameters) on consumer hardware by treating a personal machine as a unified inference platform. FreeToken optimizes computation and model state mapping across available GPU, CPU, and memory resources, enabling interactive speeds for models that would typically require datacenter-class infrastructure. The engine is open-source and available for local deployment, making advanced AI capabilities more accessible to individual developers and smaller teams, particularly in sensitive industries like healthcare and legal where data privacy is paramount. AI

IMPACT Enables deployment of large MoE models on consumer hardware, potentially lowering costs for developers and small teams.

RANK_REASON The article describes a new serving engine that enables existing models to run on consumer hardware, rather than a new model release or fundamental research breakthrough.

Read on MarkTechPost →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

FreeToken enables large MoE models on single workstation GPUs

COVERAGE [1]

  1. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Meet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU

    <p>FreeToken splits MoE cache misses between PCIe fills and CPU execution using measured bandwidths, unlocking frontier models locally</p> <p>The post <a href="https://www.marktechpost.com/2026/08/23/meet-freetoken-an-edge-native-moe-serving-engine-that-runs-753b-glm-5-2-on-a-sin…