Project Maya, a new system designed to run the 321-billion-parameter GLM5.3-Flash model, has been released. This project allows users to run such a large model on consumer-grade hardware by intelligently managing model layers across GPU, RAM, and NVMe SSD. Early users report significant speed improvements, with Maya achieving over 30 tokens/s on a single GPU and maintaining high accuracy compared to the original FP8 model, while also offering OpenAI- and Anthropic-compatible APIs for agent integration. AI
IMPACT Enables running large, high-accuracy models on consumer hardware, potentially democratizing access to advanced AI capabilities.
RANK_REASON This is a software tool that enables running a large language model on consumer hardware, rather than a new model release from a frontier lab.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →