Flyweight, an open-source C++/CUDA engine, has been released on PyPI, designed to run Mixture of Experts (MoE) models that exceed a single GPU's VRAM by utilizing system RAM. The engine supports various models including Qwen, DeepSeek-V4, and Gemma, and offers features like OpenAI/Anthropic-compatible APIs and a chat UI. Developers are seeking contributors to expand hardware support, particularly for AMD and macOS systems, and to optimize CPU expert kernels. AI
影响 Enables running larger MoE models on consumer hardware, potentially lowering the barrier to entry for advanced AI experimentation.
排序理由 Release of an open-source engine for running large models.
- CPP
- CUDA
- DeepSeek-V4
- flyweight
- Gemma
- graphics processing unit
- MoE models
- Python Package Index
- Qwen
- system RAM
- VRAM
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →