PulseAugur
中
实时 20:00:46
English(EN) "Strata" for GLM5.3 Flash is here for some! Project Maya

Project Maya 使海量 GLM5.3-Flash 模型能在消费级 GPU 上运行

旨在运行 3210 亿参数 GLM5.3-Flash 模型的新系统 Project Maya 已发布。该项目通过智能管理 GPU、RAM 和 NVMe SSD 上的模型层,使用户能在消费级硬件上运行如此大的模型。早期用户报告称速度显著提升,Maya 在单个 GPU 上实现了超过 30 tokens/s 的速度,同时保持了与原始 FP8 模型相比的高准确性,并提供了 OpenAI 和 Anthropic 兼容的 API 以便集成到代理中。 AI

影响 使大型、高精度模型能在消费级硬件上运行,可能使先进的 AI 能力的获取更加普及。

排序理由 这是一个软件工具,它使得大型语言模型能在消费级硬件上运行,而不是来自前沿实验室的新模型发布。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Project Maya 使海量 GLM5.3-Flash 模型能在消费级 GPU 上运行

本文如何被排名

Signal score
12 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是一个软件工具,它使得大型语言模型能在消费级硬件上运行,而不是来自前沿实验室的新模型发布。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/inthesearchof ·

    GLM5.3 Flash 的 "Strata" 已为部分用户推出!Project Maya

    <!-- SC_OFF --><div class="md"><p>I stumbled across this as I was currently having glm5.3 flash run only around 10tok/s basically unusable. With Maya i am running 30 tok/s now. GLM5.3 feels even with this low quant like a much more enjoyable model than qwen 3.8 flash next so far.…