PulseAugur
实时 01:31:11
English(EN) # AMD # Halo2 is pulling 200 tokens per second in your home. # Opus5 about 55, 150 tokens per second downhill with a wind in its back. For $10K, you can get a h

AMD Halo2 处理器在本地 LLM 推理中实现每秒 200 个 token

据报道,AMDHalo2 AI 处理器在本地 LLM 推理中能够达到每秒 200 个 token 的速度。相比之下,Opus5 模型大约以每秒 55 个 token 的速度运行,在最佳条件下为每秒 150 个 token。估计花费 1 万美元,就可以在家庭 AI 系统中使用这些处理器运行一个拥有 1200 亿参数的开源 LLM,而无需数据中心。 AI

影响 这一发展表明在消费级硬件上本地运行大型语言模型的可行性增加,可能减少对基于云的 AI 服务的依赖。

排序理由 该项目讨论了特定硬件产品在运行 AI 模型方面的性能,属于“工具”类别。

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

AMD Halo2 处理器在本地 LLM 推理中实现每秒 200 个 token

报道来源 [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    # AMD # Halo2 在您家中每秒可处理 200 个 token。# Opus5 在顺风顺水的情况下,每秒约可处理 55 至 150 个 token。1 万美元即可获得一台...

    # AMD # Halo2 is pulling 200 tokens per second in your home. # Opus5 about 55, 150 tokens per second downhill with a wind in its back. For $10K, you can get a home # AI brain that can run 120Billion # OpenWeight # llm model that does not need a # Datacentre https://www. amd.com/e…