PulseAugur
中
实时 16:28:59
English(EN) Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090

开源方案在消费级 GPU 上训练 2B LLM,成本低于 7000 美元

研究人员开发了一种开源预训练方案,显著降低了训练大型语言模型的成本,使其在消费级 GPU 上训练的成本低于 7000 美元。他们训练的 Puro-2B 模型在 RTX 5090 GPU 上训练,其性能与 Qwen2-1.5B 等更大模型相当。该研究还引入了“Puro 成本缩放定律”,估计达到 Qwen2-1.5B 性能的成本低于 5090 美元,并探讨了数据课程如何影响下游性能。 AI

影响 降低了学术界和开源社区在可访问硬件上训练大型语言模型的门槛。

排序理由 该条目描述了一篇研究论文中发布的新开源训练方案和模型。[lever_c_demoted from research: ic=1 ai=1.0]

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

开源方案在消费级 GPU 上训练 2B LLM,成本低于 7000 美元

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该条目描述了一篇研究论文中发布的新开源训练方案和模型。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
35 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    Puro-2B:Poor Lab 在 RTX 5090 上以 5090 美元的价格训练 Qwen2-1.5B

    A cost-efficient open-source pretraining recipe trains 2B-parameter models on consumer GPUs for under $7K, yielding performance near larger baselines while deriving cost scaling laws and studying data curricula.