PulseAugur
中
实时 08:27:46
English(EN) What Running Gradient Boosting Predictions at Batch Scale Costs

梯度提升预测成本由树的数量驱动,而非特征数量

本文详细介绍了大规模运行梯度提升预测的成本,强调模型中的树的数量是计算成本的主要驱动因素,而不是特征的数量。作者提供了一个成本模型,显示使用800棵树对5000万行数据进行评分,每晚约花费0.13美元,或每年49美元,假设了特定的计算价格和吞吐量。调整学习率等超参数会显著增加树的数量,从而增加成本,即使准确性提升很小。 AI

影响 为优化梯度提升模型的推理提供了成本模型,突出了AI运营商的关键成本驱动因素。

排序理由 文章提供了针对特定机器学习技术的分析和成本建模,而不是宣布新版本或事件。

在 dev.to — LLM tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

梯度提升预测成本由树的数量驱动,而非特征数量

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Commentary
文章提供了针对特定机器学习技术的分析和成本建模,而不是宣布新版本或事件。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
56 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Multigrid ·

    批量运行梯度提升预测的成本是多少

    <p>Scoring a boosted tree ensemble is cheap per row and expensive in aggregate, and almost nobody knows which of their parameters is driving the bill. It is the tree count. Here is the arithmetic that shows why, with every input stated so you can substitute your own.</p> <p>Every…