PulseAugur
中
实时 20:38:08
English(EN) vLLM ignores LoRA rank_pattern and alpha_pattern and serves the adapter at the wrong scale

vLLM 0.30.0 错误缩放 LoRA 适配器,导致模型性能下降

vLLM 版本 0.30.0 中的一个错误导致其错误地缩放 LoRA 适配器,从而导致性能显著下降。该问题源于 vLLM 忽略了适配器配置文件中的 `rank_pattern` 和 `alpha_pattern` 设置,这些设置旨在允许适配器内的不同模块使用不同的缩放因子。相反,vLLM 应用了仅从基础 `r` 和 `lora_alpha` 值派生的单一、统一的缩放因子。事实证明,这种错误的缩放会在混合专家模型上将困惑度(perplexity)提高约 40%,并在较小的测试模型上导致比 PEFT 库大 70 倍的提示对数概率偏差。这个问题很微妙,可能会被忽视,因为 vLLM 仍然会加载适配器而不会发出警告,并且生成的文本不会立即显得不正确。 AI

影响 vLLM 中的此错误可能导致使用 LoRA 适配器的模型性能显著下降,从而可能影响推理质量和效率。

排序理由 错误报告,详细说明了推理服务框架中的不正确行为。

在 dev.to — LLM tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

vLLM 0.30.0 错误缩放 LoRA 适配器,导致模型性能下降

本文如何被排名

Signal score
4 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
错误报告,详细说明了推理服务框架中的不正确行为。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准。

报道来源 [1]

  1. dev.to — LLM tag TIER_1 English(EN) · The Homelab Postmortem ·

    vLLM 忽略 LoRA rank_pattern 和 alpha_pattern,并以错误的比例提供适配器

    <p><strong>TL;DR</strong>: A PEFT LoRA adapter can give individual modules their own rank and alpha through <code>rank_pattern</code> and <code>alpha_pattern</code> in <code>adapter_config.json</code>, and PEFT scales each module by its own <code>alpha_m / r_m</code>. vLLM 0.30.0…