PulseAugur
中
实时 01:32:54
English(EN) Not How Many, But Which: Parameter Placement in Low-Rank Adaptation

LoRA 参数放置影响 GRPO 微调,不影响 SFT

研究人员调查了在低秩适应(LoRA)中用于微调大型语言模型的参数放置问题。他们的研究表明,对于监督微调(SFT),LoRA 适配器 B 矩阵中可训练参数的具体放置对性能没有显著影响。然而,在基于梯度的强化学习(GRPO)下,随机参数放置无法提升基础模型,而有针对性的放置则能恢复标准的 LoRA 准确率。这种差异归因于梯度结构,SFT 梯度稳定,而 GRPO 梯度接近正交,后者需要一种基于梯度的信息方法才能有效学习。 AI

影响 确定了有效的 GRPO 微调的关键参数放置,可能优化特定 LLM 适应任务的资源使用。

排序理由 该集群包含一篇详细介绍模型微调技术新发现的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

LoRA 参数放置影响 GRPO 微调,不影响 SFT

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇详细介绍模型微调技术新发现的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
148 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Charles Lovering ·

    不是数量,而是选择:低秩适应中的参数放置

    We study the \textit{parameter placement problem}: given a fixed budget of $k$ trainable entries within the B matrix of a LoRA adapter (A frozen), does the choice of which $k$ matter? Under supervised fine-tuning, random and informed subsets achieve comparable performance. Under …