PulseAugur
中
实时 05:39:55
English(EN) Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features

新的SCALE方法通过控制特征使用来增强SFT

研究人员推出了一种名为SCALE(Selective Control of Adaptation via Local Entropy,通过局部熵选择性控制适应)的新型监督微调(SFT)方法,旨在通过控制学习到的特征的使用方式来改进模型适应性。与专注于抑制或放大更新的现有方法不同,SCALE冻结了预训练模型和SFT增量,学习门控机制,根据熵减少来抑制、反转或外推特征。这种方法在多个Qwen模型(包括Qwen2.5-Math-1.5B、Qwen2.5-Math-7B和Qwen3-4B-Base)的数学推理和代码生成任务上表现出卓越的性能,优于强大的基线方法,同时保持了通用能力。 AI

影响 这项研究可能带来更有效的微调技术,从而提高模型在数学推理和代码生成等专业任务上的性能。

排序理由 该条目描述了一篇详细介绍一种新型监督微调方法的最新研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的SCALE方法通过控制特征使用来增强SFT

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该条目描述了一篇详细介绍一种新型监督微调方法的最新研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
11 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    重新思考SFT的Token重加权:抑制、反转和外推学习到的特征

    Supervised fine-tuning (SFT) learns most aggressively from tokens that the model deems least likely. This helps acquire new behaviors, but also amplifies noisy or conflicting supervision and can overwrite useful pretrained knowledge. Through a unified policy-loss view, we revisit…