PulseAugur
中
实时 22:11:40
English(EN) Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?

分离的嵌入层提升私有LLM训练的准确性和效率

一篇新发表在arXiv上的研究论文探讨了在差分隐私随机梯度下降(DP-SGD)下对解码器专用大型语言模型(LLM)进行微调时,权重绑定的影响。研究发现,在SST-2和QNLI等基准测试中,分离输入和输出嵌入层的模型始终优于权重绑定模型,准确率提升高达4.74%。此外,分离的嵌入层通过启用幽灵裁剪(ghost clipping),使得DP-SGD训练更节省内存,与权重绑定模型相比,内存使用量降低了60%以上。 AI

影响 分离的嵌入层为私有LLM微调提供了一种更有效、更高效的方法,可能影响未来的模型设计。

排序理由 该集群包含一篇详细介绍LLM架构和训练方法研究结果的论文。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

分离的嵌入层提升私有LLM训练的准确性和效率

本文如何被排名

Signal score
2 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含一篇详细介绍LLM架构和训练方法研究结果的论文。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
1 days old
Coverage has settled into its steady-state source set.

完整方法见我们的编辑标准。

报道来源 [2]

  1. arXiv cs.LG TIER_1 English(EN) · Razan El Mais, Ali Chehab, Ibrahim Issa, Razane Tajeddine ·

    在DP-SGD下,权重绑定对于私有环境中的Decoder-Only LLM是否仍然有益?

    arXiv:2609.40335v1 Announce Type: new Abstract: Differentially Private Stochastic Gradient Descent (DP-SGD) is a leading approach for privacy-preserving fine-tuning of large language models (LLMs). Many decoder-only LLMs employ weight tying between input and output embeddings, a …

  2. Mastodon — mastodon.social TIER_1 English(EN) · notatechguy ·

    权重绑定使私有LLM训练损失4.74个准确率点 解绑仅解码器LLM中的输入和输出嵌入可将准确率提高多达4.74个百分点

    Weight tying costs private LLM training 4.74 accuracy points Untying input and output embeddings in decoder-only LLMs lifts accuracy up to 4.74 percentage points under differential privacy fine-tuning. https://www. notatechguy.com/weight-tying-c osts-private-llm-training-4-74-acc…