English(EN)Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning
新研究探索自监督学习在公平性、效率和多样化输出方面的应用
作者PulseAugur 编辑部·[14 个来源]·
多篇研究论文探讨了自监督学习(SSL)的进展,这是一种在无标签数据上训练模型的技术。其中一项研究FairSSL,通过利用数据异质性和主题感知正则化,引入了一个框架来提高多模态SSL的公平性。另一篇论文研究了增加网络宽度如何使贪婪的逐层训练在SSL中能与端到端反向传播相媲美,尤其是在更宽的网络中。进一步的研究深入探讨了哪些下游任务最能从SSL中受益,发现它在异常检测和分类方面有效,但在预测方面效果不佳。此外,一种名为DRY-SFT的新方法旨在通过微调模型以生成各种正确的解决方案,来提高可验证领域(如编码)的输出多样性和覆盖范围。最后,一项研究提出了用于连续视频流的StreamMAE,通过流感知正则化来调整MAE重建目标,以实现具有竞争力的性能。
AI
arXiv:2508.16748v2 Announce Type: replace-cross Abstract: Prevalent multimodal self-supervised learning (SSL) methods rely on the redundancy assumption: that different views share substantial task-relevant information. We argue that this assumption fails in complex, real-world se…
arXiv cs.AI
TIER_1English(EN)·Syon Mansur, Joel Zylberberg·
arXiv:2610.00753v1 Announce Type: cross Abstract: End-to-end backpropagation has been the dominant mode of training in deep learning, allowing for the coordination of parameter updates across layers of a neural network. Prior studies have explored alternative -- and, in some case…
arXiv cs.AI
TIER_1English(EN)·Achleshwar Luthra, Lucas Bryant, Tracy Zhu, Tomer Galanti·
arXiv:2609.38393v1 Announce Type: cross Abstract: Same-instance self-supervised learning (SSL) learns representations by enforcing consistency across two views of the same underlying instance. This principle alone, however, does not determine which downstream tasks remain recover…
arXiv:2605.19462v2 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) assumes that solving pretext tasks on unlabeled data yields representations that transfer effectively across downstream applications via linear probing or fine-tuning. While this paradigm has…
arXiv cs.CL
TIER_1English(EN)·Eric Fithian, Kirill Skobelev, X. Y. Han·
arXiv:2609.31688v2 Announce Type: replace Abstract: In verifiable domains such as math and coding, finding one correct solution among many attempts can matter more than the pass rate of each attempt. Post-training can concentrate large language model outputs around a few modes, w…
End-to-end backpropagation has been the dominant mode of training in deep learning, allowing for the coordination of parameter updates across layers of a neural network. Prior studies have explored alternative -- and, in some cases, simpler -- training mechanisms, showing that th…
arXiv cs.AI
TIER_1English(EN)·Fabian A. Mikulasch, Friedemann Zenke·
arXiv:2609.37789v1 Announce Type: cross Abstract: Self-supervised learning (SSL) by predicting in latent space, without generating the input data itself, learns highly abstract, useful representations. Intuitively, this success is often attributed to its ability to discard nuisan…
arXiv cs.LG
TIER_1English(EN)·Thomas Deixelberger, Markus Steinberger·
arXiv:2609.37330v1 Announce Type: cross Abstract: Clearing fog, rain or snow from footage, or turning renders into photographs, must remove the source domain and keep the scene. Unpaired translators carry it through because their generator sees the source appearance (pixels, a ne…
arXiv cs.LG
TIER_1English(EN)·Akhlaqur Rahman Sabby, Yi Sui, Tongzi Wu, Jesse C. Cresswell, Ga Wu·
arXiv:2510.01345v2 Announce Type: replace Abstract: Self-supervised representation learning (SSRL) has demonstrated remarkable empirical success, yet its underlying principles remain insufficiently understood. While recent works attempt to unify SSRL methods by examining their in…
Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are independently sampled and globally shuffled across epochs. We study self-supervised learning from continuous video streams, where f…
arXiv cs.CV
TIER_1English(EN)·Qianxin Xia, Jiawei Du, Yuhan Zhang, Xin Zhang, Xuewan He, Wenbo Jiang, Jielei Wang, Tao Luo, Guoming Lu·
arXiv:2602.05391v3 Announce Type: replace Abstract: Dataset distillation seeks to synthesize a compact surrogate dataset that enables performance comparable to training on the original dataset for downstream tasks. For the scenario where pre-trained self-supervised models serve a…
arXiv:2609.40347v1 Announce Type: new Abstract: We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of relying on heavy 3D architectures or reconstruction-based autoencoders for learnin…
arXiv cs.CV
TIER_1English(EN)·Ivan Martinovi\'c, Lukas Knobel, Yuki M. Asano·
arXiv:2609.40333v1 Announce Type: new Abstract: Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are independently sampled and globally shuffled across epochs. We study self-supervised…
arXiv cs.CV
TIER_1English(EN)·Anthony Fuller, Scott C. Lowe, Daniel G. Kyrollos, Graham W. Taylor, Evan Shelhamer, James R. Green·
arXiv:2609.38278v1 Announce Type: new Abstract: Self-supervised learning (SSL) removes the need for annotations and makes models that are capable across more domains than supervised learning. The autoencoder SSL framework learns by reconstructing its own input after information l…