PulseAugur
中
实时 13:25:36

新的线性规划方法无需标签即可修剪数据集

研究人员开发了一种新颖的数据集修剪方法,该方法使用线性规划来选择代表性数据子集,而无需标签或模型训练。该方法将无偏子集选择重新表述为方差最小化问题,并从嵌入空间属性推导出几何标准。在 CIFAR-10、MNIST 和 CelebA 基准上的实验表明,该方法在各种数据预算下,在测试准确性方面可媲美甚至超越均匀采样,并且优于现有的几何方法,尤其是在较小的预算下。该框架通过增强小批量多样性,在减少随机梯度方差方面也具有优势。 AI

影响 通过在无需标签或大量计算的情况下进行更有效的数据子集选择,该方法可以提高训练效率和模型性能。

排序理由 该集群包含一篇详细介绍数据集修剪新方法的论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的线性规划方法无需标签即可修剪数据集

本文如何被排名

Signal score
7 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇详细介绍数据集修剪新方法的论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Rodrigo Schuller, Francisco Ganacim ·

    从第一性原理进行数据集剪枝:一种无标签的线性规划方法

    arXiv:2610.10347v1 Announce Type: cross Abstract: Dataset pruning reduces a large training set to a representative subset while preserving model performance. Existing geometry-based methods typically assume that nearby points in embedding space share similar properties. Rather th…