PulseAugur
实时 19:00:41
English(EN) Data Balancing Strategies: A Systematic Survey of Resampling and Augmentation Methods

数据平衡策略:重采样和增强方法的系统性调查

本文对机器学习中的数据平衡策略进行了系统性回顾,涵盖了重采样和增强技术。它将方法从 SMOTE 等基础方法归类到先进的深度生成模型和集成策略。回顾强调,最佳方法的选择高度依赖于数据集特征和评估指标,并指出了未来的研究方向,例如将基础模型适应偏斜分布。 AI

影响 全面概述了用于改善模型在不平衡数据集上性能的技术,这对于许多实际应用至关重要。

排序理由 这是一篇发表在 arXiv 上的系统性综述论文。

在 arXiv stat.ML 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

数据平衡策略:重采样和增强方法的系统性调查

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
这是一篇发表在 arXiv 上的系统性综述论文。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
120 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv stat.ML TIER_1 English(EN) · Behnam Yousefimehr, Mehdi Ghatee, Javad Fazli, Shervin Ghaffari, Zahra Rafei, Mohammad Amin Seifi, Sajed Tavakoli, Abolfazl Nikahd, Mahdi Razi Gandomani, Alireza Orouji, Ramtin Mahmoudi Kashani, Sarina Heshmati, Negin Sadat Mousavi ·

    数据平衡策略:重采样和增强方法的系统性调查

    arXiv:2505.13518v2 Announce Type: replace Abstract: Imbalanced datasets, where one class significantly outnumbers others, remain a persistent challenge in machine learning, often biasing predictions toward the majority class and degrading classifier performance. This paper provid…