PulseAugur
EN
LIVE 11:24:20

New sampling method improves multi-label data balance

A new sampling method based on the multivariate Bernoulli distribution has been proposed for multi-label datasets where labels vary significantly in frequency. This novel algorithm accounts for label dependencies by estimating distribution parameters from observed label frequencies and calculating weights for each label combination. Applied to a sample of research articles labeled with 64 biomedical topic categories, the method successfully created a more balanced sub-sample, improving the representation of less common categories while preserving the overall frequency order. AI

RANK_REASON Academic paper detailing a new statistical method for data sampling. [lever_c_demoted from research: ic=1 ai=0.4]

Read on arXiv stat.ML →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New sampling method improves multi-label data balance

COVERAGE [1]

  1. arXiv stat.ML TIER_1 English(EN) · Simon Chung, Colby J. Vorland, Donna L. Maney, Andrew W. Brown ·

    A Multivariate Bernoulli-Based Sampling Method for Multi-Label Data with Application to Meta-Research

    arXiv:2512.08371v4 Announce Type: replace-cross Abstract: Datasets may contain observations with multiple labels. If the labels are not mutually exclusive, and if the labels vary greatly in frequency, obtaining a sample that includes sufficient observations with scarcer labels to…