importance sampling
PulseAugur coverage of importance sampling — every cluster mentioning importance sampling across labs, papers, and developer communities, ranked by signal.
-
New framework improves AI-driven image analysis with statistical rigor
Researchers have developed a new framework for multi-target estimation in large image collections, addressing the bias introduced by computer vision models. This approach combines model predictions with limited human an…
-
New Kernel-WIS estimator improves off-policy evaluation for contextual bandits
Researchers have introduced Kernel-WIS, a new estimator for off-policy evaluation in contextual bandits. This method utilizes offline data and is designed to be asymptotically consistent. Kernel-WIS aims to outperform e…
-
New UP objective enhances LLM reasoning by balancing exploration and stability
Researchers have introduced Unbounded Positive Asymmetric Optimization (UP), a novel objective function designed to improve reinforcement learning (RL) for large language models (LLMs). UP addresses the exploration-stab…
-
New Selective Importance Sampling method improves LLM alignment
Researchers have introduced Selective Importance Sampling (SIS), a novel plug-in method designed to enhance the alignment of large language models (LLMs) during reinforcement learning post-training. This approach addres…
-
Paper reviews optimality in Monte Carlo importance sampling
This paper provides a comprehensive review of optimality within importance sampling techniques, a critical component for the performance of Monte Carlo sampling methods. It explores various frameworks for designing adap…
-
New research advances off-policy evaluation techniques for ML
Two new research papers explore advanced techniques for off-policy evaluation (OPE) in machine learning, a critical process for assessing the performance of new policies using existing data. The first paper introduces "…
-
New DR-IS method boosts ML robustness against adversarial label corruption
Researchers have developed a new sub-sampling method called Disagreement-Regularized Importance Sampling (DR-IS) to improve robustness against adversarial label corruption in machine learning. This method leverages the …