PulseAugur
实时 02:20:00

新研究探索先进的多臂老虎机算法 · 跟踪 8 个来源

该集群包含几篇研究论文,探讨了多臂老虎机算法的进展。主题包括表征对抗性噪声老虎机中的可学性、开发具有有限适应性的上下文轮播老虎机,以及为线性老虎机提出新的探索策略。此外,还提出了关于学习同伴影响概率、处理马尔可夫老虎机中不可观察的状态以及优化条件因果老虎机的研究。 AI

影响 这些在老虎机算法方面的理论进步可能导致各种人工智能应用中更有效和更高效的决策系统。

排序理由 集群包含多篇关于老虎机算法理论方面的学术论文。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 15 个来源。 我们如何撰写摘要 →

新研究探索先进的多臂老虎机算法 · 跟踪 8 个来源

报道来源 [15]

  1. arXiv cs.AI TIER_1 English(EN) · Yunjin Tong ·

    具有双边信息不对称的上下文老虎机监督博弈

    arXiv:2607.00155v1 Announce Type: new Abstract: We study runtime human oversight of an AI agent when private information runs in both directions: the human privately knows her reward function, while the AI privately knows the quality of the action it proposes. This is the kind of…

  2. arXiv cs.LG TIER_1 English(EN) · Arpit Agarwal, Rohan Ghuge, Viswanath Nagarajan, Zhengjia Zhuo ·

    用于单调随机优化的半老虎学习

    arXiv:2312.15427v3 Announce Type: replace Abstract: Stochastic optimization is a widely used approach for optimization under uncertainty, where uncertain input parameters are modeled by random variables. Exact or approximation algorithms have been obtained for several fundamental…

  3. arXiv cs.LG TIER_1 English(EN) · Bin Du, Chang Liu, Dingqi Zhu, Lintao Ye, Dengfeng Sun ·

    带采样违规约束的分布式在线多臂老虎机子模组最大化

    arXiv:2607.00680v1 Announce Type: new Abstract: We study distributed online submodular maximization under partition matroid constraints, in which multiple agents select a limited number of actions from their own subsets sequentially to maximize the cumulative value of a sequence …

  4. arXiv cs.LG TIER_1 English(EN) · Dengfeng Sun ·

    带采样违规约束的分布式在线多臂老虎机子模组最大化

    We study distributed online submodular maximization under partition matroid constraints, in which multiple agents select a limited number of actions from their own subsets sequentially to maximize the cumulative value of a sequence of objective functions. We develop a unified alg…

  5. arXiv cs.LG TIER_1 English(EN) · Tanmay Goyal, Sukruta Prakash Midigeshi, Gaurav Sinha ·

    具有有限适应性的上下文Slate GLM老虎机

    arXiv:2606.31449v1 Announce Type: new Abstract: We investigate the contextual slate bandit problem with generalized linear rewards under limited adaptivity. At each round, the learner is presented with $N$ sets of items, where each item is represented by a $d$-dimensional feature…

  6. arXiv cs.LG TIER_1 English(EN) · Steve Hanneke, Kun Wang ·

    对抗性噪声老虎机学习能力的完整表征

    arXiv:2605.09200v2 Announce Type: replace Abstract: We study adversarial noisy bandits given a known function class $\mathcal{F}$. In each round, the adversary selects a function $f \in \mathcal{F}$, the learner chooses an arm, and then observes a noisy reward determined by the c…

  7. arXiv cs.LG TIER_1 English(EN) · Gaurav Sinha ·

    具有有限适应性的上下文Slate GLM赌徒

    We investigate the contextual slate bandit problem with generalized linear rewards under limited adaptivity. At each round, the learner is presented with $N$ sets of items, where each item is represented by a $d$-dimensional feature vector. The learner then constructs a slate by …

  8. arXiv cs.LG TIER_1 English(EN) · Toshinori Kitamura, Shuai Liu, Csaba Szepesv\'ari ·

    通过绝对扰动实现线性老虎机随机探索

    arXiv:2606.28616v1 Announce Type: new Abstract: In stochastic linear bandits, the canonical Upper Confidence Bound (UCB) algorithm admits a simple frequentist regret analysis but can be computationally demanding, while Thompson Sampling (TS) is computationally attractive yet typi…

  9. arXiv cs.LG TIER_1 English(EN) · Ahmed Sayeed Faruk, Mohammad Shahverdikondori, Elena Zheleva ·

    使用线性上下文老虎机学习同伴影响概率

    arXiv:2510.19119v2 Announce Type: replace Abstract: In networked environments, it is common for users to share recommendations about content, products, services, and possible courses of action. Whether these recommendations are accepted and acted upon is highly context-dependent,…

  10. arXiv cs.LG TIER_1 English(EN) · Thomas Hira, Victor Boone, Urtzi Ayesta, Ina Maria Verloop ·

    具有不可观测状态和受限决策时期的马尔可夫博弈学习

    arXiv:2606.27448v1 Announce Type: new Abstract: This paper studies the problem of regret minimization in Markovian bandits with \emph{non-observable states} and possibly \emph{constrained} decision epochs. The focus is restricted to a ``pure'' regret benchmark, that compares the …

  11. Hugging Face Daily Papers TIER_1 English(EN) ·

    利用多臂老虎机中的相似性

    In many online learning and bandit problems, the actions we consider possess inherent similarities--for instance because they share latent traits, tags, or hierarchical structure. We study online learning with a similarity-structured action set, encoded by a rooted tree whose lea…

  12. arXiv stat.ML TIER_1 English(EN) · Gianmarco Genalti, Marco Mussi, Nicola Gatti, Marcello Restelli, Matteo Castiglioni, Alberto Maria Metelli ·

    利用图触发连接静息与活跃的 Bandit 算法:上升与衰退

    arXiv:2409.05980v2 Announce Type: replace Abstract: Rested and Restless Bandits are two well-known bandit settings that are useful to model real-world sequential decision-making problems in which the expected reward of an arm evolves over time due to the actions we perform or due…

  13. arXiv stat.ML TIER_1 English(EN) · Devdan Dey, Sujoy Bhore, Avishek Ghosh ·

    单索引老虎机问题的最优遗憾值

    arXiv:2605.09454v2 Announce Type: replace Abstract: We study the $\textit{single-index bandit}$ problem, where rewards depend on an unknown one-dimensional projection of high-dimensional contexts through an unknown reward function. This model extends linear and generalized linear…

  14. arXiv stat.ML TIER_1 English(EN) · Lucas L\'evy, Jean-Lou Valeau, Arya Akhavan, Patrick Rebeschini ·

    自洽扰动用于线性老虎机

    arXiv:2510.24187v3 Announce Type: replace Abstract: We consider the adversarial linear bandits setting and present a unified algorithmic framework that bridges Follow-the-Regularized-Leader (FTRL) and Follow-the-Perturbed-Leader (FTPL) methods, extending the known connection betw…

  15. arXiv stat.ML TIER_1 English(EN) · Francisco N. F. Q. Simoes, Itai Feigenbaum, Mehdi Dastani, Thijs van Ommen ·

    条件因果老虎机问题的最小搜索空间

    arXiv:2502.06577v3 Announce Type: replace-cross Abstract: Causal knowledge can be used to support decision-making problems. This has been recognized in the causal bandits literature, where a causal (multi-armed) bandit is characterized by a causal graphical model and a target var…