PulseAugur
实时 00:13:30
English(EN) Adversarial Bandit Optimization with Globally Bounded Perturbations to Convex Losses

新研究探讨了具有界限扰动和分布式智能体的对抗性赌博优化

两篇新研究论文探讨了对抗性赌博优化,这是一种机器学习技术,其损失可能不是凸的且不光滑。第一篇论文介绍了一个用于凸损失和 beta-光滑损失的全局预算扰动的框架,并建立了预期遗憾保证。第二篇论文解决了分布式对抗性赌博问题,提出了一种黑盒方法,允许智能体通过流言通信最小化全局平均损失,实现了接近最优的遗憾界限。 AI

影响 这些论文推进了强化学习的理论理解,可能导致在复杂、不确定的环境中做出更鲁棒、更有效的决策算法。

排序理由 两篇在 arXiv 上发表的学术论文,详细介绍了对抗性赌博优化方面的进展。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新研究探讨了具有界限扰动和分布式智能体的对抗性赌博优化

报道来源 [3]

  1. arXiv cs.LG TIER_1 English(EN) · Zhuoyu Cheng, Kohei Hatano, Eiji Takimoto ·

    具有全局有界扰动凸损失的对抗性赌博优化

    arXiv:2606.19891v1 Announce Type: new Abstract: We study adversarial bandit optimization in which the loss functions may be non-convex and non-smooth. In each round, the learner selects an action and observes only the loss incurred at that action. The loss consists of an underlyi…

  2. arXiv cs.LG TIER_1 English(EN) · Eiji Takimoto ·

    具有全局有界凸损失扰动的对抗性赌博优化

    We study adversarial bandit optimization in which the loss functions may be non-convex and non-smooth. In each round, the learner selects an action and observes only the loss incurred at that action. The loss consists of an underlying convex and $β$-smooth component and an advers…

  3. arXiv cs.LG TIER_1 English(EN) · Hao Qiu, Mengxiao Zhang, Nicol\`o Cesa-Bianchi ·

    分布式对抗性赌博机的近乎最优遗憾:一种黑盒方法

    arXiv:2602.06404v2 Announce Type: replace Abstract: We study distributed adversarial bandits, where $N$ agents cooperate to minimize the global average loss while observing only their own local losses. We show that the minimax regret for this problem is $\tilde{\Theta}(\sqrt{(\rh…