PulseAugur
实时 10:29:35
English(EN) Reoptimization Algorithms for Contextual Bandits with Knapsack Constraints

新算法改进了具有背包约束的上下文老虎机问题的遗憾界限

研究人员为具有背包约束的上下文老虎机问题开发了新算法,该问题涉及在资源受限和奖励不确定的情况下将客户分配给产品。所提出的算法扩展了置信上限(UCB)家族并利用了再优化技术。这些方法实现了 O((ln T)^3 / T) 的平均遗憾,相比于类似动态定价问题的现有界限有了显著改进。 AI

影响 为资源受限环境中的不确定性决策引入了改进的理论界限。

排序理由 该集群包含一篇在 arXiv 上发表的研究论文,详细介绍了针对特定机器学习问题的新算法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新算法改进了具有背包约束的上下文老虎机问题的遗憾界限

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Zhen Xu ·

    具有背包约束的上下文老虎机问题的重优化算法

    arXiv:2608.11383v1 Announce Type: new Abstract: We study new algorithms for Contextual Bandits with Knapsack. In these problems, there are finitely many types of customers, products, and resources. Each product is made from a fixed combination of resources, and resources have fin…