This paper addresses the overestimation bias in Q-learning, particularly within large discrete action spaces. The authors propose an "action intersection" strategy that semi-decouples Q-value estimation by allowing shared trajectory data between two Q-functions. This method offers fine-grained control over estimation bias, enabling it to range from underestimation to overestimation by adjusting the data sharing fraction. Experiments in both tabular and deep reinforcement learning settings demonstrate significant improvements over existing state-of-the-art baselines. AI
IMPACT Introduces a novel technique to improve Q-learning performance in complex environments, potentially enhancing agent capabilities in large-scale decision-making scenarios.
RANK_REASON The cluster contains an academic paper detailing a new method for reinforcement learning. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →