Researchers have developed a novel approach to training AI agents for complex games by utilizing a continuous Dyna loop. This method trains a policy solely within a learned world model of a ten-hero multiplayer online battle arena (MOBA) game, with the real game serving only to provide data for world model updates and policy evaluation. The trained policy achieved a 70.2% win rate in the actual game, a significant improvement from zero wins when trained only in imagination. Key findings indicate that model exploitation is difficult to assess from within the simulated environment, and a world model that is accurate on its training data may be inaccurate in the context of the policy's own games, necessitating the Dyna loop for repair. AI
IMPACT Demonstrates a novel method for training AI agents in complex environments, potentially accelerating AI development in strategy games and beyond.
RANK_REASON Academic paper detailing a new AI training methodology. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →