Researchers have developed a Wasserstein policy gradient (WPG) method for entropy-regularized linear-quadratic (LQ) control problems. This approach leverages the fact that unrestricted LQ control problems have linear-Gaussian optimal policies. The WPG method updates action laws by considering transport in the action space, and a Bellman verification argument confirms that the discounted-occupancy-weighted statewise Wasserstein gradient is tangent to this policy class. Consequently, the WPG method simplifies to a finite-dimensional ordinary differential equation (ODE) for the feedback gain and action covariance, which is proven to be globally well-posed and converges exponentially from any admissible initialization. AI
RANK_REASON Academic paper detailing a new method for a specific control problem. [lever_c_demoted from research: ic=1 ai=0.7]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →