Researchers have developed a new primal-dual algorithm that achieves optimal $\widetilde{\mathcal{O}}(\sqrt{K})$ regret and cumulative constraint violation for episodic adversarial linear CMDPs with unknown transitions. This new algorithm overcomes the limitations of previous methods, which were bound by $\widetilde{\mathcal{O}}(K^{3/4})$. The approach combines adaptive Follow the Regularized Leader (FTRL), contracted value estimation, and an exponential Lyapunov function, eliminating the need for policy mixing and ensuring policy parameters remain compatible with uniform concentration. AI
RANK_REASON The cluster contains a research paper detailing a new algorithm for a specific type of decision-making process. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- CMDPs
- DagsHub
- Follow-the-Regularized-Leader
- Gotit.pub
- Hugging Face
- IArxiv
- Influence Flower
- Lyapunov function
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →