PlanBench
PulseAugur coverage of PlanBench — every cluster mentioning PlanBench across labs, papers, and developer communities, ranked by signal.
-
AI learns faithful symbolic planning using multi-role reinforcement learning
Researchers have developed a novel multi-role reinforcement learning framework to improve the faithfulness of symbolic planning by large language models. This framework utilizes a single language model to act as an Acto…
-
New AI training method uses governance records for improved workflow repair
Researchers have developed a method called Verifier-Selected Self-Training (VSST) that uses governance records from machine-verifiable workflows to supervise AI models. These records, which include task contracts, model…
-
New CoSPlan benchmark challenges vision-language models in visual planning tasks
Researchers have introduced CoSPlan, a new benchmark designed to evaluate the sequential planning capabilities of vision-language models (VLMs) in visual domains. Unlike text-based planning, CoSPlan requires models to e…