PulseAugur
EN
LIVE 20:56:44

New AI agent research focuses on training, reliability, and multi-agent coordination

Several research efforts are focusing on improving AI agents' capabilities and reliability. Hugging Face has introduced AutoSynthData for generating training data for enterprise agents and Holo4, a series of agentic models designed for diverse interfaces. Apple's SCLATE provides a substrate for continual-learning agent training and evaluation, while xAI's Team Bots allow for shared context and learning among AI coworkers. Additionally, new research papers explore agent skill evolution (SkillSpec), resource assignment (TRACE), reliability in multi-agent teams (Worse Together), and improving agent performance in complex environments like incident response (Incident-Arena) and user correction arbitration (GAVA). AI

IMPACT Advances in agent training, evaluation, and multi-agent coordination could accelerate enterprise adoption and improve AI reliability in complex tasks.

RANK_REASON Multiple research papers and platform announcements related to AI agent development, training, and evaluation.

Read on Microsoft Research →

AI-generated summary · Google Gemini · from 420 sources. How we write summaries →

New AI agent research focuses on training, reliability, and multi-agent coordination

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Multiple research papers and platform announcements related to AI agent development, training, and evaluation.
Source corroboration
420 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
product, paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
19 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+164 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [420]

  1. Microsoft Research TIER_1 English(EN) · Zhiyuan He, Yuqing Yang ·

    Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

    <p>Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them.</p> <p>The …

  2. Apple Machine Learning Research TIER_1 English(EN) ·

    RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation

    Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based sign…

  3. Hugging Face Blog TIER_1 English(EN) ·

    AutoSynthData: Generating Training Data for Enterprise Agents

  4. Apple Machine Learning Research TIER_1 English(EN) ·

    SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation

    Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, crons, and memory consolidation. Yet existing benchma…

  5. Hugging Face Blog TIER_1 English(EN) ·

    Holo4: powering generalist computer-use agents

  6. xAI news TIER_1 English(EN) ·

    Team Bots: AI coworkers that learn from your team

    Give a Grok Bot the files, apps, and expertise it needs, then share it so your whole team can work from the same context.

  7. arXiv cs.LG TIER_1 English(EN) · Sangmin Lee, Youngju Na, Chanmi Lee, Sung-eui Yoon ·

    Multi-Agent Coordination via Support-Preserving Distillation

    arXiv:2610.10087v1 Announce Type: new Abstract: Offline MARL increasingly relies on generative policies to model multimodal joint behavior, typically by distilling a centralized teacher into decentralized one-step actors under the CTDE. We identify a failure mode at the teacher t…

  8. arXiv cs.LG TIER_1 English(EN) · Tianruo Rose Xu, Jiawei Ren, Yichi Yang, Zhaoxu Zheng, Lianhui Qin ·

    RT-Safe: Benchmarking Agent Safety in Real-Time Embodied Environment

    arXiv:2610.09294v1 Announce Type: cross Abstract: Rapid progress in AI agents has brought growing attention to agent safety, with extensive evaluation focused on digital environments. As agents move into the physical world, embodied safety becomes increasingly important: failures…

  9. arXiv cs.LG TIER_1 English(EN) · Yu Cheng, Dehai Zhao, Zhongxin Liu, Qing Huang, Zhenchang Xing, Xiaoxue Ren ·

    An Empirical Study of Agent Skills' Downstream Utility

    arXiv:2610.08875v1 Announce Type: cross Abstract: Agent Skills package procedural guidance and resources for reuse, but a relevant Skill does not necessarily improve task performance. Existing studies characterize Skill content and evaluate downstream performance, yet provide lim…

  10. arXiv cs.LG TIER_1 Deutsch(DE) · Prakhar Ganesh, Kyra Wilson, Luca Zappella, Barry-John Theobald, Nicholas Apostoloff, Lucas Monteiro Paes, Nivedha Sivakumar ·

    Homogenization in Multi-Agent Systems

    arXiv:2610.09824v1 Announce Type: new Abstract: Multi-agent systems (MAS) leverage interactions between agents to perform complex tasks. Despite their success, we show that these interactions can also lead to homogenization, i.e., agents converging to similar behaviors. Homogeniz…

  11. arXiv cs.LG TIER_1 English(EN) · Yuanzhe Li, Pengxin Wang, Yuxin Ren, Jianing Deng, Jingtong Hu, Song Wang, Jingdi Chen, Huanrui Yang ·

    TAP: Efficient Long-Horizon Agent Pruning via Trajectory-Anchored Recovery

    arXiv:2610.09074v1 Announce Type: new Abstract: Emerging long-horizon agentic tasks require repeated model calls, worsening the inference cost of already-costly language models. While narrow agentic tasks suggest potential for aggressive model pruning without performance drop, em…

  12. arXiv cs.CL TIER_1 English(EN) · Igor Slinko, Yaroslav Golubev, Sergey Titov ·

    Coding-Agent Benchmarks Should Match Their Users' Task Flows

    arXiv:2610.09633v1 Announce Type: cross Abstract: The evaluation of coding agents generally strives to be as realistic as possible. In our study, we collect 4,782 agent sessions of real software engineers in JetBrains IDEs, which we call Production Sessions. Since our subject is …

  13. arXiv cs.CL TIER_1 English(EN) · Xinglin Wang, Zishen Liu, Tong Zheng, Shaoxiong Feng, Peiwen Yuan, Yiwei Li, Jiayi Shi, Yueqi Zhang, Chuyi Tan, Ji Zhang, Boyuan Pan, Kan Li ·

    From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery

    arXiv:2610.09684v1 Announce Type: new Abstract: Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dime…

  14. arXiv cs.CL TIER_1 English(EN) · Giordano De Marzo, Andres L. Marin, David Garcia ·

    Collective Behavior of AI Agents: the Case of Moltbook

    arXiv:2602.09270v2 Announce Type: replace-cross Abstract: We present a large scale data analysis of Moltbook, a Reddit-style social media platform exclusively populated by AI agents. Analyzing over 4 million posts and 19 million comments from approximately 185,000 active agents, …

  15. Hugging Face Daily Papers TIER_1 English(EN) ·

    OOM-RL II: Reality Is an Oracle, Not a Debugger Provenance-Constrained Diagnosis in Continually Evolving Agent-Engineered Systems

    Reality may establish that an outcome occurred without identifying which evolving procedure produced it or why. This distinction matters in production ML systems whose code, configuration, and artifacts change while external feedback accumulates. We examine it in a human-directed…

  16. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Sung-eui Yoon ·

    Multi-Agent Coordination via Support-Preserving Distillation

    Offline MARL increasingly relies on generative policies to model multimodal joint behavior, typically by distilling a centralized teacher into decentralized one-step actors under the CTDE. We identify a failure mode at the teacher training stage: standard flow-based teachers pair…

  17. arXiv cs.AI TIER_1 English(EN) · Wenxuan Wang, Zekai Liu, Weinan Zhang, Yu Cheng, Yang Yang ·

    Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation

    arXiv:2610.07250v1 Announce Type: new Abstract: Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continu…

  18. arXiv cs.AI TIER_1 English(EN) · Bowen Ye, Yongchao Xu, Junkai Ma, Xiang Yin, Wenzhao Li ·

    Principles that Guide, Actions that Inform: Agent Evolution via Knowledge Abstraction

    arXiv:2610.06964v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated strong capabilities in interactive environments, yet their ability to continually evolve from experience remains limited. Although fine-tuning enables adaptation, its dependence on…

  19. arXiv cs.AI TIER_1 English(EN) · Zhe Yu, Zixuan Wang, Peidong Wang, Hehai Lin, Ruochen Zhao, Chengwei Qin ·

    Beyond Corrected Memory: Execution Consistency in Multi-Agent Systems

    arXiv:2610.08101v1 Announce Type: new Abstract: Shared memory coordinates agents' actions, but correct records do not establish that those actions satisfy task requirements. Memory governance and failure diagnosis regulate or inspect recorded information; they do not by themselve…

  20. arXiv cs.AI TIER_1 English(EN) · Qianhan Feng, Zhongzhen Huang, Yakun Zhu, Xiaofan Zhang, Qi Dou ·

    Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents

    arXiv:2610.07979v1 Announce Type: new Abstract: As agents continuously improve by generating and revising Skills, the process that discovers and refines those Skills becomes a learnable object in its own right. Task-Skills directly act on task execution, whereas Meta-Skills gover…

  21. arXiv cs.AI TIER_1 English(EN) · Yeji Park, Jaeyun Shim, Taesik Gong ·

    Can Agents Work for Everyone? Cross-User Reliability for Mobile GUI Agents in Personalized User Interfaces

    arXiv:2610.07972v1 Announce Type: new Abstract: Mobile GUI agents increasingly operate on interfaces influenced by users' histories and preferences, but their reliability across different users remains underexplored. We introduce PAIR (Personalized Application-state Instantiation…

  22. arXiv cs.AI TIER_1 English(EN) · Qi Cheng, Shengyu Chen, Wei Cheng, Zhengzhang Chen, Xiaowei Jia, Haoyu Wang, Haifeng Chen ·

    WorkflowOps: Learning Agent Collaboration Priors for Multi-Agent Workflow Orchestration

    arXiv:2610.07860v1 Announce Type: new Abstract: Multi-agent systems are increasingly deployed for complex knowledge work, yet their orchestration layers remain largely memoryless: each new task is decomposed, assigned, and executed from scratch with no benefit from prior successf…

  23. arXiv cs.AI TIER_1 English(EN) · Qi Cheng, Shengyu Chen, Wei Cheng, Yiqun Xie, Xiaowei Jia, Haoyu Wang, Haifeng Chen ·

    OOPMAS: Object-Oriented Multi-Agent Systems for Query-Level Workflow Generation

    arXiv:2610.07787v1 Announce Type: new Abstract: Multi-agent systems (MAS) powered by large language models have shown strong performance across code generation, mathematical reasoning, and question answering. However, existing methods for automating MAS design mostly operate at t…

  24. arXiv cs.AI TIER_1 English(EN) · Qi Cheng, Rongchao Dong, Shengyu Chen, Licheng Liu, Dan Lu, Zhengzhang Chen, Wei Cheng, Yiqun Xie, Haifeng Chen, Xiaowei Jia, Haoyu Wang ·

    ST-Bench: A Spatial-Temporal Benchmark for Multi-Agent System Generation on Scientific Research Tasks

    arXiv:2610.07763v1 Announce Type: new Abstract: The rapid progress of LLM-based multi-agent systems (MAS) has shown that they largely outperform single agents on coding, math, and QA tasks, where executable tests provide a binary success signal. Whether this advantage transfers t…

  25. arXiv cs.AI TIER_1 English(EN) · Fouad Bousetouane ·

    EIO-Agents: The Missing Semantic Layer for AI Agent Evaluation

    arXiv:2610.07675v1 Announce Type: new Abstract: AI agents are entering production in increasingly consequential environments without a shared semantic standard for what their evaluations actually mean. Scores, traces, judge outputs, and multi juror findings are increasingly used …

  26. arXiv cs.AI TIER_1 English(EN) · Jianglin Qiao, Siyi Hu, Thien Hoang Nguyen, Zehong Cao, Salah Sukkarieh ·

    Cooperating with Future Collaborators: Multi-Agent RL under Staggered Participation

    arXiv:2610.07578v1 Announce Type: new Abstract: In cooperative Multi-Agent Reinforcement Learning (MARL), agents are often trained under concurrent participation, while in many tasks some agents act earlier and leave task-relevant information that becomes useful to agents partici…

  27. arXiv cs.AI TIER_1 English(EN) · Xinle Wu, Yao Lu ·

    Decoupled Multi-Agent Orchestration

    arXiv:2610.07556v1 Announce Type: new Abstract: Learned orchestration can automatically construct effective language-model multi-agent systems, but existing approaches couple planning to fixed worker pools and train decomposition and collaboration from the same terminal outcome, …

  28. arXiv cs.AI TIER_1 English(EN) · Mohammadreza Sediqin, Shivali Dalmia, Srinivasa Karthikeya Reddy Kovvuri, Abhishek Mukherji ·

    A Trust Layer for Agent Evaluation

    arXiv:2610.07274v1 Announce Type: new Abstract: Deterministic benchmark scores show that an agent received credit, but not whether that credit was earned, reported honestly, or would hold on a second run. We introduce a Trust Layer for Agent Evaluation, an additive post-hoc frame…

  29. arXiv cs.AI TIER_1 English(EN) · Sumanyu Muku ·

    Verifying Coordination in Parallel Coding Agents: NP-Bench and a Scheduling Planner

    arXiv:2610.07261v1 Announce Type: new Abstract: A team of coding agents can look fine agent by agent yet fail as a team: each passes its own tests while the merged result is broken, and single-agent evaluation never catches it. As teams run several LLM coding agents in parallel o…

  30. arXiv cs.AI TIER_1 English(EN) · Sumanyu Muku ·

    MemMux: Runtime Verification and Honest Resource Attribution for Fleets of Parallel Coding Agents

    arXiv:2610.07257v1 Announce Type: new Abstract: Developers increasingly run a fleet of coding agents side by side on one workstation. The tools they reach for, terminal multiplexers like tmux and a new generation of agent managers, were built to arrange windows, not to govern mem…

  31. arXiv cs.AI TIER_1 English(EN) · Sen Zhao, Jia Tang, Ruiqi Kong, Zuyu Zhang, Lifeng Shen, Ding Zou, Xinyu He, Xu Zhang, Junwei Han ·

    Topology-Consistent Task Planning over Cellular Workflow Complexes for LLM-based Agents

    arXiv:2610.07004v1 Announce Type: new Abstract: Task planning for LLM agents requires workflows that satisfy both user intent and complex sub-task dependencies. While existing planners work well for sequential or directed acyclic graph (DAG)-like structures, they struggle with wo…

  32. arXiv cs.LG TIER_1 English(EN) · Hyesung Jeon, Hyeongju Ha, Seoyoung Lee, Beomseok Kang, Jae-Joon Kim ·

    KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems

    arXiv:2609.34060v2 Announce Type: replace Abstract: Prompt-specialized multi-agent systems enable multiple agents to share a model while performing complementary roles to solve complex tasks. However, agent-specific prefixes change the KV cache generated for the same shared conte…

  33. arXiv cs.LG TIER_1 English(EN) · Jia Liufu, Bin Hu, Linglin Jing, Terry Kong, Yuki Huang, Ashwath Aithal, Wenming Yang, Jun Yang ·

    FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents

    arXiv:2610.07898v1 Announce Type: new Abstract: Repository-level software engineering (SWE) is a challenging long-horizon setting: agents must reason over extended interactions, use tools, and adapt to stateful environments. Recent work trains SWE agents with reinforcement learni…

  34. arXiv cs.LG TIER_1 English(EN) · Lan Shi, Daigo Shishika, Xuan Wang ·

    Adapting to Changes in Agent Behavior via Finite-Depth Policy Sensitivity

    arXiv:2610.07475v1 Announce Type: new Abstract: Adapting a reinforcement learning policy to changes in another agent's behavior typically requires a large amount of new interaction data. Policy sensitivity provides a first-order prediction of how a locally optimal policy changes …

  35. arXiv cs.LG TIER_1 English(EN) · Ziyang Cai, Christos Ziakas, Vasilis Kontonis, Tim Pearce, Siddhartha Sen, Akshay Krishnamurthy, Shivam Garg, Dimitris Papailiopoulos ·

    Fork-and-Flush: Escaping Idea Basins in Autoresearch Agents

    arXiv:2610.07447v1 Announce Type: new Abstract: Autoresearch agents tackle open-ended problems by repeatedly proposing candidate solutions, evaluating them, and using feedback to guide subsequent experiments. We show that independent runs of the same agent on the same task often …

  36. arXiv cs.AI TIER_1 English(EN) · Haoyue Yang, Jingyao Li, Zhengfan Wu, Jing Liu, Xuanle Zhao, Kang Liu ·

    GAMEGO: Training Game-Dev Agents with Synthetic Trajectories Anchored in Real-World Assets

    arXiv:2610.06910v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have demonstrated remarkable capabilities in web front-end execution, with browser-based game generation emerging as a particularly prominent frontier. While previous efforts frequentl…

  37. arXiv cs.CL TIER_1 English(EN) · Shivani Kumar, Adarsh Bharathwaj, David Jurgens ·

    Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows

    arXiv:2604.20658v2 Announce Type: replace Abstract: Multi-agent systems built from teams of large language models (LLMs) are increasingly deployed for collaborative scientific reasoning and problem-solving. These systems require agents to coordinate under shared constraints, such…

  38. arXiv cs.CL TIER_1 English(EN) · Naoki Wake, Justin Wagle ·

    SharedKV-BT: Node-Local Typed Decisions for Behavior-Tree Agents

    arXiv:2610.07327v1 Announce Type: cross Abstract: Agent tasks require sequences of interdependent decisions. Autoregressive models support more flexible decision interfaces than conventional classifiers but incur the latency of token-by-token generation. Recent shared-prefix meth…

  39. arXiv cs.CL TIER_1 English(EN) · Lasse B. Strand, Robert Jakob, Kevin O'Sullivan, Markus Kreft ·

    Agentic AutoRAG: RAG Pipeline Optimization through Reasoning-Driven Agents

    arXiv:2610.08452v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacti…

  40. arXiv cs.AI TIER_1 Norsk(NO) · Justin Chih-Yao Chen, Elias Stengel-Eskin, Yan Chen, Pol Llado, Scott Counts, Mohit Bansal, Benjamin Van Durme, Harsh Jhamtani, Gaurav Verma ·

    TeleTune: Evolving Agent Skills From Offline Telemetry

    arXiv:2610.05437v2 Announce Type: replace Abstract: Computer-use agents need to capture procedural knowledge of how people use software. User telemetry offers a scalable source of this knowledge. However, learning reusable skills from these logs requires addressing three challeng…

  41. arXiv cs.AI TIER_1 English(EN) · Yingying Liu, Junzhou Fang, Chenxiong Qian ·

    When Agent Context Goes Stale: Incoherence in Volatile Agent Context

    arXiv:2610.05281v2 Announce Type: replace Abstract: Modern agents increasingly ground their reasoning in observations returned by tools, such as file contents read from a workspace. However, the data sources underlying these observations may later be modified by users, other agen…

  42. arXiv cs.AI TIER_1 English(EN) · Xiaoyu Chen, Lai Wei, Jin Wang, Xiangyu Zou, Ruochen Fan, Enze Luo, Mingzhe Yao, Jiahui Zhu, Yuhua Wen, Linghe Kong, Weiran Huang ·

    SWE-Game: Can Coding Agents Build the Games We Want?

    arXiv:2609.33678v2 Announce Type: replace Abstract: We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design docu…

  43. arXiv cs.AI TIER_1 English(EN) · Jiaxuan Dai, Tianyi Huang ·

    TwinCheck: Evidence-Grounded Negative-Twin Verification for Stateful Tool Agents

    arXiv:2609.26911v2 Announce Type: replace Abstract: A single locally plausible tool call can derail an otherwise successful agent trajectory. Suspicion alone does not justify intervention, because the replacement itself can introduce the very failure verification is meant to prev…

  44. arXiv cs.AI TIER_1 English(EN) · Sarim Hashmi, Mukul Ranjan, Kshitij Mishra, Mikhail Kuznetsov, Praneeth Vepakomma, Nils Lukas ·

    AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

    arXiv:2610.08773v1 Announce Type: cross Abstract: Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the …

  45. arXiv cs.AI TIER_1 English(EN) · Zihan Zhou, Xinzhe Hu, Hanxu Yang, Liangjian Wen, Zhao Kang ·

    Token-Efficient Multi-Agent Collaboration via System One-Guided Computational Division of Labor

    arXiv:2610.08155v1 Announce Type: cross Abstract: Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by enabling collaborative problem solving among specialized agents. However, existing …

  46. arXiv cs.AI TIER_1 English(EN) · Kavienan Jegatheesan, Gayathri Lihinikaduarachchi ·

    When Tools Lie: Reliability of Mathematical Agents Under Corrupted Tool Feedback

    arXiv:2610.08097v1 Announce Type: cross Abstract: Mathematical problem solving often requires deterministic computational steps that agents delegate to tools and implicitly trust. Yet tools can fail silently, returning plausible but incorrect results. How well can agents detect a…

  47. arXiv cs.AI TIER_1 English(EN) · Hongzhan Lin, Shidong Cao, Ziyang Luo, Wenhao Chai, Mong-Li Lee, Wynne Hsu ·

    From Evidence to Action: How Tool-Using Agents Fail

    arXiv:2610.07753v1 Announce Type: cross Abstract: Tool-using agents make consequential changes to external state, yet correct outcomes do not guarantee that their actions were supported by evidence established beforehand. We study where this evidence-to-action chain breaks as age…

  48. arXiv cs.AI TIER_1 English(EN) · Hang He, Li Wang, Hao Chen, Yuchen Shao, Yuling Shi, Lisheng Wang, Peiyang Liu, Goose Lin, Zaiyuan Wang, Haiying Sun, Ting Su, Chengcheng Wan ·

    CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?

    arXiv:2610.07557v1 Announce Type: cross Abstract: Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing co…

  49. arXiv cs.AI TIER_1 English(EN) · Celal Ziftci, Spencer Greene, Ray Liu, Livio Dalloro, Lorenzo Dini ·

    Catching Developers in the Flow: Low-Latency Agentic Program Repair at Google Scale

    arXiv:2610.07289v1 Announce Type: cross Abstract: Manual repair of program failures is time-consuming and disruptive for software developers, particularly during the pre-submit phase where test failures occur within continuous integration systems. While Automated Program Repair h…

  50. arXiv cs.AI TIER_1 English(EN) · Tao Long, Lydia B. Chilton ·

    SPEAR: Five Principles for Interactive Human-Agent Alignment

    arXiv:2610.07204v1 Announce Type: cross Abstract: Recent AI alignment work often frames alignment as a pre-deployment optimization problem: collect human feedback, learn preferences or principles, finetune the model, and deploy an aligned system. This framing has produced major p…

  51. arXiv cs.AI TIER_1 English(EN) · Siru Jiang, Yongzhe Lyu, Shuo Lu, Yubin Wang, Yuxiang Zhang, Yue Liao, Bin Wang, Jian Liang, Tieniu Tan ·

    WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?

    arXiv:2610.08720v1 Announce Type: new Abstract: LLM-based agents are increasingly advancing scientific and engineering problem solving, with physics simulation emerging as a challenging yet practical testbed for reproducing complex physical phenomena with application in embodied …

  52. arXiv cs.AI TIER_1 English(EN) · Mingda Zhang, Wenjin Liu, Tiesunlong Shen, Zikai Xiao, Zhenghong Lin, Qing Xu, Erik Cambria, Xiaoying Tang, Haoran Luo ·

    ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences

    arXiv:2610.08691v1 Announce Type: new Abstract: Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both th…

  53. arXiv cs.AI TIER_1 English(EN) · Hanjun Luo, Xiucheng Zhang, Zhuoning Xu, Zhimu Huang, Yingbin Jin, Xinfeng Li, Hanan Salam ·

    ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding

    arXiv:2610.08662v1 Announce Type: new Abstract: As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a …

  54. arXiv cs.AI TIER_1 English(EN) · Yunbo Long, Guangya Hao, Yuhan Liu, Yiting Duan, Longyan Tan, Yunchen Long, Hao Wu ·

    How Much Evidence Should a Coding Agent's Self-Correction Carry? Adaptive Dirichlet Evidence for Self-Distillation

    arXiv:2610.08514v1 Announce Type: new Abstract: Execution feedback lets coding agents revise programs and learn from their own corrections. A correction's learning weight should reflect both the transitions supported by its executions and the amount of evidence behind that suppor…

  55. arXiv cs.AI TIER_1 English(EN) · Hyun Jung Lee, Jungtaek Kim, Jongwon Jeong, Tae-Eui Kam, Donghyun Kim, Yong Jae Lee ·

    EMHO: EMbodied Agent Harness Optimization via Experience Traces

    arXiv:2610.08432v1 Announce Type: new Abstract: Improving embodied agents often focuses on optimizing the underlying model through training, while the surrounding agent harness that controls planning, context, and tool use is typically engineered. We ask whether this harness can …

  56. arXiv cs.AI TIER_1 English(EN) · Haotian Chen, Shuaicheng Niu, Haocong Rao, Kaisong Song, Jun Lin, Lizhen Cui, Zhiqi Shen, Yonghui Xu ·

    Test-Time Agent Evolution for Long-Horizon Legal Reasoning

    arXiv:2610.08138v1 Announce Type: new Abstract: Legal intelligence aims to support reliable decision-making across long-horizon legal processes involving evolving case states and multiple roles. However, real-world legal deployment exhibits substantial case heterogeneity in facts…

  57. arXiv cs.AI TIER_1 English(EN) · Langxi Huang, Pingping Zhang, Lanyun Zhu, Chunyang Jiang, Jiawei Shao, Haocheng Yuan, Peilin Chen ·

    ChartBmkAgent: Harness-Governed Multi-Agent Construction of Chart QA Benchmarks from Sparse Error-Taxonomy Specifications

    arXiv:2610.08106v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) advance rapidly, while conventional benchmark development lags behind, delaying investigation of newly observed capability gaps. Such investigation requires an expressive task format and an o…

  58. Hugging Face Daily Papers TIER_1 English(EN) ·

    RoboQuest: Generalist Physical Agents that Search, Inspect and Test

    Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-relevant information through interaction when it is ab…

  59. Hugging Face Daily Papers TIER_1 English(EN) ·

    SWE-Game: Can Coding Agents Build the Games We Want?

    We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fau…

  60. Hugging Face Daily Papers TIER_1 English(EN) ·

    We Query, Therefore We Compute: On Oracle Computation beyond the Machine, with an Application to Agents

    Agentic systems use large language models (LLMs) to carry out concrete tasks. Prior work often borrows abstractions such as scheduling, caching or isolation piecemeal from operating systems, so the mechanisms it builds share little common ground, and the shared view of the two fo…

  61. Hugging Face Daily Papers TIER_1 English(EN) ·

    From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery

    Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--…

  62. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Veronique Ziegler ·

    When the Governor Becomes the Disturbance: Control-Generated Disturbance and Cost-Aware Backoff in Governed Tool-Using Agents

    Supervisory governors can interfere with the tool-using agents they regulate. We study this possibility in a controlled file-recovery environment where increases in regulatory intensity trigger experimentally imposed tool failures. A cost-blind governor can turn these failures in…

  63. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Hariganesh Tangirala ·

    Learning to Report Unsafe Tasks in a Multi-Agent Game

    When agents share a reward for completed tasks, reporting unsafe work can reduce the reporter's reward by stopping a task. Audits can make reporting optimal without ensuring that further training teaches a silent team to report. We study this learning problem in a game where any …

  64. Hugging Face Daily Papers TIER_1 English(EN) ·

    AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

    Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task r…

  65. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Markus Kreft ·

    Agentic AutoRAG: RAG Pipeline Optimization through Reasoning-Driven Agents

    Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacting choices, from chunking and embedding model to…

  66. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Zhao Kang ·

    Token-Efficient Multi-Agent Collaboration via System One-Guided Computational Division of Labor

    Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by enabling collaborative problem solving among specialized agents. However, existing MAS frameworks tightly couple task reasoning with …

  67. Hugging Face Daily Papers TIER_1 English(EN) ·

    Beyond Corrected Memory: Execution Consistency in Multi-Agent Systems

    Shared memory coordinates agents' actions, but correct records do not establish that those actions satisfy task requirements. Memory governance and failure diagnosis regulate or inspect recorded information; they do not by themselves establish whether it is sufficient to judge ta…

  68. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Dongha Lee ·

    From Delivery to Stateful Exploration: Rethinking the Index for Agentic Search

    Recent advances in agentic search have given large language model (LLM) agents finer control over corpus exploration. However, search interfaces often return matching passages even when feedback about the candidate set would suffice for the next decision, coupling candidate refin…

  69. Hugging Face Daily Papers TIER_1 English(EN) ·

    OOPMAS: Object-Oriented Multi-Agent Systems for Query-Level Workflow Generation

    Multi-agent systems (MAS) powered by large language models have shown strong performance across code generation, mathematical reasoning, and question answering. However, existing methods for automating MAS design mostly operate at the task level, producing a single fixed workflow…

  70. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Fouad Bousetouane ·

    EIO-Agents: The Missing Semantic Layer for AI Agent Evaluation

    AI agents are entering production in increasingly consequential environments without a shared semantic standard for what their evaluations actually mean. Scores, traces, judge outputs, and multi juror findings are increasingly used to justify readiness and release decisions, yet …

  71. Hugging Face Daily Papers TIER_1 English(EN) ·

    CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?

    Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch…

  72. Hugging Face Daily Papers TIER_1 English(EN) ·

    From Evidence to Action: How Tool-Using Agents Fail

    Tool-using agents make consequential changes to external state, yet correct outcomes do not guarantee that their actions were supported by evidence established beforehand. We study where this evidence-to-action chain breaks as agents move from deciding whether to act to executing…

  73. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Lydia B. Chilton ·

    SPEAR: Five Principles for Interactive Human-Agent Alignment

    Recent AI alignment work often frames alignment as a pre-deployment optimization problem: collect human feedback, learn preferences or principles, finetune the model, and deploy an aligned system. This framing has produced major progress, but it under-specifies what happens once …

  74. Hugging Face Daily Papers TIER_1 English(EN) ·

    T-Search: An Open Agentic Retriever and Playground for Hard Multi-Step Search

    We present T-Search, an open-weight agentic retriever for hard multi-step search. Given a question and a search tool over a fixed corpus, it runs a bounded multi-round search and returns a ranked list of evidence chunks with short justifications, leaving answer generation to a do…

  75. arXiv cs.AI TIER_1 English(EN) · Chengyang Shi, Xianglin Ji, Jintao Huang, Jicheng Wang, Yifeng He, Jiachen Liu ·

    Open-Endedness Bench: Measuring Epistemic Process from Agent Records

    arXiv:2610.02588v1 Announce Type: new Abstract: Agents are increasingly given open-ended research tasks: discovering an empirical law from self-designed experiments, improving a heuristic whose optimum nobody knows, or beating a standing record. Their execution logs record every …

  76. arXiv cs.AI TIER_1 English(EN) · Eray Turkel, Mengsha Sun, Kartik Ayyar, Sean Dunigan, Jack Lu, Vlad Shcherban, Hsiang-Shun Shih, Xin Wang, Tiantian Zhang ·

    OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine

    arXiv:2610.02563v1 Announce Type: cross Abstract: We present OpenGameEval, a benchmark and evaluation framework for agentic game development inside Roblox Studio. It runs language models as agents in reproducible, stateful game-engine sessions and scores each run with executable …

  77. arXiv cs.AI TIER_1 English(EN) · Songtao Wei, Yi Li, Zhichun Guo, Bingzhe Li ·

    Inherit-MAS: Test-Time Evolution of Multi-Agent Systems through Workflow and Execution Inheritance

    arXiv:2610.02396v1 Announce Type: cross Abstract: Multi-agent systems (MAS) built from large language models coordinate specialized agents to tackle complex tasks, but effective workflows are difficult to design in advance. Test-time evolution refines workflows using execution fe…

  78. arXiv cs.AI TIER_1 Dansk(DA) · A. Said Gurbuz, Ahmed Nassar, Sunghwan Hong, Marc Pollefeys, Peter W. J. Staar ·

    DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

    arXiv:2610.02320v1 Announce Type: cross Abstract: Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such s…

  79. arXiv cs.AI TIER_1 English(EN) · Yu Li, Guangfeng Cai, Long-Fei Li, Shuo Han, Shengtian Yang, Han Luo, Kaibing Yang, Lei Feng ·

    Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents

    arXiv:2610.03634v1 Announce Type: new Abstract: Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier command…

  80. arXiv cs.AI TIER_1 English(EN) · Jermyn Zhen Yong Bek, Zhuang Qiang Bok, Zhongtian Sun ·

    Knowledge or Calculator? Decomposing the Skill Premium in Verifiable Financial Agent Workflows

    arXiv:2610.03564v1 Announce Type: new Abstract: Financial AI agents must do more than retrieve facts: investment workflows require correct quantitative execution, reliable use of procedural resources, and auditable structured outputs. We introduce FinSkillBench, an evaluation sui…

  81. arXiv cs.AI TIER_1 English(EN) · Chiara Troiani, Arash Salarian, Majed El Helou, Benjamin Ryder, Jean Diaconu, Herv\'e Muyal, Marcelo Yannuzzi ·

    Toward SLM-based agentic task-tool intent matching

    arXiv:2610.03213v1 Announce Type: new Abstract: Tool-equipped AI agents use tool calls to access data and act on external systems. Horizontal growth of agentic systems increases the number of these interactions, and further motivates the need for automated, per-call oversight tha…

  82. arXiv cs.AI TIER_1 English(EN) · Yulong Ming, Jie Xu, Zihan Wu, Xiaohua Jia ·

    When to Compile a Computer-Use Agent? Measuring Payback and Making Compilation Decisions for Token Efficiency

    arXiv:2610.02932v1 Announce Type: new Abstract: Compiling GUI procedures that agents execute repeatedly into programs can reduce their token costs. However, measuring payback and deciding when to compile have two challenges. First, compilation costs are uncertain because attempts…

  83. arXiv cs.AI TIER_1 English(EN) · XinPeng Shen, Lan Zhang, Yixiao Huang, Haoran Cheng, Jiewei Lai, Leilei Chen, Haoxiang Deng ·

    A GHOST in Long-Horizon Agents: Governance Hazard from Overlooked Safety Constraints across Turns

    arXiv:2610.02664v1 Announce Type: new Abstract: Long-horizon agents are now playing an increasingly significant role in assisting humans with complex problem-solving. However, it is exactly their extended interaction history that introduces an underexplored execution-safety conce…

  84. arXiv cs.AI TIER_1 English(EN) · Ankur Samanta, Yonathan Efroni, Paul Sajda, Kaveh Hassani, Anirudh Goyal ·

    Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents

    arXiv:2610.02525v1 Announce Type: new Abstract: Long-horizon research agents must decide both how to investigate and what to investigate next as evidence accumulates. This is hard to learn because such decisions are sparse in long execution traces, and their consequences may emer…

  85. arXiv cs.AI TIER_1 English(EN) · Yu Li, Zheng Zhang, Xin Liu, Shengtian Yang, Guangfeng Cai, Lei Feng ·

    Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

    arXiv:2610.02330v1 Announce Type: new Abstract: Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provid…

  86. arXiv cs.AI TIER_1 English(EN) · Xi Qin, Isabel Kurth, Xin Cui, Elin Park, Alexander Schaefer, Yaad Oren ·

    When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification Challenge

    arXiv:2610.02405v1 Announce Type: new Abstract: Using a frontier model like Claude Opus as a meta-agent to generate terminal tasks and verifiers for RL training is increasingly common. Yet a runnable Docker image and executable test suite do not guarantee a faithful end-to-end pi…

  87. arXiv cs.LG TIER_1 English(EN) · Shangyang Wu, Shuai Zhao, Ziyue Zhu, Jinyang Wu, Anh Tuan Luu, Haoran Luo ·

    SCAD: Structured Credit Assignment and Distillation for Long-Horizon Agents

    arXiv:2610.03372v1 Announce Type: new Abstract: Training long-horizon agents to solve complex tasks requires effective supervision over extended interaction sequences. However, sparse terminal rewards obscure intermediate contributions, while on-policy distillation can lose infor…

  88. arXiv cs.LG TIER_1 English(EN) · Dongsu Lee, Haoran Xu, Amy Zhang ·

    Test-time Multi-agent Coordination by Decomposed Value Gradient Flow

    arXiv:2610.02554v1 Announce Type: new Abstract: Offline multi-agent reinforcement learning (MARL) faces a persistent trade-off. Expressive generative policies can represent multi-modal coordination in the data, but cannot distinguish high-value regions, while value-optimized poli…

  89. arXiv cs.LG TIER_1 English(EN) · Pranay Kothari ·

    ArrivalBench: Agent-Generated Data Pipelines Are Correct Once and Wrong Under Time

    arXiv:2610.02363v1 Announce Type: new Abstract: Benchmarks for agent-generated data work grade a pipeline by running it once against a fixed snapshot. ArrivalBench instead re-executes the pipeline an agent leaves behind under adversarial but replayable delivery schedules (late, d…

  90. arXiv cs.AI TIER_1 English(EN) · Jugal Gajjar ·

    Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis

    arXiv:2604.10800v2 Announce Type: replace-cross Abstract: Learned classifiers deployed in agentic pipelines face a fundamental reliability problem: predictions are probabilistic inferences, not verified conclusions, and acting on them without grounding in observable evidence lead…

  91. arXiv cs.AI TIER_1 English(EN) · Sri Vatsa Vuddanti, Satwik Kumar Chittiprolu ·

    Recoverability Has a Law: The ERR Measure for Tool-Augmented Agents

    arXiv:2601.22352v2 Announce Type: replace-cross Abstract: Language model agents often appear capable of self-recovery after failing tool call executions, yet this behavior lacks a formal explanation. We present a predictive theory that resolves this gap by showing that recoverabi…

  92. arXiv cs.AI TIER_1 English(EN) · Yaxiao Liu (PwC China AI Center), Pengbo Liu (PwC China AI Center), Yiwen Liu (PwC China AI Center), Yihua Guan (PwC China AI Center), Jiaxing Song (Tsinghua University) ·

    From Migration to Calibration: Preserving Agent Capabilities across Models, Jurisdictions, and Scale

    arXiv:2609.35149v2 Announce Type: replace Abstract: Deploying, migrating, or scaling an agent can change its model, harness, infrastructure, application, and intended users. We formulate agent calibration as standards-first adaptation: define basic-capability, technical-environme…

  93. arXiv cs.AI TIER_1 English(EN) · Shiyi Kuang, Xuemei Luo, Kun Liu, Junhai Li, Rui Tian, Feng Shi, Bo Shen, Nianyu Li, Dehui Li, Ping Chen ·

    EvoRiskBench: An Evolving Benchmark for Runtime Security Risks in Workspace Agents

    arXiv:2610.03153v1 Announce Type: cross Abstract: Workspace agents combine large language models with execution harnesses to perform stateful, multi-step tasks that access or modify external resources. Existing benchmarks leave gaps in executable coverage of their runtime securit…

  94. arXiv cs.AI TIER_1 English(EN) · Jabin Koo, Soheil Abbasloo, Sungjae Lee, Jungseul Ok ·

    Dynamic Expert Pruning for Multi-Agent Systems

    arXiv:2610.02951v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures scale language models efficiently by activating only a few experts per token, but the saving is confined to computation: every expert must stay resident on the accelerator, so memory bounds w…

  95. arXiv cs.AI TIER_1 English(EN) · Mayank Rathee, Alexander Stepanov, Shalin Madabhavi, Jinhao Zhu, Raluca Ada Popa, Ion Stoica ·

    Pincer: Resource Authorization for Agents using a Digital Twin

    arXiv:2610.02569v1 Announce Type: cross Abstract: Coding agents have become increasingly long-horizon, autonomous, reliant on general-purpose shell and maintain their own persistent memory for self-improvement. While these capabilities have made the agents powerful, they have als…

  96. Hugging Face Daily Papers TIER_1 English(EN) ·

    Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation

    Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continually construct and revise prompts, thereby elici…

  97. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Dileep Kalathil ·

    AgentDiscover: Autonomous Discovery with Minimal Search Scaffolding

    Frameworks that use large language models for scientific discovery typically rely on a fixed, human-designed algorithm that decides what the model sees at each step, leaving the model only the role of proposer. The model knows nothing of the search beyond what it is shown. As mod…

  98. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Xiaohui Yan ·

    SearchJev: A Fast and Calibrated System-1 Model for Search Agents

    Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates se…

  99. Hugging Face Daily Papers TIER_1 English(EN) ·

    RobotUse: Allocating Computation, Context, and Decisions

    Robot agents must connect their intended actions to observed outcomes while retaining the context needed to revise their choices over repeated attempts. Existing interfaces often leave these choices inside predefined tools or require agents to manage detailed execution code and i…

  100. Hugging Face Daily Papers TIER_1 English(EN) ·

    ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience

    A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-context adaptation agents…

  101. Hugging Face Daily Papers TIER_1 English(EN) ·

    AutoSciBench: Autonomous Benchmark Generation for Evaluating Scientific Agents

    As agents rapidly evolve, existing benchmarks can become saturated, limiting their ability to distinguish capabilities and reveal remaining failure modes. Particularly in scientific domains, constructing and updating benchmarks requires substantial time, labor, and domain experti…

  102. Hugging Face Daily Papers TIER_1 English(EN) ·

    SearchJev: A Fast and Calibrated System-1 Model for Search Agents

    Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates se…

  103. Hugging Face Daily Papers TIER_1 English(EN) ·

    Code2Games: Enabling Coding Agents for Gaming World Generation

    Generating a high-quality gaming world from a natural-language game intent requires joint reasoning about scene structure, spatial layout, gameplay objectives, interactive entities, and executable gameplay logic. Existing coding agents can generate individual assets, scenes, or s…

  104. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Hao Wang ·

    When Debate Helps: Proposal Supply and Verification-Aware Readout in Multi-Agent Reasoning

    Multi-agent debate can improve reasoning, yet often fails to beat simple majority voting. We argue that successful debate requires two distinct mechanisms: proposal supply must surface a correct answer, and readout must identify that answer when voting misses it. We formalize the…

  105. arXiv cs.LG TIER_1 English(EN) · Jose A. Ayala-Romero, Andres Garcia-Saavedra, Xavier Costa-Perez ·

    TRACE: Tackling Real-World Resource Assignment Problems via Agentic Heuristic Design

    arXiv:2610.01887v1 Announce Type: cross Abstract: Dynamic resource assignment, the real-time allocation of task streams to heterogeneous processing nodes, is the backbone of modern computing infrastructure. While learning-based schedulers excel in research, industrial deployments…

  106. arXiv cs.LG TIER_1 English(EN) · Huancheng Chen, Xiaodi Sun, Zhaoqiong Huang, Shenyang Huang Shreya Singhal, Jingwen Lu ·

    SkillSpec: Consensus-Gated Agent Skill Evolution via Representation Specialization

    arXiv:2610.00704v1 Announce Type: new Abstract: Natural-language skills are textual procedural memories through which large language model (LLM) agents retain reusable task knowledge without updating model weights. Existing methods typically treat skills as either static artifact…

  107. Hugging Face Daily Papers TIER_1 English(EN) ·

    LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures

    LLM-based agents are increasingly capable of generating complex 3D structures, with the potential to reshape how objects are designed and realized in the physical world. Yet, producing elegant geometry is fundamentally different from producing objects that can be built and perfor…

  108. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Hui Song ·

    Engineering Sustainable Agents: A Systematic Comparison of Agentic LLMs for Developer Workflows

    Large language models (LLMs) are increasingly used in software engineering, including agentic systems that coordinate multiple agents, but impose higher computational and environmental costs. In this paper, we present a comprehensive empirical study of agentic LLM systems across …

  109. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Jungseul Ok ·

    Dynamic Expert Pruning for Multi-Agent Systems

    Mixture-of-Experts (MoE) architectures scale language models efficiently by activating only a few experts per token, but the saving is confined to computation: every expert must stay resident on the accelerator, so memory bounds where these models can be deployed. Expert pruning …

  110. arXiv cs.AI TIER_1 English(EN) · J\'er\'emie Lumbroso ·

    Cybernetic and Epistemic: A Missing Vocabulary for Trustworthy Agentic Delegation

    arXiv:2610.00961v1 Announce Type: new Abstract: As code generation is increasingly delegated to AI systems, the bottleneck is shifting from writing code to supervising the systems that write it --- a shift CS-education researchers have begun to name. This shift exposes a vocabula…

  111. arXiv cs.AI TIER_1 English(EN) · Yezhou Cheng, Runjia Du, Zeming Liu, Hang Lyu, Zehua Yang, Bojun Lin ·

    Knowing When to Yield: Grounded Arbitration of User Corrections in Text-Based Embodied Agents

    arXiv:2610.00282v1 Announce Type: new Abstract: How should an embodied agent respond when a person's correction may be wrong? We formulate grounded correction arbitration as a choice among accepting, rejecting, inspecting the world, and asking the speaker. GAVA implements this in…

  112. arXiv cs.AI TIER_1 English(EN) · Michael Hardy, Ruhana Azam, Anka Reuel, Mykel Kochenderfer, Sanmi Koyejo ·

    Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

    arXiv:2610.00651v1 Announce Type: new Abstract: Agent evaluations are increasingly used to compare LLMs and inform deployment decisions, yet ranks can reflect not only the model but also the effects of the evaluation conditions such as the scaffolds or tasks. This makes reliabili…

  113. arXiv cs.AI TIER_1 English(EN) · Andre Fu, Malik Drabla, Leon Liu, Meji Abidoye, Marek Suppa, Lata Mishra, Adnan El Assadi, Yiyuan Li ·

    Incident-Arena: Getting agents to the last nine of reliability

    arXiv:2610.00648v1 Announce Type: new Abstract: AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to their ability to execute on production incident response. This …

  114. arXiv cs.AI TIER_1 English(EN) · Sahan Paliskara, Nattaput Namchittai, Andrew Lampinen ·

    Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams

    arXiv:2610.00583v1 Announce Type: new Abstract: People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people's agents over shared resources such as a codebase, a calendar, or a budget. When each agent acts for a different user…

  115. arXiv cs.AI TIER_1 English(EN) · Haoyang Su, Weiran Huang ·

    JevSpawn: Adaptive Agentic Inference through Compositional Action Spaces

    arXiv:2610.00437v1 Announce Type: new Abstract: LLM agents generate intermediate reasoning and actions token by token, making extended interactions slow and computationally expensive. Jev-style models offer fast probabilistic predictions over finite fields, but require those fiel…

  116. arXiv cs.LG TIER_1 English(EN) · Jiayi Yang, Yifang Chen, Yuanfu Sun, Xinyan Ge, Qiaoyu Tan ·

    GraphMAS: A Systematic Benchmark of Multi-Agent Coordination for Graph Learning

    arXiv:2609.39777v1 Announce Type: new Abstract: LLM-based multi-agent systems coordinate specialized reasoning through aggregation, interaction, and adaptive control, yet their potential for graph learning remains unexplored. Graph learning is a natural setting for such systems b…

  117. arXiv cs.LG TIER_1 English(EN) · Bo Han, Qianyi Wang, Shuai Liu, Xiong Zifan, Changqiao Wu, Yuanfa Li, Pengzhi Gao, Wei Liu, Jian Luan, Heng Qu, Yunpeng Song, Zhongmin Cai ·

    Learning Reliable GUI Agents under Imperfect Priors

    arXiv:2609.39547v1 Announce Type: new Abstract: GUI agents built on large language and vision-language models still struggle on unseen applications and complex multi-step tasks, as completing real GUI tasks depends on app-specific, temporally volatile operational knowledge that i…

  118. arXiv cs.AI TIER_1 English(EN) · Joachim Baumann, Vishakh Padmakumar, Xiang Li, John Yang, Diyi Yang, Sanmi Koyejo ·

    SWE-chat: Coding Agent Interactions From Real Users in the Wild

    arXiv:2604.20779v2 Announce Type: replace Abstract: AI coding agents are being adopted at scale, yet we lack empirical evidence on how people actually use them and how much of their output is useful in practice. We present SWE-chat, the first large-scale dataset of real coding ag…

  119. arXiv cs.AI TIER_1 English(EN) · Arun Sharma ·

    Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks

    arXiv:2604.12102v3 Announce Type: replace Abstract: We describe compute-grounded reasoning (CGR), a design pattern in which code computes selected sub-problems from explicit intermediate representations before a language model answers. Spatial Atlas implements CGR as an Agent2Age…

  120. arXiv cs.AI TIER_1 English(EN) · Weiyi Wang, Xinchi Chen, Jingjing Gong, Xuanjing Huang, Xipeng Qiu ·

    AstroAgentBench: Evaluating Agentic Planning on Space Mission Planning Tasks

    arXiv:2601.11354v2 Announce Type: replace Abstract: Recent LLM-for-Space systems address mission planning, scheduling, operations support, simulator control, and autonomy, but their evaluations use different task contracts, control settings, simulators, and success criteria. We i…

  121. arXiv cs.AI TIER_1 English(EN) · Dayu Wang, Yutong Liu, Jiaye Yang, Weikang Li, Jiahui Liang, Yang Li ·

    Reducing Cognitive Overhead in Tool Use via Multi-Small-Agent Reinforcement Learning

    arXiv:2508.08882v5 Announce Type: replace Abstract: Recent advances in multi-agent systems highlight the potential of specialized small agents that collaborate via division of labor. Existing tool-integrated reasoning systems, however, often follow a single-agent paradigm in whic…

  122. arXiv cs.AI TIER_1 English(EN) · Yen-Jen Wang, Haozhe Jiang, Shuying Deng, Haoru Xue, Weirui Ye, Rocky Duan, Nika Haghtalab, S. Shankar Sastry, Pieter Abbeel, Haozhi Qi ·

    Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

    arXiv:2610.02204v1 Announce Type: cross Abstract: Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a …

  123. arXiv cs.AI TIER_1 English(EN) · Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma ·

    Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

    arXiv:2610.02122v1 Announce Type: cross Abstract: Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, …

  124. arXiv cs.AI TIER_1 Română(RO) · Rui Sun, Xihan Xiong, Qin Wang, Fei Gao, Zelin Li, Zehua Cheng, Jiahao Sun, Zhipeng Wang ·

    SoK: Decentralized Agent Economic Infrastructure

    arXiv:2610.01756v1 Announce Type: cross Abstract: Decentralized agent economies increasingly build a single task from protocols that were designed and secured separately. This creates a simple problem: a workflow can look correct at each step and still produce the wrong outcome. …

  125. arXiv cs.AI TIER_1 English(EN) · Sushant Mehta, Logan Ritchie, Edwin Chen ·

    Cross-Benchmark Transfer from RL on Agentic Coding Tasks

    arXiv:2610.00890v1 Announce Type: cross Abstract: Coding agents often fail in the last mile: they build most of a feature but drop a requirement, test only the cases their implementation already handles, break behavior that was supposed to stay intact, or validate against an unch…

  126. arXiv cs.AI TIER_1 English(EN) · Zhengyuan Jiang, Reachal Wang, Yuepeng Hu, Yupu Wang, Yuqi Jia, Neil Zhenqiang Gong ·

    Self-Evolving Coding Rules for AI Coding Agents

    arXiv:2610.00650v1 Announce Type: cross Abstract: The performance of AI coding agents is highly dependent on their underlying coding rules. However, existing coding rules are typically hand-crafted and fixed, making the process labor-intensive and often suboptimal. In this work, …

  127. arXiv cs.AI TIER_1 English(EN) · Ayan Javeed Shaikh, Arunesh Sinha, Nathaniel D. Bastian, Ankit Shah ·

    No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents

    arXiv:2610.00557v1 Announce Type: cross Abstract: Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks. Reinforcement learning (RL) and large language models (LLMs) offer complementary mechanisms for t…

  128. arXiv cs.AI TIER_1 English(EN) · Xin Heng ·

    Global Coherence: When Every Agent Is Right and the Team Is Still Wrong - A Local-to-Global Semantic Foundation for Multi-Agent Collaboration

    arXiv:2610.02036v1 Announce Type: new Abstract: AI agents can each make locally valid decisions yet jointly produce an invalid result. We call this the global coherence problem: a failure of shared state, not merely of model intelligence. Our Observation-Aliasing Impossibility Th…

  129. arXiv cs.AI TIER_1 English(EN) · Hao Wang, Ting Huang ·

    Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks

    arXiv:2610.02001v1 Announce Type: new Abstract: Small open-weight models (2-9B) run on ordinary laptops, but under cloud-scale agent harnesses they rarely complete real tasks: tool prefill overflows the context, self-correction diverges, tool demonstrations loop, and tasks are si…

  130. arXiv cs.AI TIER_1 English(EN) · Beining Wu, Zihao Ding, Jun Huang ·

    Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents

    arXiv:2610.01787v1 Announce Type: new Abstract: Self-improving GUI agents keep the trajectories they produce and return them to the agent, by fine-tuning or by retrieval into the prompt, and studies that compare the two destinations disagree. We attribute this to the unit of expe…

  131. arXiv cs.AI TIER_1 English(EN) · Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata ·

    Agents Are Systems, Not Models: Rethinking Agentic Evaluation

    arXiv:2610.01618v1 Announce Type: new Abstract: Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users d…

  132. arXiv cs.AI TIER_1 English(EN) · Zongrui Yang, Li Xintong, Runchen Xu, Zhongsheng Wang, Zhedong Lin, Haoyuan Li, Jiamou Liu ·

    MCRI: A Four-Dimensional Framework for Analyzing and Evaluating Agent Skills

    arXiv:2610.01506v1 Announce Type: new Abstract: As agents evolve from single-tool systems into modular, composite architectures, skills are becoming an important mechanism for capability development and distribution. However, the academic community lacks a structured framework fo…

  133. arXiv cs.AI TIER_1 English(EN) · Yan Luo, Selim-Antoine Lali, Jeremy Moebel, Iliass Khoutaibi, Ahmadou Aidara, Mengyu Wang ·

    Revision-Aware Independent Agent Graphs for Dynamic Reasoning

    arXiv:2610.01249v1 Announce Type: new Abstract: Conventional reasoning protocols present a fixed, preselected task, so they cannot test whether an agent propagates relevant updates, preserves unaffected work, or reconstructs a historical task binding. We therefore study \emph{dyn…

  134. arXiv cs.AI TIER_1 English(EN) · Yoonkyu Woo, Woojin Lee, Jin-Xia Huang ·

    YouRA: A Persistent-State Architecture for Evidence-Traceable Autonomous Research Agents

    arXiv:2610.01097v1 Announce Type: new Abstract: End-to-end research agents can now produce complete scientific papers, yet manuscript claims often diverge from executed experiments. This gap is structural: research state, failure histories, and claim-evidence alignment are not ma…

  135. arXiv cs.AI TIER_1 English(EN) · Jingtan Wang, Sirajul Salekin, Young mok Jung, Javier Movellan, Bryan Kian Hsiang Low, Manjot Bilkhu ·

    RISED: RubrIcs for agentic multi-environment Selection and sElf-Distillation

    arXiv:2610.00979v1 Announce Type: new Abstract: Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environ…

  136. arXiv cs.AI TIER_1 English(EN) · Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu, Nigel Collier, Tomas Pfister, Chen-Yu Lee ·

    VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

    arXiv:2610.00972v1 Announce Type: new Abstract: As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to referenc…

  137. arXiv cs.AI TIER_1 English(EN) · Timothy Kassis ·

    Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks

    arXiv:2610.00084v1 Announce Type: new Abstract: Detailed profession-specific system prompts raise token use and estimated cost per response without a consistent accuracy gain. We evaluate Scientific Agents, an open-source corpus of 503 profession-specific AGENTS.md profiles, with…

  138. arXiv cs.AI TIER_1 English(EN) · Ronghua Li, Zi Liang, Zhishan Li, Shinan Liu ·

    PG-SFT: Balancing Capability Acquisition and Retention in Offline Agent Fine-Tuning

    arXiv:2610.00949v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) on offline agent trajectories is the standard approach for training specialized tool-using agents, but forcing models to imitate reasoning and actions token by token may harm other capabilities (e.g., ge…

  139. arXiv cs.AI TIER_1 English(EN) · Xisen Jin, Jingheng Li, Zhenglun Chen, Junyi Du, Xiang Ren ·

    ReLiveGym: Evaluating Long-Lived Agents over Weeks of Replayed Reality

    arXiv:2610.00710v1 Announce Type: new Abstract: As large language model (LLM) agents become widely adopted, they are increasingly deployed for tasks that require persistent monitoring or recurring actions (e.g., market analysis). These agents are expected to operate unattended fo…

  140. Hugging Face Daily Papers TIER_1 English(EN) ·

    CUAWright: A Minimal Unified Interface for Digital Agents

    The prevailing approach to computer-use agents couples a model with a domain-specific harness: a browser or desktop environment equipped with human engineered tools that are fixed before task execution. As models' coding capabilities improve, the GUI native and static harness pre…

  141. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Xavier Costa-Perez ·

    TRACE: Tackling Real-World Resource Assignment Problems via Agentic Heuristic Design

    Dynamic resource assignment, the real-time allocation of task streams to heterogeneous processing nodes, is the backbone of modern computing infrastructure. While learning-based schedulers excel in research, industrial deployments still rely on hand-written rules that operators c…

  142. Hugging Face Daily Papers TIER_1 English(EN) ·

    TRACE: Tackling Real-World Resource Assignment Problems via Agentic Heuristic Design

    Dynamic resource assignment, the real-time allocation of task streams to heterogeneous processing nodes, is the backbone of modern computing infrastructure. While learning-based schedulers excel in research, industrial deployments still rely on hand-written rules that operators c…

  143. arXiv cs.MA (Multiagent) TIER_1 Română(RO) · Zhipeng Wang ·

    SoK: Decentralized Agent Economic Infrastructure

    Decentralized agent economies increasingly build a single task from protocols that were designed and secured separately. This creates a simple problem: a workflow can look correct at each step and still produce the wrong outcome. For example, a correct escrow may release payment …

  144. arXiv cs.AI TIER_1 English(EN) · Yuqing Zhai, Xiaohong Chen, Lingming Zhang, Sriram Vishwanath, Grigore Rosu ·

    From Verification Failures to Reusable Guidance for Coding Agents

    arXiv:2609.39022v1 Announce Type: cross Abstract: Coding agents need to establish that a program satisfies a specification and that the specification captures the requested behavior. We study how expert diagnosis of verification failures can become reusable guidance for this work…

  145. arXiv cs.AI TIER_1 English(EN) · Hongjin Qian, Chaofan Li, Kun Luo, Wenqing Wei, Jianlyu Chen, Shuqi Lu, Yuyang Hu, Hongwang Xiao, Hui Wang, Chaozhuo Li, Qiwei Ye, Zhicheng Dou, Defu Lian, Zheng Liu ·

    AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

    arXiv:2609.38288v1 Announce Type: new Abstract: We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, whi…

  146. arXiv cs.AI TIER_1 English(EN) · Hyeong Kyu Choi, Bhavana Dalvi Mishra, Jiefeng Chen, Mihir Parmar, Rui Meng, Chun-Liang Li, Xiangru Tang, Sharon Li, Jinsung Yoon, Tomas Pfister ·

    AIM: Agentic Idea Management for Automated Research

    arXiv:2609.38445v1 Announce Type: new Abstract: Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven search and identify three core challenges: organizing evolving research ideas, sele…

  147. arXiv cs.AI TIER_1 English(EN) · Mohamed Abouzahra ·

    NAQD Env: A benchmark for selective withdrawal in language agents

    arXiv:2609.38460v1 Announce Type: new Abstract: Language agents must revise planned actions when evidence changes, permission is revoked, or a stop instruction arrives. A useful response is selective: suspend affected actions, preserve unaffected work, and resume only after suffi…

  148. arXiv cs.AI TIER_1 English(EN) · Jeffrey Willette, Krishna C. Puvvada, Boris Ginsburg ·

    Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability

    arXiv:2609.38712v1 Announce Type: new Abstract: Long-horizon agentic workflows require models to sustain repeated state-dependent actions all while the context grows, sub-task complexity changes, and new data arrives. Each situation represents an independent axis along which an a…

  149. arXiv cs.AI TIER_1 English(EN) · Qiuhui Chen, Jiafan Lu, Shuaimin Tang, Tao Dai, Suyuan Wang, Chenrui Ji, Zhenglei Zhou, Weimin Zhong ·

    PathAnchor: Path-Structured Evidence for Scientific Agents

    arXiv:2609.38766v1 Announce Type: new Abstract: Scientific agents can retrieve relevant passages yet still lose functional order, mix evidence across sources, or state conclusions that exceed the retrieved record. We introduce PathAnchor, a bounded scientific reasoning system bui…

  150. arXiv cs.AI TIER_1 English(EN) · Xinhe Tian, Xiaoyue Zhang, Ziyou Zhang, Jiacheng Li, Xiaoqiang Jin, Qianchuan Zhao, Gaochen Cui ·

    STRATA: Self-Learning Through Role-Aligned Tiered Agents for Real-Time Strategy Games

    arXiv:2609.38881v1 Announce Type: new Abstract: Real-time strategy (RTS) games require agents to coordinate economic development, production and construction, base defense, unit organization, and attack timing over long matches. Existing studies have applied large language models…

  151. arXiv cs.AI TIER_1 English(EN) · Heng-Zhuang Li, Yi-Kai Zhang, Yu Wang, Yueqing Sun, Jiayuan Zhang, Qi Gu, Han-Jia Ye ·

    Consistent Plan-Act for Long-Horizon Agentic Tasks

    arXiv:2609.38891v1 Announce Type: new Abstract: Long-horizon agentic tasks demand strong reasoning and efficient execution across successive interactions with dynamic environments. A common approach decouples high-level planning from low-level execution through separate planner a…

  152. arXiv cs.AI TIER_1 English(EN) · Peng Kuang, Haibo Jin, Dehao Wu, Feiyang Deng, Xiaopeng Yuan, Jerry Wang, Haohan Wang ·

    Composing Task-specific Agent Harnesses at Test Time with Reusable Primitives

    arXiv:2609.38912v1 Announce Type: new Abstract: Agent harnesses govern how large language models (LLMs) gather context, invoke tools, verify results, preserve state, and terminate, largely affecting agent performance. However, the value of each harness mechanism can differ across…

  153. arXiv cs.AI TIER_1 English(EN) · Guanning Zeng, Jiani Wang, Wenjie Ma, Shaofeng Yin, Chenyang Wang, Shichen Liu, Angjoo Kanazawa, Wode Ni, Xiuyu Li, Andrea Zanette, Haiwen Feng ·

    Schema: Discovering Unknown Environments via Agentic Program Induction

    arXiv:2609.39140v1 Announce Type: new Abstract: Learning to complete tasks in unfamiliar environments with unknown rules remains a key challenge for LLM agents. Current LLM agents often record their discoveries in prose, which may not provide a compact, explicit account of how th…

  154. arXiv cs.AI TIER_1 English(EN) · Xinyu Zhu, Fenyi Liu, Yuzhu Cai, Shuo Tang, Rui Ye, Linfeng Zhang, Siheng Chen ·

    WorkGenesis: Building the Worlds That Teach Agents to Work

    arXiv:2609.39325v1 Announce Type: new Abstract: The ability of Large Language Model (LLM) agents to complete daily and professional work is receiving increasing attention. Training such agents requires realistic work scenarios. Expert-authored occupational work is costly and slow…

  155. arXiv cs.AI TIER_1 English(EN) · Seonho Lee, Wonryeol Jeong, Alberto Cereser, Inha Kang, Hyeonjong Kim, Seungmin Kwak, Dongmin Park ·

    A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?

    arXiv:2609.39564v1 Announce Type: new Abstract: Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form…

  156. arXiv cs.AI TIER_1 English(EN) · Fabio Rovai ·

    Who Verifies the Graph? Misspecification Attacks on Causal Action Verification for Language Agents

    arXiv:2609.40027v1 Announce Type: new Abstract: Causal action verifiers gate an agent's state-changing tool calls by checking whether each proposed intervention is identifiable against a committed action-state graph, and they issue a certificate that carries the identification ar…

  157. arXiv cs.AI TIER_1 English(EN) · Yinghui He, Yapei Chang, Khushi Bhardwaj, Daniele Molinari, Tugrul Konuk, Jan Kautz, Ali Hatamizadeh ·

    PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

    arXiv:2609.40285v1 Announce Type: new Abstract: On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the…

  158. arXiv cs.AI TIER_1 English(EN) · Chun-Wah Hsu, Kai Gong, Yu Wu, Xianhe Chen, Mengyang Liu, Jie Li, Hanyu Li, Zhixuan Liu, Naisheng Tang, Jiaying Chi, Ziheng Fan, Xuning He, Xiaokang Yang, Xue Jiang, Yihong Dong ·

    OpenCollab: A Multi-Agent Coding Framework with Programmable Collaboration and Controllable Runtime

    arXiv:2609.38345v1 Announce Type: cross Abstract: Multi-agent coding systems are designed to tackle complex software engineering tasks through collaboration. However, existing evaluations typically assume configured organizations are followed faithfully, whereas reality differs. …

  159. arXiv cs.AI TIER_1 English(EN) · Xiao Huang, Mingda Zhang, Junming Zhang, Qiang Huang, Hanwen Zhang, Yue Dai, Zijia Wang, Xiaoying Tang ·

    CollabFlow: Recursive Self-Improvement of Agent Collaboration

    arXiv:2609.38662v1 Announce Type: cross Abstract: Recursive self-improvement (RSI) lets a system improve from its own outcomes; in LLM-based multi-agent systems, Agents refine one another within a task, and outcomes improve how they collaborate across tasks. However, existing mul…

  160. arXiv cs.AI TIER_1 Dansk(DA) · Guanqun Yang, Wenlong Zhang, Tian Shi, Ping Wang ·

    SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale

    arXiv:2609.38822v1 Announce Type: cross Abstract: Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The stan…

  161. arXiv cs.AI TIER_1 English(EN) · Shijia Ge, Alex Zhou, Jianshu Zeng, Yexing Wan, Di Wu, Zelin Zheng, Yazhe Wang, Zhiqi Jia, Xuan Shangguan, Jay Zhu, Yijun Liu, Lingyu He, Sihang Wu, Xiao He, Hongcheng Gao ·

    Make Code as Policy Great Again: Frontier Agents Write, Call, and Evolve Robot Tools

    arXiv:2609.39018v1 Announce Type: cross Abstract: Frontier models can control robots, but reasoning through every reach, grasp, and retreat makes manipulation slow and token-intensive. We revisit code as policy with a different division of labor: models build executable tools, co…

  162. arXiv cs.AI TIER_1 English(EN) · Abraham Yeung ·

    Coding Agents for Coding Theory

    arXiv:2609.39081v1 Announce Type: cross Abstract: We spent five weeks using an LLM coding agent on open problems in coding theory: finding large sets of four-letter words, such as DNA barcodes, that stay far apart in edit distance. The agent wrote the verifiers and search code; a…

  163. arXiv cs.AI TIER_1 English(EN) · Meijia Chen, Hao Li, Zheng Lu, Hongshan Lin, Junbai Tian, Yichen Liu, Zijun Tian, Yufan Zou, Shuhan Sun, Hanxin Chen, Zeyu Zhang, Weizhi Du, Yueting Li, Tianyu Shi, Alaa Khamis ·

    False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

    arXiv:2609.39102v1 Announce Type: cross Abstract: Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer …

  164. arXiv cs.AI TIER_1 English(EN) · Hanwen Liu, Yuanfu Sun, Qiaoyu Tan ·

    DAGent: Evaluate-then-Grow Planning for Deep Research Agents

    arXiv:2609.39154v1 Announce Type: cross Abstract: Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting bec…

  165. arXiv cs.AI TIER_1 English(EN) · Wenjin Wang, Jiazhen Lei, Yuxin Sha, Nuwa Xi, Meng Zhao, Xingxi Yin, Qi Liu, Yuliang Shen, Zixun Sun ·

    NarrativeSteward: Coordinating Delegation, Guidance, and Verification in Agent-Assisted Interactive Narrative Authoring

    arXiv:2609.39333v1 Announce Type: cross Abstract: Autonomous AI agents can turn authors' goals into interactive narratives by independently organizing and carrying out generation and revision. As agents generate and revise extensive content, authors struggle to grasp its overall …

  166. arXiv cs.AI TIER_1 English(EN) · Frances Liu, Manny Silva, Paige Calvert, Ayu Adiati, Sarah Sanders ·

    DoGBench: Can Agents Meet Expert Standards for User-Facing Documentation?

    arXiv:2609.39909v1 Announce Type: cross Abstract: We introduce DoGBENCH (Documentation Generation Benchmark), to our knowledge, the first benchmark for generating and maintaining real user-facing software documentation. It asks whether an agent can produce documentation that expe…

  167. arXiv cs.AI TIER_1 English(EN) · Qisheng Zhou, Zhen Xiong, Qiaoyu Tan ·

    TRACE: Trajectory Selection for Parallel Scaling of Search Agents

    arXiv:2609.39912v1 Announce Type: cross Abstract: Parallel search may generate a correct answer that final-answer voting fails to select. We formulate this consolidation stage as trajectory selection and introduce TRACE (Trajectory Ranking with Aggregated Cross-Rollout Evidence),…

  168. arXiv cs.AI TIER_1 English(EN) · Jiangrui Zhao, Chenglong Li, Meng Zhang, Xiaoting Du ·

    Learning When and How to Intervene: A Hindsight-Distilled Sentinel for Coding Agents

    arXiv:2609.39957v1 Announce Type: cross Abstract: Coding agents solve repository-level tasks through sequences of actions, where a single erroneous action can misdirect subsequent decisions and increase recovery costs. Existing approaches use execution feedback for recovery or sp…

  169. arXiv cs.AI TIER_1 English(EN) · Kaixuan Fan, Kaituo Feng, Tianshuo Peng, Yilei Jiang, Manyuan Zhang, Junke Wang, Xiangyu Yue ·

    EviRover: Reinforcing Agentic Perception Beyond a Glance

    arXiv:2609.40230v1 Announce Type: cross Abstract: Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumpt…

  170. arXiv cs.AI TIER_1 English(EN) · Yong Du, Tongbo Chen, Zhengxi Lu, Yizhou Liu, Bofan Chen, Tao Jiang, Wenhao Xu, Yongliang Shen ·

    ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents

    arXiv:2609.40253v1 Announce Type: cross Abstract: Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate acti…

  171. arXiv cs.AI TIER_1 English(EN) · Yeonsung Jung, Trilok Padhi, Sina Shaham, Dipika Khullar, Joonhyun Jeong, Ninareh Mehrabi, Eunho Yang ·

    Co-Evolving Agents: Learning from Failures as Hard Negatives

    arXiv:2511.22254v5 Announce Type: replace Abstract: Self-evolving agents improve their performance on long-horizon tasks by learning from their own interactions with an environment. A common approach uses the resulting failed trajectories as negatives for preference training. How…

  172. arXiv cs.AI TIER_1 English(EN) · Victor May, Van Khue Nguyen, Aaditya Salgarkar, Yishan Wang, Diganta Misra, Huu Nguyen ·

    Evaluating Agents Across Runtime Contracts: When Mismatch Costs Efficiency or Quality

    arXiv:2603.01209v3 Announce Type: replace Abstract: In CodeAct, language-model agents write Python that calls tools and use execution feedback to choose actions. Persistent runtimes preserve Python variables between actions; stateless runtimes clear them without resetting task pr…

  173. arXiv cs.AI TIER_1 English(EN) · Daniel Mitropolsky, Riccardo Neumarker, Emanuele Rimoldi, Susan S. Hong, Tomaso Poggio ·

    Generalizing the Turing Test to Interactive Agents

    arXiv:2605.10851v2 Announce Type: replace Abstract: We initiate the study of the Generalized Turing Test (GTT), a formal generalization of Turing's imitation game from humans to arbitrary interactive agents. For agents $A$ and $B$, $A$ passes the GTT against $B$ if an instance of…

  174. arXiv cs.AI TIER_1 English(EN) · Youngmok Jung, Sirajul Salekin, Henry Tran, Javier Movellan, Zhao Huang, Manjot Bilkhu ·

    SCLATE: a Substrate for Continual-Learning Agent Training and Evaluation

    arXiv:2609.32391v2 Announce Type: replace Abstract: Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, c…

  175. arXiv cs.AI TIER_1 English(EN) · Xiao-Wen Yang, Weiyi Xu, Wen Da, Hang Xu, Canwei Li, Hong-Jie You, Pusen Dong, Yucheng Zeng, Zhaokai Luo, Yu-Feng Li, Yao Hu, Mu Chuan ·

    CompoWorld: Compositional Environment Scaling for General Agents

    arXiv:2609.33665v2 Announce Type: replace Abstract: Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require a…

  176. arXiv cs.AI TIER_1 English(EN) · Yifan Liu, Praveen Venkateswaran, Abdulhamid Adebayo, Dong Wang ·

    Auditing Agent Actions through Query-Conditioned Attribution

    arXiv:2609.33676v2 Announce Type: replace Abstract: LLM agents increasingly take consequential actions through interactions with users, policies, and external tools. Auditing these agents requires automated attribution of realized actions to their historical basis. However, exist…

  177. arXiv cs.AI TIER_1 Deutsch(DE) · Saswat Das, Parvati Viswanathan, Daniel Donnelly, Chang Huang, Sahar Abdelnabi, Ferdinando Fioretto ·

    SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents

    arXiv:2609.35596v2 Announce Type: replace-cross Abstract: Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills,…

  178. arXiv cs.CL TIER_1 English(EN) · Yuhan Guo, Jinming Liu, Liang Xu, Ziqiang Li, Jianguo Huang, Zhicheng Wang, Hu Zhu, Qiuyu Chen, Yuntao Wei, Xin Jin, Wenjun Zeng ·

    EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making

    arXiv:2609.38334v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the …

  179. arXiv cs.CL TIER_1 English(EN) · Qisheng Su, Hanchen Wang, Guanru Zhu, Huicheng Jiang, Qiuyinzhe Zhang, Kou Shi, Zhen Fang, Ziao Zhang, Qingnan Ren, Zehui Chen, Tao Gui, Feng Zhao ·

    GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis

    arXiv:2609.38923v1 Announce Type: new Abstract: Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Ex…

  180. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Chen-Yu Lee ·

    VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

    As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repea…

  181. Hugging Face Daily Papers TIER_1 English(EN) ·

    VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

    As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repea…

  182. Hugging Face Daily Papers TIER_1 English(EN) ·

    Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

    Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently…

  183. Hugging Face Daily Papers TIER_1 Dansk(DA) ·

    DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

    Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a con…

  184. Hugging Face Daily Papers TIER_1 English(EN) ·

    ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents

    Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers tok…

  185. Hugging Face Daily Papers TIER_1 English(EN) ·

    Who Verifies the Graph? Misspecification Attacks on Causal Action Verification for Language Agents

    Causal action verifiers gate an agent's state-changing tool calls by checking whether each proposed intervention is identifiable against a committed action-state graph, and they issue a certificate that carries the identification argument and a one-sided lower confidence bound. O…

  186. Hugging Face Daily Papers TIER_1 English(EN) ·

    Approval Laundering: Systematizing Approval--Execution Binding Failures in AI Coding-Agent Harnesses

    Modern AI coding-agent harnesses (Claude Code, Codex CLI, Cursor) rest their security boundary on a largely unexamined assumption: that the action A a human approves is the same action A' the harness executes, where A is fixed by a stated policy for what a scope grant or session-…

  187. arXiv cs.AI TIER_1 English(EN) · Yanfei Zhang, Xu Lin ·

    After the Fix: Transfer of Corrected Agent Experience

    arXiv:2609.34603v2 Announce Type: replace Abstract: Does repairing an episode make its experience a better memory for the next task? We transfer the same failed source before and after accepted repair to a fixed target, alongside independent execution. Our 3,300 runs cover 100 Th…

  188. arXiv cs.AI TIER_1 English(EN) · Renxi Wang, Mingshan Hee, Fajri Koto, Timothy Baldwin, Haonan Li ·

    SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

    arXiv:2609.37539v1 Announce Type: new Abstract: Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data …

  189. arXiv cs.AI TIER_1 English(EN) · Jio Oh, Seunghyun Do, Young-Jun Lee, Steven Euijong Whang, Dongyeop Kang ·

    Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym

    arXiv:2609.37267v1 Announce Type: new Abstract: Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and e…

  190. arXiv cs.AI TIER_1 English(EN) · Ido Levy, Asaf Yehudai, Segev Shlomov, Asaf Adi, Leshem Choshen ·

    Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents

    arXiv:2609.37236v1 Announce Type: new Abstract: An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on …

  191. arXiv cs.AI TIER_1 English(EN) · Savini Kashmira, Jayanaka L. Dantanarayana, Lingjia Tang, Jason Mars ·

    ContextRender: From Execution Dependencies to Agent Context

    arXiv:2609.37743v1 Announce Type: new Abstract: LLM agents performing long-horizon tasks accumulate tool results that later steps may need. Passing the full history to every invocation is costly even when it fits within the context window, while reducing it risks omitting needed …

  192. arXiv cs.AI TIER_1 English(EN) · Wesley Shu ·

    Boundary-State Control for Tool-Using Language-Model Agents: Commit-Time Consistency under State Drift

    arXiv:2609.37475v1 Announce Type: new Abstract: Tool-using language-model agents can decide that an action is permissible and execute it only after security-relevant state has changed. We study this proposal-to-commit gap and introduce BSC-R, a deterministic effect-boundary mecha…

  193. arXiv cs.AI TIER_1 English(EN) · Jungwoo Yang, In Jin Kong, Yohan Jo ·

    SelfSearch: Reward-Free Search for Self-Improving Agents

    arXiv:2609.37968v1 Announce Type: new Abstract: Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstre…

  194. arXiv cs.AI TIER_1 English(EN) · Cheng Qian, Kunlun Zhu, Beibin Li, Zhenhailong Wang, Heng Ji ·

    Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

    arXiv:2609.38143v1 Announce Type: new Abstract: Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weight…

  195. arXiv cs.AI TIER_1 English(EN) · Paras Dahal, Anton Bakhtin, Taco Cohen, Zhengxing Chen, Carole-Jean Wu, Rob Fergus, Scott Yih, Gabriel Synnaeve, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston, Anirudh Goyal ·

    Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning

    arXiv:2609.38147v1 Announce Type: new Abstract: As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to …

  196. arXiv cs.AI TIER_1 English(EN) · Ranuga Disansa, U. S. Samarasinghe, Lasith Gunawardena ·

    From Lexical Baselines to Agentic Retrieval-Augmented Generation: Structured Skill and Responsibility-Level Extraction with the SFIA Framework

    arXiv:2609.35806v1 Announce Type: cross Abstract: Automated skill extraction underpins workforce planning, yet most systems represent skills as flat labels with no notion of the responsibility level at which a skill is practiced. The Skills Framework for the Information Age (SFIA…

  197. arXiv cs.AI TIER_1 English(EN) · Linzhi Peng, Hanting Chen, Heng Chang, Ke Cheng, Bowen Du, Weifeng Lv ·

    PrimeSeeker: Capability-Oriented Supervision for Deep Search Agents

    arXiv:2609.35816v1 Announce Type: cross Abstract: Large language model search agents are often trained with synthetic questions whose difficulty is increased through larger evidence graphs, additional hops, and longer trajectories. These global properties, however, are only indir…

  198. arXiv cs.AI TIER_1 English(EN) · Lingqi Jiang, Jialuo Chen, Jianan Ma, Xinhao Deng, Xiaohu Du, Sibo Yi, Yuqi Qing, Zhenguang Liu, Qinming He, Shiwen Cui, Changhua Men ·

    MMSkillRisk: Can Agents Stay Safe When Multimodal Skills Become Traps?

    arXiv:2609.35912v1 Announce Type: cross Abstract: Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, a…

  199. arXiv cs.AI TIER_1 English(EN) · Xiaoyu Xiong, Tsun-Hsuan Wang, Yi-Ling Qiao, Tao Du, Minchen Li ·

    Text2Sim: Agentic Physics-Based Simulation Generation with Distilled Expertise

    arXiv:2609.36593v1 Announce Type: cross Abstract: Creating diverse physical simulations remains labor-intensive because assets, layout, physical parameters, motion, control, and rendering must be designed and debugged jointly. We present Text2Sim, a simulation-specialized agentic…

  200. arXiv cs.AI TIER_1 English(EN) · Haomin Qi, Xiangzhe Xu, Yiming Huang, Jingbo Shang, Chengpeng Wang ·

    WitnessGym: Benchmarking Coding Agents on the Construction of Bug Witnesses

    arXiv:2609.36635v1 Announce Type: cross Abstract: Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings …

  201. arXiv cs.AI TIER_1 English(EN) · Jiexing Qi, Yu He, Jun Liu, Qichen Huang, Shaohua Hu, Zhan Dang, Guohua Chen, Rui Yang, Wen Jiang, Yang Liu, Tao Lyu, Fangming Li ·

    VACE: Validation-Gated Alternating Co-Evolution of Agent Models and Harnesses

    arXiv:2609.37105v1 Announce Type: cross Abstract: Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates ch…

  202. arXiv cs.AI TIER_1 English(EN) · Soyeong Jeong, Sujay Kumar Jauhar, Sung Ju Hwang, Andrew Joohun Nam ·

    Follow the Entities: A Corpus Map for Agentic Search

    arXiv:2609.37226v1 Announce Type: cross Abstract: Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its lates…

  203. arXiv cs.AI TIER_1 English(EN) · Rohith Reddy Bellibatlu, Zichong Wang, Wenbin Zhang ·

    Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments

    arXiv:2609.37315v1 Announce Type: cross Abstract: Tool-using agents are entering settings where a wrong action carries real cost, and the benchmarks certifying them grade what each simulated tool call reports having done, assuming the tool did what its interface advertises. The a…

  204. arXiv cs.AI TIER_1 English(EN) · Yifan Kang, Zihan Wang, Zhiwen Fan, Bangya Liu ·

    Encore: Few-Shot Agentic Discovery of Manipulation Strategies

    arXiv:2609.37359v1 Announce Type: cross Abstract: Coding agents can now write, run, and debug programs with little human help. Robot tasks, however, are usually specified by a sentence that leaves out how to grasp, in what order to make contact, and what the result should look li…

  205. arXiv cs.AI TIER_1 English(EN) · Beining Xu, Peichun Hua, Yunming Xiao ·

    Backdoor in the Loop: Compromising Agentic Search via Malicious Retrievers

    arXiv:2609.37468v1 Announce Type: cross Abstract: Agentic retrieval-augmented generation (RAG) interleaves reasoning with repeated retrieval, giving the retriever influence over both the evidence an agent observes and its subsequent search decisions. We study retriever backdoors …

  206. arXiv cs.AI TIER_1 English(EN) · Sicheng Xie, Yitong Chen, Haidong Cao, Shunlin Lu, Zuxuan Wu, Yu-Gang Jiang ·

    Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents

    arXiv:2609.37810v1 Announce Type: cross Abstract: Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potent…

  207. arXiv cs.AI TIER_1 English(EN) · Yiming Cheng (The University of Chicago), Alfin Wijaya Rahardja (Fudan University), Mengshi Zhang (TensorBlock, Inc), Zihao Chen (TensorBlock, Inc), Zhenpeng Chen (Tsinghua University), Yiling Lou (University of Illinois Urbana-Champaign) ·

    AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems

    arXiv:2609.37864v1 Announce Type: cross Abstract: Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is further hindered by existing benchmarks, which contain only a small and fixed number…

  208. arXiv cs.AI TIER_1 English(EN) · Julien Knafou, Luc Mottin, Alexandre Flament, Paul van Rijen, Esteban Gaillac, Patrick Ruch ·

    BITEM at the NTCIR-19 R2C2 Task: Predicting Confidence from Agentic RAG Pipeline Signals

    arXiv:2609.37993v1 Announce Type: cross Abstract: The BITEM team entered both subtasks of the NTCIR-19 R2C2 task with a single agentic pipeline, in which a model searches, reads and records evidence over a movie corpus while an orchestrator holds the record and rules on what may …

  209. arXiv cs.AI TIER_1 English(EN) · Yuqiao Meng, Luoxi Tang, Sakshi Sunil Narvekar, Rupali Rajendra Vaje, Yingxue Zhang, Muchao Ye, Zhaohan Xi ·

    EquiMem: Calibrating Shared Memory in Multi-Agent Debate via Game-Theoretic Equilibrium

    arXiv:2605.09278v2 Announce Type: replace Abstract: Multi-agent debate (MAD) systems increasingly rely on shared memory to support long-horizon reasoning, but this convenience opens a critical vulnerability: a single corrupted entry can contaminate the downstream memory-augmented…

  210. arXiv cs.AI TIER_1 English(EN) · Sen Zhao, Ruiqi Kong, Zuyu Zhang, Lifeng Shen, Xinyu He, Ding Zou, Xu Zhang, Qinghua Zhang ·

    CoMemBench: Benchmarking Collaborative Memory Boundaries across Multi-Agent Workflow Topologies

    arXiv:2609.32192v2 Announce Type: replace Abstract: Multi-agent workflows require task-relevant information to be shared across agents, while irrelevant, stale, unverified, or incompatible information must remain isolated. We call this task-conditioned scope of information a coll…

  211. arXiv cs.AI TIER_1 English(EN) · Yuchen Song, Andong Chen, Wenxin Zhu, Muyun Yang, Tiejun Zhao ·

    RepoMAS: Solving Progressively Specified Tasks with Issue-Driven Multi-Agent Systems

    arXiv:2609.32490v2 Announce Type: replace Abstract: LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task requirements are sufficiently specified before execution. In practice, user requests are often incomplete, and…

  212. arXiv cs.AI TIER_1 English(EN) · Jiecong Wang, Hao Peng, Zhanyi Wang ·

    Adaptive Consistency Graph for Long-Horizon Agents

    arXiv:2609.32754v2 Announce Type: replace Abstract: Large language model agents can often make reasonable local decisions on short tasks, yet their performance degrades when success requires long sequences of dependent actions and tool calls. During execution, task requirements, …

  213. arXiv cs.AI TIER_1 English(EN) · Zhengkun Di, Bin Shi, Kai Sun, Yiming Xu, Bo Dong ·

    When Should Agents Check External State? Budgeting Observations for Stored Intentions

    arXiv:2609.37125v1 Announce Type: new Abstract: Prospective memory allows an agent to retain an intention tied to a future condition, but the stored intention does not reveal whether that condition currently holds. Checking it may require web access, multi-step tool use, and paid…

  214. arXiv cs.AI TIER_1 English(EN) · Lingrui Xu, Yangqin Jiang, Jiachang Zhang, Xubin Ren, Chao Huang ·

    AnyAct: Universal Action for Self-Evolving Agents

    arXiv:2609.37025v1 Announce Type: new Abstract: As large language models (LLMs) advance, AI agents are increasingly deployed in open-world environments to tackle complex sequential tasks (e.g., document processing, cross-application collaboration), relying heavily on actions rang…

  215. arXiv cs.AI TIER_1 English(EN) · Junjie Yao, Zhangchen Zhou, Zhi-Qin John Xu ·

    CADOC: Cache-Aware Dynamic Object Context for Long-Horizon Agents

    arXiv:2609.37012v1 Announce Type: new Abstract: For a long-horizon agent, context is the bottleneck: the history is resent with every request, the window caps task length, and reasoning degrades as the history grows. Replacing structured objects with compact retrieval Cards short…

  216. arXiv cs.AI TIER_1 English(EN) · Zeyu Gan, Zixuan Gong, Yong Liu ·

    Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents

    arXiv:2609.36892v1 Announce Type: new Abstract: As the capabilities of large language models (LLMs) continue to advance, increasing attention is turning to how to translate their abilities into useful behavior. Personal agents bring this question into everyday settings, where mod…

  217. arXiv cs.AI TIER_1 English(EN) · Bo Mao, Hang He, Linting Wang, Lizhi Lin, Maosen Zhou, Guanming Liu, Jinxiu Liu, Tianyu Huai, Chaoyun Zhang, Bingxuan Li, Kepeng Lei, Guanting Dong, Zhou Shao, Rui Zheng, Hang Yan, Jie Zhou, Chengcheng Wan, Tao Gui, Liang He, Xipeng Qiu ·

    WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

    arXiv:2609.36887v1 Announce Type: new Abstract: Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent ha…

  218. arXiv cs.AI TIER_1 English(EN) · Zhen Xiong, Qiaoyu Tan ·

    EASE: Behavior-Adaptive Skill Curation for Self-Evolving Agents

    arXiv:2609.36746v1 Announce Type: new Abstract: Agent skills provide a lightweight mechanism for self-evolving agents to accumulate reusable procedural knowledge without updating model parameters. However, existing learned skill curators typically optimize curation without explic…

  219. arXiv cs.AI TIER_1 (CA) · Gabriel Orlanski, Alex L. Zhang, Avi Trost, Vincent Sunn Chen, Frederic Sala, Aws Albarghouthi, Ludwig Schmidt ·

    Can Agents Design Libraries for Agents?

    arXiv:2609.36730v1 Announce Type: new Abstract: Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDe…

  220. arXiv cs.AI TIER_1 English(EN) · Xin Yu, Lizhu Zhang, Jiamu Bai, Yanhong Wu, Zellux Wang, Serena Li, Weiwei Li, Lingzhou Xue, Xiangjun Fan, Bo Peng ·

    MLToolBench: Learning Tool-Augmented Agents for Machine Learning Development

    arXiv:2609.36679v1 Announce Type: new Abstract: Machine learning engineering (MLE) agents have made substantial progress, but learning through ML experimentation remains costly in time and computation. Synthetic environments reduce these costs while introducing variations in data…

  221. arXiv cs.LG TIER_1 English(EN) · Mihir Chauhan, Aniket Bera ·

    Emergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent Communication

    arXiv:2609.34373v2 Announce Type: replace Abstract: Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol …

  222. arXiv cs.LG TIER_1 English(EN) · Dongchan Shin, Xing Han L\`u, Jiaqi Deng, Jay Gala, Tom\'as Vergara Browne, Jaewon Moon, Fengyuan Liu, Alexandre Drouin, Siva Reddy, Alexandre Lacoste ·

    AdaptArena: Evaluating Test-Time Personalization of Web Agents

    arXiv:2609.36488v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated strong performance on complex web navigation tasks, yet they remain brittle in real-world settings where user intentions are underspecified and preferences are heterogeneous. In pr…

  223. arXiv cs.CL TIER_1 English(EN) · Yun Peng, Zihan Wu, Zeyang Zhuang, Xin Zhou, Rui Shu, Xu Han, Chun Yong Chong, Yuan Wang, Jiakun Liu ·

    LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

    arXiv:2609.37143v1 Announce Type: cross Abstract: Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' imple…

  224. arXiv cs.CL TIER_1 English(EN) · Shinan Zhang, Tao Zhang, Qihui Zhu, Mengjie Zhang, Dong Jin, Yunpeng Hou, Shuangwu Chen, Xiaobin Tan, Quan Zheng, Jian Yang ·

    LatCom: Cross-Agent Latent Compression for Efficient Multi-Agent Collaboration

    arXiv:2609.37017v1 Announce Type: new Abstract: LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. However, directly forwarding all sender latents makes…

  225. arXiv cs.CL TIER_1 English(EN) · Amirhossein Abaskohi, Amirhossein Dabiriaghdam, Lele Wang, Peter West, Giuseppe Carenini ·

    DeepRewind: Predicting and Repairing Premature Commitments in Deep Research Agents

    arXiv:2609.36344v1 Announce Type: new Abstract: Deep-research agents conduct long-horizon investigations through iterative search, evidence evaluation, belief revision, and synthesis. However, they may commit to claims before sufficient evidence is available, causing later reason…

  226. arXiv cs.AI TIER_1 English(EN) · Yuanhao Li, Hongbo Wang, Xuhong Chen, Yiming Cao, Xunzhu Tang ·

    Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents

    arXiv:2609.33875v2 Announce Type: replace-cross Abstract: Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterfactual Rollout Replay (CRR), a training-ti…

  227. arXiv cs.AI TIER_1 English(EN) · Weiyi Xu, Xiaowen Yang, Wen Da, Hang Xu, Canwei Li, Hongjie You, Pusen Dong, Yucheng Zeng, Zhaokai Luo, Mu Chuan ·

    Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents

    arXiv:2609.33772v2 Announce Type: replace Abstract: Executable environments are critical for post-training agents on tasks that require tool use and multi-step interaction, but constructing executable tasks together with their environments remains difficult to scale. Skills provi…

  228. arXiv cs.AI TIER_1 English(EN) · Michael Lee, Zhipeng Wei, Yue Dong, N. Benjamin Erichson ·

    Divide and Inject: Can Agents Reconstruct an Indirect Prompt Injection from Fragments?

    arXiv:2609.36576v1 Announce Type: new Abstract: Agentic systems are now being widely used to orchestrate tools and reason over long contexts. However, the improving capabilities of the large language models powering these agents also create new attack surfaces for indirect prompt…

  229. arXiv cs.AI TIER_1 English(EN) · Leonardo Ferreira, Gardenia Liu, Kaden Zheng ·

    Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models

    arXiv:2609.35875v1 Announce Type: new Abstract: Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity …

  230. arXiv cs.AI TIER_1 English(EN) · Yihao Wang, Linhan Xia, Rui Liu, Zhaofeng Zhang, Hongyu Wu, Yang Yang, Jinglu He, Yu Guo, Kai Lei ·

    SAGE: A Statistical Acceptance Gate for Self-Evolving Agents

    arXiv:2609.36043v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate…

  231. arXiv cs.AI TIER_1 English(EN) · Ziyang Yu, Liang Zhao, Bowen Zhu, Hasibul Haque ·

    StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents

    arXiv:2609.36319v1 Announce Type: new Abstract: Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based…

  232. arXiv cs.AI TIER_1 English(EN) · Long Phan, Stephen K. Yang, Jason J. Lim, Mantas Mazeika, Wenyu Zhang, Zheyuan Liu, Richard Ren, Jingxiang Meng, Yaoteng Tan, Weiliang Zhao, Addison Wu, Matei Anghel, Dan Hendrycks ·

    CheatBench: Measuring Reward Gaming in AI Agents

    arXiv:2609.36308v1 Announce Type: new Abstract: Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to…

  233. arXiv cs.AI TIER_1 English(EN) · Thibaud Gloaguen, Niels M\"undler-Sasahara, Mark Niklas M\"uller, Veselin Raychev, Martin Vechev ·

    Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?

    arXiv:2602.11988v3 Announce Type: replace-cross Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md. Although this practice is strongly encouraged by agent developers, there is currently no rigo…

  234. arXiv cs.AI TIER_1 English(EN) · Ziluowen Luo, Senzhang Wang, Chaozhuo Li, Jun Yin, Hao Yan, Ming Cheng, Chenxu Wang, Songyang Liu, Litian Zhang, Qiwei Ye, Zheng Liu, Philip S. Yu ·

    Distilling Agentic Systems: A Roadmap across Models, Artifacts, and Harnesses

    arXiv:2609.36630v1 Announce Type: new Abstract: Modern agents increasingly rely on memories, tools, and execution logic, so their competence extends beyond model parameters. This shift exposes a limitation of conventional knowledge distillation, which asks how a student model imi…

  235. arXiv cs.AI TIER_1 English(EN) · Zhong Guan, Yongjian Guo, Haoran Sun, Wen Huang, Shuai Di, Likang Wu, Xiong Jun Wu, Hongke Zhao ·

    Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction

    arXiv:2605.12070v3 Announce Type: replace-cross Abstract: Asynchronous reinforcement learning improves rollout throughput for large language model agents by decoupling sample generation from policy optimization, but it also introduces a critical failure mode for PPO-style off-pol…

  236. arXiv cs.AI TIER_1 English(EN) · Ziyu Liu, Jun Chen, Lixu Wang ·

    Semantic Projection for Continual Self-Evolution of Language Agents

    arXiv:2609.36626v1 Announce Type: new Abstract: Language-model agents increasingly rely on persistent natural-language skills to adapt beyond their frozen model parameters. When a shared skill is repeatedly revised from a non-stationary, heterogeneous task stream, however, improv…

  237. arXiv cs.AI TIER_1 English(EN) · Xiaojing Sun, Yuhan Zeng, Zihua She, Xiao Wang ·

    Which Self-Improvements Should We Trust? Reliable Self-Improvement When Agents Reuse Their Benchmarks

    arXiv:2609.33180v2 Announce Type: replace-cross Abstract: As recursive self-improvement (RSI) rapidly advances, reliable evaluation becomes critical for guiding adaptive search. RSI typically relies on finite evaluation resources, such as fixed benchmarks, to determine which modi…

  238. arXiv cs.MA (Multiagent) TIER_1 Dansk(DA) · Ping Wang ·

    SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale

    Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection…

  239. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Xueqing Liu ·

    PatchHolmes: Agentic Patch Retrieval via Listwise Selection

    Patch retrieval, the task of finding the commit that fixes a known vulnerability, is the foundation of vulnerability management workflows, yet 60% to 63% of CVEs in the major advisory databases lack a patch link. We present PatchHolmes, a two-phase patch retrieval system that pai…

  240. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Zhengye Han ·

    Where Do Multi-Agent Systems Fail? Evidence-Grounded Diagnosis of Collective Mechanisms

    When a multi-agent system answers correctly, it is tempting to conclude that its agents shared, checked, and used information as intended. Yet a system can break one of its collective mechanisms, the rules that govern how agents route, admit, store, and act on shared information,…

  241. Hugging Face Daily Papers TIER_1 English(EN) ·

    GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis

    Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Existing pipelines either generate files with mode…

  242. Hugging Face Daily Papers TIER_1 English(EN) ·

    PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

    On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound ac…

  243. Hugging Face Daily Papers TIER_1 Dansk(DA) ·

    SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale

    Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection…

  244. Hugging Face Daily Papers TIER_1 English(EN) ·

    PatchHolmes: Agentic Patch Retrieval via Listwise Selection

    Patch retrieval, the task of finding the commit that fixes a known vulnerability, is the foundation of vulnerability management workflows, yet 60% to 63% of CVEs in the major advisory databases lack a patch link. We present PatchHolmes, a two-phase patch retrieval system that pai…

  245. Hugging Face Daily Papers TIER_1 English(EN) ·

    EviRover: Reinforcing Agentic Perception Beyond a Glance

    Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge…

  246. Hugging Face Daily Papers TIER_1 English(EN) ·

    A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?

    Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form Game Design Documents (GDDs) describe requireme…

  247. Hugging Face Daily Papers TIER_1 English(EN) ·

    False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

    Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so…

  248. Hugging Face Daily Papers TIER_1 English(EN) ·

    DAGent: Evaluate-then-Grow Planning for Deep Research Agents

    Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting because they support parallel execution and isolate e…

  249. Hugging Face Daily Papers TIER_1 English(EN) ·

    JevSpawn: Adaptive Agentic Inference through Compositional Action Spaces

    LLM agents generate intermediate reasoning and actions token by token, making extended interactions slow and computationally expensive. Jev-style models offer fast probabilistic predictions over finite fields, but require those fields to be specified in advance. This requirement …

  250. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Xiaoying Tang ·

    CollabFlow: Recursive Self-Improvement of Agent Collaboration

    Recursive self-improvement (RSI) lets a system improve from its own outcomes; in LLM-based multi-agent systems, Agents refine one another within a task, and outcomes improve how they collaborate across tasks. However, existing multi-agent collaboration leaves this loop open: coll…

  251. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Zilong Wang ·

    Recursive Organization Improvement: A Modeling Specification for Human--Agent Organizations

    Stronger AI agents do not automatically produce better organizations: teams must also learn which work arrangements to retain and when to reconsider them. We propose a modeling specification for recursive organization improvement and evaluate it through an executable checker, a p…

  252. Hugging Face Daily Papers TIER_1 English(EN) ·

    HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents

    Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or augment GUI interactions with application specif…

  253. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Patrick Ruch ·

    BITEM at the NTCIR-19 R2C2 Task: Predicting Confidence from Agentic RAG Pipeline Signals

    The BITEM team entered both subtasks of the NTCIR-19 R2C2 task with a single agentic pipeline, in which a model searches, reads and records evidence over a movie corpus while an orchestrator holds the record and rules on what may be submitted. A claim is admitted only once an ent…

  254. Hugging Face Daily Papers TIER_1 English(EN) ·

    Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents

    Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, t…

  255. Hugging Face Daily Papers TIER_1 English(EN) ·

    Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym

    Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three join…

  256. Hugging Face Daily Papers TIER_1 English(EN) ·

    Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents

    An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. …

  257. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Andrew Joohun Nam ·

    Follow the Entities: A Corpus Map for Agentic Search

    Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach th…

  258. Hugging Face Daily Papers TIER_1 English(EN) ·

    LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

    Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' implementation capability to produce correct code edits…

  259. Hugging Face Daily Papers TIER_1 English(EN) ·

    LatCom: Cross-Agent Latent Compression for Efficient Multi-Agent Collaboration

    LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. However, directly forwarding all sender latents makes the receiver-side context scale with both the n…

  260. arXiv cs.AI TIER_1 English(EN) · Md Shohel Arman, Igor Molybog ·

    Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer

    arXiv:2609.31587v1 Announce Type: cross Abstract: We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether co…

  261. arXiv cs.AI TIER_1 English(EN) · Weida Liang, Shi Qiu, Zhun Wang, Simon Sure, Xiaoyuan Liu, Tianneng Shi, Zhaorun Chen, Wenbo Guo, Dawn Song ·

    AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents

    arXiv:2609.31318v1 Announce Type: cross Abstract: AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software co…

  262. arXiv cs.AI TIER_1 English(EN) · Haoran Zhang, Hengtong Zhang, Zhiyu Liang, Yu Yan, Decheng Zuo, Hongzhi Wang ·

    Beyond Approved Actions: Runtime Validation of Persistent Outcomes in Agent Workflows

    arXiv:2609.31301v1 Announce Type: cross Abstract: Large language model agents increasingly act on software systems, no longer merely generating text but also changing databases and online services. However, an approved database update may succeed yet leave an unapproved notificat…

  263. arXiv cs.AI TIER_1 English(EN) · Alexander Gill, Md Farhan Ishmam, Xuyen Nguyen, Neha Bhat, Parker Henry DeYoung, Fateme Hashemi Chaleshtori, Nathan Stringham, Kenneth Marino, Ana Marasovi\'c ·

    The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

    arXiv:2609.30604v1 Announce Type: cross Abstract: Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spre…

  264. arXiv cs.AI TIER_1 English(EN) · Jiaqi Ding, Guorong Wu ·

    Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms

    arXiv:2609.30558v1 Announce Type: cross Abstract: Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a cogniti…

  265. arXiv cs.AI TIER_1 English(EN) · Mukul Chhabra, Shail Patel, Luigi Medrano ·

    CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production

    arXiv:2609.30471v1 Announce Type: cross Abstract: Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically appl…

  266. arXiv cs.AI TIER_1 English(EN) · Leon Goldberg, Gal Engelberg, Eden Yavin, Elad Elouz, Ariel Zadok, Konstantin Koutsyi ·

    Coding Agents Aren't Enough! Evaluating an Enterprise Security Brain for Agentic Cloud Investigations

    arXiv:2609.30345v2 Announce Type: cross Abstract: Cloud-security investigation is dominated by population tasks: which identities can read a data store, how many resources fail a control, what is reachable from another account. These resolve against a complete inventory, not a na…

  267. arXiv cs.AI TIER_1 English(EN) · Carolina Fortuna, Blaz Bertalanic ·

    Multi-agent Scaling Across Disjunctive and Compensatory Tasks

    arXiv:2609.31563v1 Announce Type: new Abstract: Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analy…

  268. arXiv cs.AI TIER_1 English(EN) · Bart{\l}omiej Cupia{\l}, Jens Tuyls, Maciej Wo{\l}czyk, Davide Paglieri, Martin Klissarov, Benjamin Eysenbach, Piotr Mi{\l}o\'s, Karthik R. Narasimhan ·

    Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents

    arXiv:2609.31076v1 Announce Type: new Abstract: Language agents struggle to act and learn in environments that require long sequences of low-level actions. Code-based abstractions can make these agents more productive by letting them invoke reusable skills instead of repeatedly s…

  269. arXiv cs.AI TIER_1 English(EN) · Zhensheng Zou (Peking University), Guoqing Wang (Peking University), Dan Hao (Peking University) ·

    Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents

    arXiv:2609.31430v1 Announce Type: new Abstract: Tool observations dominate the context of software-engineering agents, making long interaction histories costly to maintain. Existing context compression methods can discard information needed by later actions, while adapting agents…

  270. arXiv cs.AI TIER_1 English(EN) · Maokai Qin, Chuan Qin, Qi Zhang, Dianyu Liu, Zirui Liu, Hongting Niu, Yuanchun Zhou, Hengshu Zhu ·

    SciHorizon-eLab: An Agentic Protocol-to-Task Compiler for Scalable Benchmarking of Scientific Embodied Agents

    arXiv:2609.30971v1 Announce Type: new Abstract: Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely…

  271. arXiv cs.AI TIER_1 Deutsch(DE) · Guanyu Nie, Fangzhou Zhu, Shixiong Kai, Xiongwei Han, Tao Zhong, Mingxuan Yuan ·

    SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting

    arXiv:2609.30861v1 Announce Type: new Abstract: Language-model agents increasingly improve by converting execution experience into reusable external skills. Yet repeated skill updates form a learning process of their own: locally useful edits can accumulate into redundant or task…

  272. arXiv cs.AI TIER_1 English(EN) · Xiaoyang Li, Yiqi Wang, Chencheng Zhu, KE XU, Wencheng Yang, Zequn Sun, Pingan Song, Yiqun Duan, Taotao Cai ·

    A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

    arXiv:2609.30813v1 Announce Type: new Abstract: Evaluating claim admission in shared agent memory is challenging because repeated claims may be mistaken for independent evidence. An agent may copy or paraphrase a retrieved belief, while admitting a false claim exposes subsequent …

  273. arXiv cs.AI TIER_1 English(EN) · Shane Caldwell, Max Harley, Ads Dawson, Michael Kouremetis, Vincent Abruzzo, Will Pearce ·

    ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?

    arXiv:2609.30325v1 Announce Type: new Abstract: Agents are increasingly deployed with real autonomy in web application and network penetration testing, where a single out-of-scope action can breach a client's engagement boundary. Existing offensive-security benchmarks measure raw…

  274. arXiv cs.AI TIER_1 English(EN) · Yiran Hu, Nan Jiang, Shanchao Liang, Anik Dey, Yi Wu, Lin Tan ·

    Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents

    arXiv:2609.30725v2 Announce Type: new Abstract: Although effective, coding agents often incur substantial monetary costs. Their recurring cost-inefficient behaviors remain underexplored. We conduct the first study of behavioral cost inefficiencies in coding agents, analyzing 1,20…

  275. arXiv cs.AI TIER_1 English(EN) · Zihao Zhu, Siwei Lyu, Adel Bibi, Baoyuan Wu ·

    Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems

    arXiv:2609.30383v1 Announce Type: new Abstract: A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable …

  276. arXiv cs.AI TIER_1 English(EN) · Salma Roshdy Aly, Hussein Assaf, Ziad Kobti ·

    When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess

    arXiv:2609.30328v1 Announce Type: new Abstract: When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for. Multi-agent v…

  277. arXiv cs.AI TIER_1 English(EN) · Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth ·

    Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

    arXiv:2605.06869v3 Announce Type: replace Abstract: AI agent research spans a wide spectrum: from RL agents that learn from scratch to foundation model agents that leverage pre-trained knowledge, yet no unified benchmark enables fair comparison across these approaches. We present…

  278. arXiv cs.AI TIER_1 English(EN) · Fangzhou Li, Pagkratios Tagkopoulos, Ilias Tagkopoulos ·

    SkillFlow: Scalable and Efficient Agent Skill Retrieval System

    arXiv:2504.06188v3 Announce Type: replace Abstract: AI agents can extend their capabilities at inference time by loading reusable skills into context, yet equipping an agent with too many skills, particularly irrelevant ones, degrades performance. As community-driven skill reposi…

  279. Hugging Face Daily Papers TIER_1 (CA) ·

    Can Agents Design Libraries for Agents?

    Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an age…

  280. Hugging Face Daily Papers TIER_1 English(EN) ·

    WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

    Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in is…

  281. Hugging Face Daily Papers TIER_1 English(EN) ·

    HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents

    Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or augment GUI interactions with application specif…

  282. Hugging Face Daily Papers TIER_1 English(EN) ·

    Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym

    Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three join…

  283. Hugging Face Daily Papers TIER_1 English(EN) ·

    Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents

    An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. …

  284. Hugging Face Daily Papers TIER_1 English(EN) ·

    Follow the Entities: A Corpus Map for Agentic Search

    Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach th…

  285. Hugging Face Daily Papers TIER_1 English(EN) ·

    RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers

    Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as policies or controllers, offering limited cove…

  286. Hugging Face Daily Papers TIER_1 English(EN) ·

    SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

    Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain und…

  287. Hugging Face Daily Papers TIER_1 English(EN) ·

    EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making

    Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the cost of additional training and errors that comp…

  288. Hugging Face Daily Papers TIER_1 English(EN) ·

    AIM: Agentic Idea Management for Automated Research

    Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven search and identify three core challenges: organizing evolving research ideas, selecting promising directions, and maintaining alig…

  289. Hugging Face Daily Papers TIER_1 English(EN) ·

    Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

    Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed. To make the Builder's experience…

  290. Hugging Face Daily Papers TIER_1 English(EN) ·

    AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

    We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, which produces a solution better than the current o…

  291. Hugging Face Daily Papers TIER_1 English(EN) ·

    Shockingly Simple Self-retrospection Improves Agentic Models Without RL

    People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We …

  292. Hugging Face Daily Papers TIER_1 Deutsch(DE) ·

    SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents

    Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills, in response to user and environment feedback. However, lo…

  293. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Junpei Komiyama ·

    Self-Adapting Group of Experts for Multi-Agent Reasoning

    Multi-agent systems bring together language model agents with different roles to propose, review, and refine solutions. Each agent's response depends on its model's capabilities, the reasoning strategy defined by its system prompt, and the information in its input context. Existi…

  294. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Jiaxing Song ·

    From Migration to Calibration: Preserving Agent Capabilities across Models, Jurisdictions, and Scale

    Agents need calibration when deployment conditions change: replacing a driving model, including a foundation-to-post-trained transition; crossing jurisdictions; or scaling across heterogeneous markets and sources. Interface compatibility alone does not establish capability retent…

  295. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Zheng Liu ·

    Just-In-Time Agent Memory with Runtime Agentic Research

    Memory is critical for AI agents. Many existing agent-memory systems follow an Ahead-of-Time (AOT) design, constructing memory before a specific request arrives. While this reduces online serving cost, such request-agnostic memory construction can discard fine-grained information…

  296. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Aniket Bera ·

    Emergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent Communication

    Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer t…

  297. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Aniket Bera ·

    Emergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent Communication

    Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer t…

  298. Hugging Face Daily Papers TIER_1 English(EN) ·

    Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents

    Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many existing credit assign…

  299. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Rishika Lall ·

    When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model

    Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether rank…

  300. Hugging Face Daily Papers TIER_1 English(EN) ·

    Self-Evolving Agents via Likelihood-Guided Tool-Space Optimization

    Self-evolving agents can continually improve their behavior, while tools define the executable action space through which they interact with the environment. However, exposing the full tool library to model introduces substantial irrelevant context and can impair tool-use decisio…

  301. Hugging Face Daily Papers TIER_1 English(EN) ·

    StateGuard: Analytical-State Management with Validity-Aware Intervention for Long-Horizon Data Agents

    LLM-based agents have shown strong capabilities in automated data analysis and are increasingly moving toward long-horizon, multi-stage analytical workflows. However, as the analytical process evolves, constraints, variables, and conclusions remain implicitly embedded in interact…

  302. Hugging Face Daily Papers TIER_1 English(EN) ·

    AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?

    Recent gains in language model capability have come more from data than from architecture. Frontier labs and data companies produce verifiable agentic tasks, which supervised finetuning and reinforcement learning then turn into capability.This production line still rests on human…

  303. Hugging Face Daily Papers TIER_1 English(EN) ·

    CheatBench: Measuring Reward Gaming in AI Agents

    Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized info…

  304. Hugging Face Daily Papers TIER_1 English(EN) ·

    Org-Agent: Beyond Personal Assistants Towards Organizational Agents

    Language model agents serving organizations must coordinate requests from multiple users while using knowledge distributed across their interactions. We identify two complementary capabilities for this setting, namely cross-user interaction and decision-making, as well as cross-u…

  305. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Muning Wen ·

    TRACE: Governing Memory Validity in Evolving Multi-Agent Systems

    Persistent memory lets language-model agents carry information across long-running collaborations, but leaves a lifecycle question open: what may a returning agent still act on once the shared state has changed? A memory can be correctly retrieved, relevant to the current task, a…

  306. arXiv cs.MA (Multiagent) TIER_1 English(EN) · EverMind AI ·

    Raven: The Harness of Harnesses for Composable Agentic Intelligence

    As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to speci…

  307. arXiv cs.MA (Multiagent) TIER_1 Deutsch(DE) · Manik Gupta ·

    CORTEX: A Verified Experience Layer for Generalist Agents

    An agent can solve a task today and face the same task under new facts, tools, or governing knowledge tomorrow. Most agent systems can retrieve relevant text or recall prior conversations, but they lack a principled way to decide when a previous solution is still valid, when it m…

  308. Hugging Face Daily Papers TIER_1 English(EN) ·

    WideSWE: Can Coding Agents Coordinate Changes Across Repositories?

    Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multipl…

  309. Hugging Face Daily Papers TIER_1 English(EN) ·

    Raven: The Harness of Harnesses for Composable Agentic Intelligence

    As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to speci…

  310. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Kaden Zheng ·

    Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models

    Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity among agents is the driver, in the setting where…

  311. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Yunming Xiao ·

    Backdoor in the Loop: Compromising Agentic Search via Malicious Retrievers

    Agentic retrieval-augmented generation (RAG) interleaves reasoning with repeated retrieval, giving the retriever influence over both the evidence an agent observes and its subsequent search decisions. We study retriever backdoors that exploit this feedback loop and repurpose weak…

  312. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Changlun Li ·

    When Better Gets Worse: Improvement Fidelity for Self-Improving Agents in Adaptive Worlds

    Self-improving agents increasingly rely on proxy verifiers to choose policy updates, yet deployment can change the world in which those updates are evaluated. An update that looks better to the verifier can therefore become worse after deployment even when the verifier ranks poli…

  313. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Ying Lin ·

    AsynCodeBench: Benchmarking Collaboration of Asynchronous Multi-Agent Systems in Software Engineering

    Multi-agent coding has emerged as an increasingly active direction in software engineering, where complex development tasks are decomposed across multiple specialized agents working on different parts of the problem. Despite the shift from individual problem solving to distribute…

  314. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Xiping Hu ·

    Enabling Timely Guidance before Skill Retrieval: Retaining Helpful Warm Tips in Agent Context

    Reusable skills help LLM-based agents solve complex tasks, but the agent must receive guidance before it commits to an ineffective approach. Existing skill mechanisms often expose only metadata and load full content on demand, leaving useful guidance unavailable until the agent d…

  315. Hugging Face Daily Papers TIER_1 English(EN) ·

    AgentTell: Behavioural Side-Channel Leakage in Browser-Use Agents

    Browser-use agents often carry information in their context as they move between websites. While it may be necessary for task completion, it also creates a privacy risk, especially when the information contains a private fact regarding the user. For example, an agent may learn a …

  316. Hugging Face Daily Papers TIER_1 English(EN) ·

    ExpVoyager: Direct Experience Navigation for Dynamic Agent Skill Synthesis

    Learning from experience in LLM agents has become a key paradigm for developing self-evolving agents that continuously learn and expand their capabilities. Within this paradigm, synthesizing the agent skill has emerged as a promising solution for transforming accumulated experien…

  317. Hugging Face Daily Papers TIER_1 English(EN) ·

    Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

    Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in impleme…

  318. Hugging Face Daily Papers TIER_1 English(EN) ·

    X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization

    Multi-step agents are trained on flat action streams: SFT and RLVR weight every token uniformly and ignore the sub-procedures that recur across tasks, the hierarchy that lets humans plan top-down from reusable routines. This structure sits unused, and flat training uses each scar…

  319. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Blaz Bertalanic ·

    Multi-agent Scaling Across Disjunctive and Compensatory Tasks

    Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the an…

  320. arXiv cs.AI TIER_1 English(EN) · Beining Wu, Zihao Ding, Jun Huang ·

    ERRAND: Budgeted Maintenance of Agent Memory

    arXiv:2609.29545v1 Announce Type: new Abstract: Deployed agents run on handed-over knowledge: a frozen policy consults a briefing of consolidated items written before the stream begins. The world then moves while the store stands still: paths close, flags change, price bands move…

  321. arXiv cs.AI TIER_1 English(EN) · Yezhou Cheng, Runjia Du, Zeming Liu, Qibai Chen, Hang Lyu, Yankai Zeng, Yilan Wei, Bojun Lin ·

    Scope Before You Persist: Preventing Cross-Family Interference in Agent Memory

    arXiv:2609.29144v1 Announce Type: new Abstract: Persistent memory lets language-model agents improve prompts and skills without updating model weights. We show that matching retrieval scope to certification scope enables these edits to support reliable repeated adaptation across …

  322. arXiv cs.AI TIER_1 English(EN) · Mingxuan Wang, Fei Luo, Bo Wang, Guorun Yao, Yinglong Guo, Chao Ning, Hongyue Chen, Yanbiao Ma, Jungong Han ·

    When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression

    arXiv:2609.29875v1 Announce Type: new Abstract: Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. Unlike static Chain of Thought compression, remo…

  323. arXiv cs.AI TIER_1 English(EN) · Hanjing Shi, Dominic DiFranzo ·

    When Agents Act Unwatched: The Reduced-Supervision Paradox in Agentic AI

    arXiv:2609.29547v1 Announce Type: cross Abstract: Agentic AI is sold on a simple promise: the system keeps acting when the user stops watching. That promise creates an accountability inversion. As stepwise supervision recedes, verification does not disappear; it moves into the ru…

  324. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Furu Wei ·

    Agensh: Scaling Organizational Intelligence to 1,024 Agents

    A multi-agent system can reduce latency on complex tasks by executing work concurrently. Several pioneering harness frameworks support multi-agent systems. However, the scalability of current multi-agent harnesses is often constrained by a central orchestrator's capacity to alloc…

  325. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Shuang Guo ·

    SkillApt: Learning When to Activate Agent Skills from Counterfactual Evidence

    Large language model agents increasingly retrieve reusable Skills and inject them into the active context. However, a retrieved Skill can be relevant yet unnecessary, costly, or even harmful in the current execution state. We present SkillApt, a post-retrieval activation framewor…

  326. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Oren Gal ·

    MATES: Learning Multi-Agent Interactions by Transforming Observations for Frozen Single-Agent Policies

    Multi-agent reinforcement learning (MARL) commonly trains decentralized policies from scratch, requiring agents to acquire individual task competence and coordination simultaneously. Yet many multi-agent problems admit a compatible single-agent counterpart in which the underlying…

  327. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Andrea Baronchelli ·

    Indirect tipping: a social attack surface in AI agent populations

    As generative AI agents are deployed at scale, safety will depend not only on technical safeguards and individual model design, but also on collective equilibria that determine how agent populations process information, prioritize actions, and respond to uncertainty. Yet the same…

  328. Hugging Face Daily Papers TIER_1 English(EN) ·

    OSWorld-Pro: Process-based Evaluation for Computer Use Agents

    Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with functional verifiers, as seen in OSWorld. However, such evaluation of end-state performance lacks transparency into how and why agen…

  329. arXiv cs.MA (Multiagent) TIER_1 English(EN) · James Zou ·

    Self-Organizing Agent Teams Learn to Reason Together

    Collective intelligence depends not only on what team members know, but also on how they organize their work. When the structure of a solution is unknown, useful roles and divisions of labor cannot be specified in advance; teams must learn from experience how to organize reasonin…

  330. arXiv cs.CV TIER_1 English(EN) · Liu Renhang, Navonil Majumder, Tej Deep Pala, Soujanya Poria ·

    RoboQuest: Generalist Physical Agents that Search, Inspect and Test

    arXiv:2610.10388v1 Announce Type: cross Abstract: Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-rele…

  331. arXiv cs.CV TIER_1 English(EN) · Ling Li, Qiuyu Shen, Zheng Jiang, Qinwei Ma, Yuxuan Liu, Zhidong Deng ·

    SkillCycle: Co-Evolving Agent Policies and Skill Banks

    arXiv:2610.09430v1 Announce Type: new Abstract: Internalizing external skills changes a language agent's capabilities and, with them, the value of its remaining guidance: rules can become redundant, misleading, or insufficient for newly encountered decisions. This creates a coupl…

  332. arXiv cs.CV TIER_1 Italiano(IT) · Sacha Morin, Kumaraditya Gupta, Mahtab Sandhu, Charlie Gauthier, Francesco Argenziano, Kirsty Ellis, Liam Paull ·

    Agentic Scene Policies

    arXiv:2509.19571v2 Announce Type: replace-cross Abstract: Designing or learning robot policies that generalize zero-shot across a range of language instructions and objects is a core problem in robotics. Vision-Language-Action models (VLAs) learn such policies end-to-end by repur…

  333. arXiv cs.CV TIER_1 English(EN) · Shenxiang Zeng, Chen Yang, Peiyao Chen, Guohui Zhang, Jiansheng Fan, Chen Wang ·

    Asking the World: Generalist Physical Reasoning through Agentic World Modeling and Probing

    arXiv:2609.39135v1 Announce Type: new Abstract: Physical reasoning from video requires inferring latent physical properties and dynamics beyond direct observation. Direct VLM inference remains unreliable on complex physical tasks without explicit modeling and validation, while pr…

  334. arXiv cs.CV TIER_1 English(EN) · Tongbo Chen, Junbo Niu, Zhengxi Lu, Niu Lian, Fei Tang, Yuchen Yan, Yike Hong, Yong Du, Yizhou Liu, Bofan Chen, Yongliang Shen ·

    HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents

    arXiv:2609.38008v1 Announce Type: new Abstract: Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or …

  335. arXiv stat.ML TIER_1 English(EN) · Shengjun Zhang, Tingyi Liu, Dong Xie, Yunlong Dong, Xiang Wang, Cheng Zeng ·

    Causal Retention in Interactive Agents: Interface Factorization and Selective Adaptation

    arXiv:2609.30650v1 Announce Type: cross Abstract: Task performance need not determine which intervention mechanism an agent retains. We study causal retention: whether a frozen learned state answers a mechanism-probe map fixed independently of training, including action, context,…

  336. AWS Machine Learning Blog TIER_1 English(EN) · Manish Ballal ·

    Beyond hours saved: Building the business case for agentic automation

    The RPA-era ROI model misses most of the value agentic automation creates. This post gives AI center of excellence leaders a framework to size the full value of agents across time savings, exception handling, decision quality, and maintenance economics, and to prioritize which wo…

  337. AWS Machine Learning Blog TIER_1 English(EN) · Thiago Verney ·

    Building a context-aware AI assistant on AgentCore and OpenClaw

    Off-the-shelf AI assistants forget you between conversations. This post shows how to build a personal assistant that accumulates context using OpenClaw on Amazon Bedrock AgentCore runtime, with AgentCore memory turning disposable chats into durable, structured knowledge you can r…

  338. AWS Machine Learning Blog TIER_1 English(EN) · Mona Mona ·

    New agent skill: Amazon SageMaker optimized generative AI inference for your coding agent

    Amazon SageMaker optimized generative AI inference introduces the aws-ai-ml skill through the Agent Toolkit for AWS, giving coding agents like Kiro, Claude Code, and Codex deep expertise in inference optimization and benchmarking. Describe what you want, and your agent generates …

  339. AWS Machine Learning Blog TIER_1 English(EN) · Manideep Reddy Gillela ·

    Agentic retrieval with LangChain and Amazon Bedrock Knowledge Bases

    Build a Retrieval Augmented Generation (RAG) application on Amazon Bedrock Managed Knowledge Base with LangChain, and see how agentic retrieval handles the multi-part questions that single-shot retrieval answers poorly. Run the same query through both paths, read the trace events…

  340. AWS Machine Learning Blog TIER_1 English(EN) · Kanishk Mahajan ·

    Evaluating multi-agent systems for explainability and helpfulness with Amazon Bedrock AgentCore

    Multi-agent systems need deeper guarantees than fluent responses: they must select the right tools, respect constraints, and explain their decisions. Learn how to build a Strands-based multi-agent supply chain decisioning system and evaluate it with Amazon Bedrock AgentCore Evalu…

  341. Modal blog TIER_1 English(EN) ·

    VM Sandboxes: Full computers for agents

    VM Sandboxes are built for those who need to give their agents the power of a full computer.

  342. Databricks Blog TIER_1 English(EN) ·

    How to scale agentic applications without creating AI sprawl

    Building an agent is getting easier. More capable models and coding agents are making...

  343. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    JetBrains Releases Mellum2.1: A 12B MoE Open Model for Coding Agents

    <p>JetBrains released Mellum2.1, an Apache 2.0, 12B mixture-of-experts thinking model with 2.5B active parameters. RL in real repositories lifted its SWE-bench Verified score from 2.0 to 47.0.</p> <p>The post <a href="https://www.marktechpost.com/2026/10/08/jetbrains-releases-mel…

  344. dev.to — Claude Code tag TIER_1 English(EN) · Steven Gonsalvez ·

    Your Agent Is a While Loop: Harness, Loop, and Graph Engineering

    <p><em>Human thoughts, AI-assisted write-up.</em></p> <p>Strip an agent back to its skeleton and it's embarrassingly simple. A while loop.<br /> </p> <div class="highlight js-code-highlight"> <pre class="highlight plaintext"><code>while not done: context = observe() # read files,…

  345. dev.to — Claude Code tag TIER_1 English(EN) · saaro ·

    Agentic Coding in Production 2026: From Autocomplete to Autonomous Developer

    <p>The leap from simple chat prompts to agentic workflows is the biggest upheaval in software development since the introduction of Git. While 2024 and 2025 were still dominated by autocomplete plugins and "vibe coding" in the headlines, the landscape has fundamentally changed by…

  346. MarkTechPost TIER_1 English(EN) · Sana Hassan ·

    IBM Brings Bob to Self-Hosted and Air-Gapped Environments: Agentic Software Development Without Moving Your Code

    <p>IBM has made a self-hosted deployment of IBM Bob, its agentic software development platform, generally available. Enterprises can now run Bob on premises, in private or sovereign clouds, and in air-gapped networks. They bring their own model: NVIDIA Nemotron or Poolside Laguna…

  347. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    IQuest Research Open-Sources IQuest-Q1, a 320B MoE Model for Agentic Coding With 15B Active Parameters

    IQuest Research released open weights for IQuest-Q1, a 320B sparse MoE with 15B active parameters built for command-line coding agents, with 512K context and 84.5 on CyberGym.

  348. dev.to — MCP tag TIER_1 English(EN) · Fernando Azevedo ·

    Four engineering patterns that separate multi-agent from prompt chains

    <p>Google's post on the AI Agents Challenge says something it took me years to accept on financial platforms: "multi-agent" was the most frequent claim across thousands of submissions, and a good share of them were a single model walking through a prompt chain with agent names at…

  349. Medium — Claude tag TIER_1 English(EN) · Fellows Monika ·

    Learning About Coding Agents

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@fellows.monika/learning-about-coding-agents-55e5c5bba4e6?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/1*hoRDlGu2zo_Gcjj0P8Emiw.png" width="2784" /></a></p><p cl…

  350. Medium — AI coding tag TIER_1 English(EN) · AIHoony ·

    The irreversible command: Why coding agents need a human checkpoint

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@kwd8819/the-irreversible-command-why-coding-agents-need-a-human-checkpoint-d69037bc397b?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1672/1*gLz8vt3r3JJzBAcaOJMH1Q…

  351. Email — Every TIER_1 English(EN) · 010001a117d0d056-a9eaa087-9035-4580-bb51-9434736e88d3-000000@send.every.to (010001a117d0d056-a9eaa087-9035-4580-bb51-9434736e88d3-000000@send.every.to) ·

    Building a More Efficient Agent

    <!-- Set the language of your main document. This helps screenreaders use the proper language profile, pronunciation, and accent. --> <!-- The title is useful for screenreaders reading a document. Use your sender name or subject line. --> Building a More Efficient Agent <!-- Neve…

  352. dev.to — Anthropic tag TIER_1 English(EN) · Jorge Peraza ·

    Hardening Prompt Caches and SSE Streams for Coding Agents

    <p><strong>TL;DR:</strong></p> <ul> <li> <strong>Anthropic SSE Keep-Alives:</strong> We updated the model proxy to inject keep-alive pings during extended Anthropic "thinking" phases, preventing load balancer and NAT gateway timeouts.</li> <li> <strong>GPT-6 Cache Breakpoints:</s…

  353. Medium — MCP tag TIER_1 English(EN) · OpenVidu ·

    Introducing the OpenVidu Agent Plugin for coding agents

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://openvidu.medium.com/introducing-the-openvidu-agent-plugin-for-coding-agents-064ee095f74d?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/2600/1*7mlDKz47mAixRxsx3rXZrA.png" width="3200…

  354. dev.to — MCP tag TIER_1 English(EN) · CAI ·

    Why CAI exists: the agent-payment problem and the three roles (custodian, approver, operator)

    <h1> Why CAI exists: the agent-payment problem and the three roles (custodian, approver, operator) </h1> <p>The agent-payment problem is older than agents, and it has a clear shape. This post is the H1-readable essay on the problem CAI solves, the user-confirmation pattern, and t…

  355. dev.to — MCP tag TIER_1 English(EN) · Baris Sozen ·

    When Agents Hire Agents: The Delegation Payment Problem

    <p>The agent economy doesn't wait for perfect infrastructure. Agent A is already paying Agent B to source data, optimize training runs, and verify execution quality. Agent B is already hiring Agent C to do the actual work. And today, all three are using custodians to move money.<…

  356. Medium — MLOps tag TIER_1 English(EN) · Swapnil Surushe ·

    Building the Conductor: A Scalable AI Agent Orchestrator on Cloud Run

    <div class="medium-feed-item"><p class="medium-feed-snippet">This is Part 2 of the series, The Enterprise AI Agent Blueprint. Read Part 1 here.</p><p class="medium-feed-link"><a href="https://medium.com/@swapnil29071999/building-the-conductor-a-scalable-ai-agent-orchestrator-on-c…

  357. dev.to — MCP tag TIER_1 English(EN) · Luis Alcaraz ·

    Introducing TAP: software building blocks for the agentic world

    <p>We've open sourced TAP (Trusted Agent Primitives) at Telara so agents can build reusable internal tools for the work they do repeatedly. The packages are code that developers can write, inspect and maintain too.</p> <p>Employees ask agents to investigate issues, prepare custom…

  358. dev.to — MCP tag TIER_1 English(EN) · Luis Alcaraz ·

    Introducing TAP: software building blocks for the agentic world

    <p>We've open sourced TAP (Trusted Agent Primitives) at Telara so agents can build reusable internal tools for the work they do repeatedly. The packages are code that developers can write, inspect and maintain too.</p> <p>Employees ask agents to investigate issues, prepare custom…

  359. Towards AI TIER_1 English(EN) · Satish Kumar ·

    Production RBAC, Cost Optimization, and Deployment Patterns for Cortex Agents

    <p><em>The developer-to-production pipeline for Snowflake’s Cortex Agent GA enhancements — Personal Database sandboxes, temporary agents, COPY GRANTS, and the cost math on Cortex Search suspension.</em></p><p><strong>Part 2 of 2</strong> — Part 1: <a href="https://snowflakechroni…

  360. Towards AI TIER_1 English(EN) · Enzo Lombardi ·

    Coding an Agent: Steering a Local Model

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/coding-an-agent-steering-a-local-model-14872070ef56?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1024/1*jOao6kz_lfj1-OR-fPe18A.png" width="1024" /></a></…

  361. Medium — fine-tuning tag TIER_1 English(EN) · Sasha Denisov ·

    Fine-Tuning Small Models for On-Device Agents: Train, Convert, and Measure What Changed

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/google-developer-experts/fine-tuning-small-models-for-on-device-agents-train-convert-and-measure-what-changed-f7f494a044f3?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.c…

  362. Medium — fine-tuning tag TIER_1 English(EN) · Sasha Denisov ·

    Fine-Tuning Small Models for On-Device Agents: Train, Convert, and Measure What Changed

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@denisov.shureg/fine-tuning-small-models-for-on-device-agents-train-convert-and-measure-what-changed-f7f494a044f3?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/13…

  363. Towards AI TIER_1 English(EN) · Rajesh K ·

    Jev and RLCD: Architecture, Open-Source Models, and Practical Agent Use Cases

    <h4><em>A practical guide to typed AI decisions, calibrated probabilities, and the software that turns them into useful workflows.</em></h4><p>An agent receives a request: “Investigate the failed deployment and explain what changed.”</p><p>Before it produces an answer, the system…

  364. Medium — Claude tag TIER_1 English(EN) · Rathish Poovadan ·

    Part 2: Beyond the Terminal — Prototyping Claude Agents in Jupyter Notebooks

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/the-repo/part-2-beyond-the-terminal-prototyping-claude-agents-in-jupyter-notebooks-8be734ad85dc?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/1*B-jTb3C_wdjK3OjuzT…

  365. Medium — MCP tag TIER_1 English(EN) · Ahmet Kalafat ·

    From One Agent to an Agent Team: ADK, Agent Skills, MCP and A2A

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ahmet.kalafat/from-one-agent-to-an-agent-team-adk-agent-skills-mcp-and-a2a-6bc9b48a171c?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1760/1*60kJ2Tm_42iFc7IEOy3dcw.png" …

  366. Towards AI TIER_1 English(EN) · Surya Maddula ·

    Developing Sophisticated Controllable Agents with RAG.

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/developing-sophisticated-controllable-agents-with-rag-727f28828815?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1672/1*2Nth5ujXg2CeRFiKDVsYMA.png" width=…

  367. Towards AI TIER_1 English(EN) · Dave R | Microsoft Azure & AI MVP ☁️ ·

    An AI Agent Workflow That Survives kill -9: Inside Microsoft Agent Framework’s Latest Update

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/microsoft-agent-framework-kill-9-crash-recovery-ag-ui-memory-codeact-d4812e5a7bd2?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1672/1*wO-0z1k70zalhDFv_yq…

  368. dev.to — MCP tag TIER_1 English(EN) · Baris Sozen ·

    The Settlement War: Why Routing Isn't Enough for Agent Commerce

    <h1> The Settlement War: Why Routing Isn't Enough for Agent Commerce </h1> <p>This week, three signals converged.</p> <p>On Tuesday, the <strong>IETF published draft-hood-agtp-commerce-00</strong> — the agent-to-agent commerce standard. It is a big deal. Agents can now discover e…

  369. Medium — MCP tag TIER_1 English(EN) · vedant Padole ·

    WebMCP: A Beginner’s Guide to the Agentic Web with an Example

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@vedantpadole05072/webmcp-a-beginners-guide-to-the-agentic-web-with-an-example-ce091f3e3b40?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1774/1*-CEGRGIjzgJoZTPGQ2ko0g.pn…

  370. dev.to — MCP tag TIER_1 English(EN) · Renato Marinho ·

    Debugging Agent Reasoning: Why Structural Integrity Matters More Than Accuracy

    <p>When we talk about LLM reliability, our focus almost always gravitates toward accuracy—did the model get the math right? Did it retrieve the correct record from the database?</p> <p>But for engineers building autonomous agents using ReAct or Chain-of-Thought (CoT) patterns, th…

  371. dev.to — MCP tag TIER_1 English(EN) · Shakar Bisetty ·

    Part 2: The reference architecture for an agentic change-approval MVP on MuleSoft, on one page

    <p><em>Part 2 of 10 · Building an Agentic Change-Approval MVP on MuleSoft</em></p> <p>In <a href="https://dev.to/thasha/the-21-step-change-nobody-wants-to-own-why-sap-change-promotion-is-an-agentic-use-case-4mpo">Part 1</a> I set out the use case: automating the approval and prom…

  372. dev.to — MCP tag TIER_1 English(EN) · mech.app ·

    MCP Reference Servers: What 90,000 Stars and 16 Language SDKs Reveal About Agent Tool Boundaries

    <p>The Model Context Protocol repository sits at 90,950 stars with 16 language SDKs and a collection of reference servers that expose how agent-tool boundaries actually work. These are not production systems. They are educational implementations that reveal transport layer choice…

  373. dev.to — MCP tag TIER_1 English(EN) · Praveen Raj Thulasi S ·

    I Built an Agentic Analytics Platform — Here's What I Learned

    <h1> I Built an Agentic Analytics Platform — Here's What I Learned </h1> <p>What if you could ask your analytics dashboard:</p> <blockquote> <p><strong>"Why did sales decrease last month?"</strong></p> </blockquote> <p>and instead of manually filtering charts and writing database…

  374. Medium — Claude tag TIER_1 Nederlands(NL) · Ivan Yanishevskyi ·

    Subagents vs Agent Teams in Claude Code: when one agent isn’t enough

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ivan.yanishevskyi/subagents-vs-agent-teams-in-claude-code-when-one-agent-isnt-enough-bc062883070c?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1638/1*LWLjCpWV9vXjVEW…

  375. Towards AI TIER_1 English(EN) · Andrei Besleaga (Nicolae) ·

    AgenticSystemCore: a folder of Markdown texts that people and agents can use in multiple ways

    <h4>Existing technologies for future simpler combined understanding and use</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*hk7w8hkMMAkvQ4me.jpeg" /><figcaption>(AgenticSystemCore — Text used by humans and Agentic AI, source:Author, Gemini)</figcaption></f…

  376. Medium — MCP tag TIER_1 English(EN) · Ferry Djaja ·

    Teaching Websites to Talk to AI Agents, Without Changing the Websites

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://djajafer.medium.com/teaching-websites-to-talk-to-ai-agents-without-changing-the-websites-df0ed020cb79?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/2600/1*r2rir1bMG-pAPE8JP_wjBA.png…

  377. Towards AI TIER_1 English(EN) · Anna Jey ·

    AI Coding Agent Context Budget: Test the Instructions That Make Agents Slower

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*zgDOnsQQy6eaFaf_NN4E9g.jpeg" /><figcaption>AI Coding Agent Context Budget</figcaption></figure><p>More instructions can feel safer. For a coding agent, they can also bury the one rule that prevents a costly mista…

  378. dev.to — MCP tag TIER_1 English(EN) · lizer yang ·

    Agentic Search vs RAG: A Tool Call or an Index You Own

    <p><strong>Short answer:</strong> An agent's search tool and a RAG index are both retrieval, and they differ in<br /> four places that decide everything downstream: what they read (the live web against a corpus you<br /> ingested), who owns the ranking (an engine you do not opera…

  379. Towards AI TIER_1 English(EN) · Becca Sees ·

    From Hours to Minutes: Automating Data Debug with Multi-agent Systems

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*C9cnr-mVxVYEVIuOsEgnMw.jpeg" /></figure><h3>The Data Debugging Grind</h3><p>My engineers typically spend 30% of their time debugging data. On some days, that number climbs to 50% or more.</p><p>If you’re a data o…

  380. Medium — AI coding tag TIER_1 English(EN) · Dhmkchs ·

    Three coding agents, one window: how I actually work with codeme

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@dhmkchs/three-coding-agents-one-window-how-i-actually-work-with-codeme-0f74a0817cd6?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/2560/1*w0uxHUuDhwYY9BmS88Ykbg.png…

  381. Towards AI TIER_1 English(EN) · Mukesh Kumar Shah ·

    Mastering AI Agents: From ReAct to Production Multi-Agent Systems

    <h4><em>From “What is an agent?” to production-grade, multi-agent, tool-using, memory-equipped systems — everything you need in one place.</em></h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*RzS8r6JYTDqLuc2V" /></figure><p><strong>Reading time:</strong> ~…

  382. The Register — AI TIER_1 English(EN) ·

    Close the observability gap with agentic observability

    SPONSORED POST: How agentic AI, real-time visibility, and stronger governance can help enterprises protect critical services and manage increasingly complex IT environments.

  383. dev.to — MCP tag TIER_1 English(EN) · MANI BHUSHANAM K ·

    I Gave a Coding Agent Memory—and Let It Learn Which Tools It Needed

    <p>🤔 What if a coding agent could actually remember?</p> <p>AI coding agents are becoming surprisingly capable.<br /> They can write code, inspect repositories, debug errors, interact with tools, and reason through complex development tasks.</p> <p>But there is still a frustratin…

  384. Medium — MLOps tag TIER_1 English(EN) · Pavan Gosangi ·

    Enterprise Agentic RAG: The Dual-Database Dilemma — Synchronizing State & Vectors in Agentic RAG

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/agentic-architecture/enterprise-agentic-rag-the-dual-database-dilemma-synchronizing-state-vectors-in-agentic-rag-8c164c15e16e?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/…

  385. dev.to — Anthropic tag TIER_1 English(EN) · Anshul Rajpal ·

    Claude Opus 5.5: Anthropic's Latest Model for Agentic Coding

    <blockquote> <p><strong>Key Takeaways</strong></p> <ul> <li>Claude Opus 5.5 launched September 22, 2026</li> <li>Anthropic's first model in the new Claude 5.5 family</li> <li>40% cheaper to run than Opus 5</li> <li>Optimized for agentic coding and knowledge work</li> <li>Availabl…

  386. dev.to — MCP tag TIER_1 English(EN) · qianqiuwanzi ·

    One Memory Layer, Thirteen Agents: How We Share Context Across a Multi-Agent Team

    <p>Running one AI agent is easy. Running thirteen that actually cooperate is where memory becomes the real bottleneck.</p> <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A…

  387. dev.to — LLM tag TIER_1 English(EN) · Eryk Kubiak ·

    Local Coding Agent: Weak Model? No, Bad Plumbing

    <p> </p> <p><strong>In this video:</strong></p> <p>0:00 Why agents fail around turn ten<br /> 0:19 It runs but is useless<br /> 1:54 The base-URL swap<br /> 3:26 Where compatibility breaks<br /> 5:28 From tokens to a tool call<br /> 8:31 Context, speed and turn count<br /> 12:25 …

  388. dev.to — LLM tag TIER_1 English(EN) · Priyesh Dave ·

    How Context Window Degradation Breaks Long-Running Agents in Production—and How to Engineer Your Way Out

    <p>Liquid syntax error: Unknown tag 'endraw'</p>

  389. r/LocalLLaMA TIER_1 English(EN) · /u/Low_Bad_6585 ·

    Running an LLM-driven town with 800+ persistent agents: concurrency, context caching, and inference costs

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1x0ms0m/running_an_llmdriven_town_with_800_persistent/"> <img alt="Running an LLM-driven town with 800+ persistent agents: concurrency, context caching, and inference costs" src="https://preview.redd.it/71a7tj…

  390. dev.to — LLM tag TIER_1 English(EN) · HiDevs ·

    Building a Multi-Agent Load Simulator: 50 Agents, One Collection

    <p>As AI applications evolve from simple chatbots into complex multi-agent architectures, the demands placed on the vector database change fundamentally. In a single-agent loop, context retrieval is predictable and mostly sequential: an agent sends a search request, waits for a r…

  391. dev.to — LLM tag TIER_1 English(EN) · Ameer Mavia ·

    How Multi-Agent LLM Orchestration Supports Complex AI Workflows

    <p><a href="https://nextigent.ai/services/model-orchestration/" rel="noopener noreferrer">Multi-agent LLM orchestration</a> coordinates multiple AI agents, language models, tools, and data sources so they can contribute to a shared workflow. Instead of expecting one model to inte…

  392. dev.to — LLM tag TIER_1 English(EN) · mech.app ·

    OpenAI and Ironclad: Turning Contract Workflows Into Agent Evals

    <p>OpenAI and Ironclad just published a case study on using production contract workflows as both training data and evaluation benchmarks for computer-use agents. This is not a demo. It is a partnership where a SaaS company opens its workflow engine to become an agent training gr…

  393. dev.to — LLM tag TIER_1 English(EN) · Pratik ·

    How to Build Resilient AI Agents with Search Fallback Loops

    <p>Building autonomous AI agents is incredibly rewarding until you deploy them to production and real-world data breaks your clean pipelines.<br /> A common bottleneck is the tool execution layer. When your agent invokes a vector DB search or a live web API, it assumes it will re…

  394. dev.to — LLM tag TIER_1 English(EN) · mech.app ·

    Sarvam Arya: Production Agent Orchestration Stack Exposes State, Routing, and Recovery Plumbing

    <p>Sarvam just released Arya, an orchestration stack designed to move agent systems from prototype to production. The announcement centers on a concrete problem: frontier models can write compilers and rebuild rendering pipelines in demos, but they fail silently and inconsistentl…

  395. dev.to — LLM tag TIER_1 English(EN) · Abdulmuiz Adebayo ·

    Smallops benchmark report · MD Can Small Local Models Be Agentic? A 6-Round Benchmark of 4 Ollama Models

    <h2> A build-in-public deep dive from the smallOps project </h2> <p>Why this benchmark exists<br /> .<br /> smallOps is an experiment in giving small, locally-run language models — the kind that fit comfortably on a laptop with no GPU — the ability to act as coding agents. Tools …

  396. r/MachineLearning TIER_1 English(EN) · /u/heyitsdannyle ·

    SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]

    <table> <tr><td> <a href="https://www.reddit.com/r/MachineLearning/comments/1wyw0my/swerace_a_codingagent_benchmark_of_188_real/"> <img alt="SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]" src="https://preview.redd.it/0uopztlmp…

  397. dev.to — LLM tag TIER_1 English(EN) · mech.app ·

    Time-Travel Debugging for LLM Agents: Burr's Counterfactual Replay Architecture

    <p>Every agent trace tool shows you a waterfall of steps. You spot that step six produced garbage. Now what? You re-run the entire pipeline and hope it lands in the same place. With a non-deterministic model, it doesn't. You can never separate your change from model jitter.</p> <…

  398. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    Holo4: How H Company Built a Single Open-Weight Agent That Clicks, Codes, and Calls APIs

    <h1> Holo4: How H Company Built a Single Open-Weight Agent That Clicks, Codes, and Calls APIs </h1> <p>Most agentic AI models are specialists. A GUI-focused model can click through a browser but is lost when the task requires an API call. A tool-calling model can chain function c…

  399. dev.to — LLM tag TIER_1 English(EN) · VIBHEESHAN NK ·

    Agentic GraphRAG: From RAG to Graph-Based Agentic Reasoning

    <p>Large Language Models (LLMs) can answer many questions, but answering questions over a large and complex corpus becomes difficult when the required information is spread across multiple documents or depends on relationships between different entities.<br /> To address this cha…

  400. dev.to — LLM tag TIER_1 English(EN) · zhijie ·

    Tuning a Local Qwen Coding Agent: What Changed, What Still Fails

    <p>On eight development coding cases repeated three times, a local Qwen3.8-27B agent went from <strong>5/24 to 23/24 functional and delivered successes</strong> across sequential output-budget and reasoning rounds. That is promising development evidence. It does not establish gen…

  401. dev.to — LLM tag TIER_1 English(EN) · Rashid Mahmood ·

    Eight broken tool calls: how six agent frameworks recover

    <p>Models send bad tool calls: malformed JSON, a tool name that does not exist, a missing argument. What happens next is decided by the agent framework, not the model. In <a href="https://github.com/code-with-rashid/agentic-arena" rel="noopener noreferrer">agentic-arena</a> I mea…

  402. dev.to — LLM tag TIER_1 Nederlands(NL) · mech.app ·

    iFixAi: Independent Agent Auditing in 120 Seconds

    <p>Production agents fail in ways that runtime guardrails cannot catch. They hallucinate plausible-sounding API calls, leak context across tool invocations, and drift from their original task specification without triggering a single exception. iFixAi is a Python CLI that audits …

  403. dev.to — LLM tag TIER_1 English(EN) · Aleksei Romanov ·

    The Specialized Agent Shift: How Open Models Are Beating Frontier APIs in Production

    <p>For the past two years, the standard enterprise AI strategy looked suspiciously like a parlor trick: take an 8,000-token system prompt packed with JSON schemas, API documentation, and polite formatting rules, stuff it into the context window of a giant 400-billion-parameter fr…

  404. dev.to — LLM tag TIER_1 English(EN) · Aleksei Romanov ·

    Long-Horizon Agents: Why Multi-Turn Reasoning Breaks and the Practical Training Tricks That Fix It

    <p>Almost every modern language model looks impressive on a two-step demo. You ask it to check a database or summarize a document, it calls the right tool, formats the answer, and looks like an autonomous engineer.</p> <p>The illusion falls apart the moment you ask that same mode…

  405. dev.to — LLM tag TIER_1 English(EN) · lizer yang ·

    Agentic Retrieval: The Loop That Decides When to Search

    <blockquote> <p>Originally published at <a href="https://smartgate.network/industry/agentic-retrieval?utm_source=devto&amp;utm_medium=syndication" rel="noopener noreferrer">Agentic Retrieval: The Loop That Decides When to Search</a> on smartgate.network.</p> </blockquote> <p>A sh…

  406. dev.to — LLM tag TIER_1 English(EN) · astronaut ·

    Agent Context in 2026: The Whole Map on One Page

    <p>There are dozens of posts about <code>CLAUDE.md</code> best practices. Each one is a list of tips, and none of them explains <em>why</em>. We stopped guessing and went to the primary sources instead: Anthropic's own guidance and the 2026 papers on context engineering.</p> <p>T…

  407. dev.to — LLM tag TIER_1 English(EN) · mech.app ·

    PhantomEnvironments: How Fictional Worlds Solve the Agent Training Bottleneck

    <p>Training LLM agents with reinforcement learning hits a hard wall: you need environments that provide verifiable rewards, support long-horizon interaction, and scale without burning budget. Human-curated data is expensive. LLM-generated environments hallucinate and leak benchma…

  408. dev.to — LLM tag TIER_1 English(EN) · mech.app ·

    Whiteboard IDE: What a Canvas-Based Design Tool Reveals About Agent Context Management

    <p>Whiteboard is a YC W26-backed open-source IDE that replaces the file tree with a spatial canvas. Instead of navigating folders, you arrange components on a 2D plane. The project (422 HN points, 142 comments) exposes a different set of plumbing decisions for agent context manag…

  409. dev.to — LLM tag TIER_1 English(EN) · Akhil Belide ·

    Building a Deal Intelligence Agent with Persistent Memory

    <p>Sales conversations rarely happen in isolation.</p> <p>A customer might discuss pricing in one call, raise an objection in another, introduce a new stakeholder later, and finally ask for a specific next step several conversations after the first meeting.</p> <p>The problem is …

  410. dev.to — LLM tag TIER_1 English(EN) · Vardhini ·

    Why Stateless AI Agents Fail at Enterprise Negotiations — And How Episodic Memory Fixes It

    <p>When teams deploy large language models into enterprise workflows, the default architectural pattern connects a model to a user interface with high-level system instructions. In brief, deterministic tasks like code refactoring or single-turn customer support, this pattern work…

  411. r/LocalLLaMA TIER_1 English(EN) · /u/jonas__m ·

    Speculative reward hacking in coding agents

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wsuag0/speculative_reward_hacking_in_coding_agents/"> <img alt="Speculative reward hacking in coding agents" src="https://preview.redd.it/7pvdij77hcsh1.png?width=640&amp;crop=smart&amp;auto=webp&amp;s=0a5fbcf…

  412. dev.to — LLM tag TIER_1 English(EN) · AI Frontier Post ·

    OKF Agent Memory: give your coding agents a git-native memory that survives every session

    <p><em>Originally published at <a href="https://aifrontierpost.com/articles/okf-agent-memory-git-native-project-memory/" rel="noopener noreferrer">AI Frontier Post</a>.</em></p> <p>You know the feeling. You spend an hour with a coding agent establishing the architecture: why the …

  413. dev.to — LLM tag TIER_1 English(EN) · Manidhar Bheempadu ·

    When AI Agent Memory Learns What Not to Reuse

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frh9n3neia8zhvf4hia93.png"><img alt=" " height="380" …

  414. Mastodon — mastodon.social TIER_1 English(EN) · Moltbookpulse ·

    Concurrency, Forecasts, and Web Trust Boundaries "Multi-agent evals serialize the race condition they claim to measure" (general) + "An expectation without a se

    Concurrency, Forecasts, and Web Trust Boundaries "Multi-agent evals serialize the race condition they claim to measure" (general) + "An expectation without a settle date is not a forecast" (agents) This + more in today's Moltbook Pulse (Edition #83): https:// superagent-ebe00561.…

  415. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Building multi-agent systems in .NET? Explore a structured approach to coordinating specialized agents, managing workflows, and creating maintainable AI applica

    Building multi-agent systems in .NET? Explore a structured approach to coordinating specialized agents, managing workflows, and creating maintainable AI applications with familiar .NET patterns. # dotnet # AI https:// isaacl.dev/hbw

  416. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Docker Agent: a Tool for building, running and sharing AI Agents using a simple YAML configuration, rich tool Ecosystem and multi-agent Orchestration # AI # Age

    Docker Agent: a Tool for building, running and sharing AI Agents using a simple YAML configuration, rich tool Ecosystem and multi-agent Orchestration # AI # Agent https:// github.com/docker/docker-agent

  417. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    🤖 Long-running agents: is the bottleneck the model or the scaffolding around it? Something I keep noticing with agent setups: every individual step is easy for

    🤖 Long-running agents: is the bottleneck the model or the scaffolding around it? Something I keep noticing with agent setups: every individual step is easy for the model, but the full chain still falls apart on long tasks. I think it comes down to three things: Error compoundin..…

  418. Mastodon — mastodon.social TIER_1 English(EN) · schuler ·

    An independent test found Hindsight, a memory bank for coding agents, retained unwritten project rules across sessions without explicit instruction. The tool dr

    An independent test found Hindsight, a memory bank for coding agents, retained unwritten project rules across sessions without explicit instruction. The tool draws on git history to help agents maintain conventions. Teams should verify storage location and check for empty knowled…

  419. r/ClaudeAI TIER_2 English(EN) · /u/Disastrous_Exam9484 ·

    Repos & Dungeons: watch your agents fight their way through your codebase

    <table> <tr><td> <a href="https://www.reddit.com/r/ClaudeAI/comments/1wzt7ei/repos_dungeons_watch_your_agents_fight_their_way/"> <img alt="Repos &amp; Dungeons: watch your agents fight their way through your codebase" src="https://external-preview.redd.it/aWZqZHE3YTM4MHVoMdDq__bK…

  420. r/ClaudeAI TIER_2 English(EN) · /u/george-lin ·

    Agent Communication & Orchestration in VelaTerm

    <table> <tr><td> <a href="https://www.reddit.com/r/ClaudeAI/comments/1wzkh5s/agent_communication_orchestration_in_velaterm/"> <img alt="Agent Communication &amp; Orchestration in VelaTerm" src="https://external-preview.redd.it/dmE2aHF0bDFieXRoMapbV4K0o4qagdBawe0kyN1CzVcnkshsHO15v…