<p>Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them.</p> <p>The …
Apple Machine Learning Research
TIER_1English(EN)·
Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based sign…
Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, crons, and memory consolidation. Yet existing benchma…
arXiv:2610.10087v1 Announce Type: new Abstract: Offline MARL increasingly relies on generative policies to model multimodal joint behavior, typically by distilling a centralized teacher into decentralized one-step actors under the CTDE. We identify a failure mode at the teacher t…
arXiv:2610.09294v1 Announce Type: cross Abstract: Rapid progress in AI agents has brought growing attention to agent safety, with extensive evaluation focused on digital environments. As agents move into the physical world, embodied safety becomes increasingly important: failures…
arXiv:2610.08875v1 Announce Type: cross Abstract: Agent Skills package procedural guidance and resources for reuse, but a relevant Skill does not necessarily improve task performance. Existing studies characterize Skill content and evaluate downstream performance, yet provide lim…
arXiv cs.LG
TIER_1Deutsch(DE)·Prakhar Ganesh, Kyra Wilson, Luca Zappella, Barry-John Theobald, Nicholas Apostoloff, Lucas Monteiro Paes, Nivedha Sivakumar·
arXiv:2610.09824v1 Announce Type: new Abstract: Multi-agent systems (MAS) leverage interactions between agents to perform complex tasks. Despite their success, we show that these interactions can also lead to homogenization, i.e., agents converging to similar behaviors. Homogeniz…
arXiv:2610.09074v1 Announce Type: new Abstract: Emerging long-horizon agentic tasks require repeated model calls, worsening the inference cost of already-costly language models. While narrow agentic tasks suggest potential for aggressive model pruning without performance drop, em…
arXiv:2610.09633v1 Announce Type: cross Abstract: The evaluation of coding agents generally strives to be as realistic as possible. In our study, we collect 4,782 agent sessions of real software engineers in JetBrains IDEs, which we call Production Sessions. Since our subject is …
arXiv:2610.09684v1 Announce Type: new Abstract: Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dime…
arXiv cs.CL
TIER_1English(EN)·Giordano De Marzo, Andres L. Marin, David Garcia·
arXiv:2602.09270v2 Announce Type: replace-cross Abstract: We present a large scale data analysis of Moltbook, a Reddit-style social media platform exclusively populated by AI agents. Analyzing over 4 million posts and 19 million comments from approximately 185,000 active agents, …
Reality may establish that an outcome occurred without identifying which evolving procedure produced it or why. This distinction matters in production ML systems whose code, configuration, and artifacts change while external feedback accumulates. We examine it in a human-directed…
Offline MARL increasingly relies on generative policies to model multimodal joint behavior, typically by distilling a centralized teacher into decentralized one-step actors under the CTDE. We identify a failure mode at the teacher training stage: standard flow-based teachers pair…
arXiv:2610.07250v1 Announce Type: new Abstract: Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continu…
arXiv:2610.06964v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated strong capabilities in interactive environments, yet their ability to continually evolve from experience remains limited. Although fine-tuning enables adaptation, its dependence on…
arXiv:2610.08101v1 Announce Type: new Abstract: Shared memory coordinates agents' actions, but correct records do not establish that those actions satisfy task requirements. Memory governance and failure diagnosis regulate or inspect recorded information; they do not by themselve…
arXiv:2610.07979v1 Announce Type: new Abstract: As agents continuously improve by generating and revising Skills, the process that discovers and refines those Skills becomes a learnable object in its own right. Task-Skills directly act on task execution, whereas Meta-Skills gover…
arXiv cs.AI
TIER_1English(EN)·Yeji Park, Jaeyun Shim, Taesik Gong·
arXiv:2610.07972v1 Announce Type: new Abstract: Mobile GUI agents increasingly operate on interfaces influenced by users' histories and preferences, but their reliability across different users remains underexplored. We introduce PAIR (Personalized Application-state Instantiation…
arXiv:2610.07860v1 Announce Type: new Abstract: Multi-agent systems are increasingly deployed for complex knowledge work, yet their orchestration layers remain largely memoryless: each new task is decomposed, assigned, and executed from scratch with no benefit from prior successf…
arXiv:2610.07787v1 Announce Type: new Abstract: Multi-agent systems (MAS) powered by large language models have shown strong performance across code generation, mathematical reasoning, and question answering. However, existing methods for automating MAS design mostly operate at t…
arXiv:2610.07763v1 Announce Type: new Abstract: The rapid progress of LLM-based multi-agent systems (MAS) has shown that they largely outperform single agents on coding, math, and QA tasks, where executable tests provide a binary success signal. Whether this advantage transfers t…
arXiv:2610.07675v1 Announce Type: new Abstract: AI agents are entering production in increasingly consequential environments without a shared semantic standard for what their evaluations actually mean. Scores, traces, judge outputs, and multi juror findings are increasingly used …
arXiv:2610.07578v1 Announce Type: new Abstract: In cooperative Multi-Agent Reinforcement Learning (MARL), agents are often trained under concurrent participation, while in many tasks some agents act earlier and leave task-relevant information that becomes useful to agents partici…
arXiv:2610.07556v1 Announce Type: new Abstract: Learned orchestration can automatically construct effective language-model multi-agent systems, but existing approaches couple planning to fixed worker pools and train decomposition and collaboration from the same terminal outcome, …
arXiv:2610.07274v1 Announce Type: new Abstract: Deterministic benchmark scores show that an agent received credit, but not whether that credit was earned, reported honestly, or would hold on a second run. We introduce a Trust Layer for Agent Evaluation, an additive post-hoc frame…
arXiv:2610.07261v1 Announce Type: new Abstract: A team of coding agents can look fine agent by agent yet fail as a team: each passes its own tests while the merged result is broken, and single-agent evaluation never catches it. As teams run several LLM coding agents in parallel o…
arXiv:2610.07257v1 Announce Type: new Abstract: Developers increasingly run a fleet of coding agents side by side on one workstation. The tools they reach for, terminal multiplexers like tmux and a new generation of agent managers, were built to arrange windows, not to govern mem…
arXiv:2610.07004v1 Announce Type: new Abstract: Task planning for LLM agents requires workflows that satisfy both user intent and complex sub-task dependencies. While existing planners work well for sequential or directed acyclic graph (DAG)-like structures, they struggle with wo…
arXiv:2609.34060v2 Announce Type: replace Abstract: Prompt-specialized multi-agent systems enable multiple agents to share a model while performing complementary roles to solve complex tasks. However, agent-specific prefixes change the KV cache generated for the same shared conte…
arXiv cs.LG
TIER_1English(EN)·Jia Liufu, Bin Hu, Linglin Jing, Terry Kong, Yuki Huang, Ashwath Aithal, Wenming Yang, Jun Yang·
arXiv:2610.07898v1 Announce Type: new Abstract: Repository-level software engineering (SWE) is a challenging long-horizon setting: agents must reason over extended interactions, use tools, and adapt to stateful environments. Recent work trains SWE agents with reinforcement learni…
arXiv cs.LG
TIER_1English(EN)·Lan Shi, Daigo Shishika, Xuan Wang·
arXiv:2610.07475v1 Announce Type: new Abstract: Adapting a reinforcement learning policy to changes in another agent's behavior typically requires a large amount of new interaction data. Policy sensitivity provides a first-order prediction of how a locally optimal policy changes …
arXiv cs.LG
TIER_1English(EN)·Ziyang Cai, Christos Ziakas, Vasilis Kontonis, Tim Pearce, Siddhartha Sen, Akshay Krishnamurthy, Shivam Garg, Dimitris Papailiopoulos·
arXiv:2610.07447v1 Announce Type: new Abstract: Autoresearch agents tackle open-ended problems by repeatedly proposing candidate solutions, evaluating them, and using feedback to guide subsequent experiments. We show that independent runs of the same agent on the same task often …
arXiv:2610.06910v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have demonstrated remarkable capabilities in web front-end execution, with browser-based game generation emerging as a particularly prominent frontier. While previous efforts frequentl…
arXiv cs.CL
TIER_1English(EN)·Shivani Kumar, Adarsh Bharathwaj, David Jurgens·
arXiv:2604.20658v2 Announce Type: replace Abstract: Multi-agent systems built from teams of large language models (LLMs) are increasingly deployed for collaborative scientific reasoning and problem-solving. These systems require agents to coordinate under shared constraints, such…
arXiv:2610.07327v1 Announce Type: cross Abstract: Agent tasks require sequences of interdependent decisions. Autoregressive models support more flexible decision interfaces than conventional classifiers but incur the latency of token-by-token generation. Recent shared-prefix meth…
arXiv cs.CL
TIER_1English(EN)·Lasse B. Strand, Robert Jakob, Kevin O'Sullivan, Markus Kreft·
arXiv:2610.08452v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacti…
arXiv cs.AI
TIER_1Norsk(NO)·Justin Chih-Yao Chen, Elias Stengel-Eskin, Yan Chen, Pol Llado, Scott Counts, Mohit Bansal, Benjamin Van Durme, Harsh Jhamtani, Gaurav Verma·
arXiv:2610.05437v2 Announce Type: replace Abstract: Computer-use agents need to capture procedural knowledge of how people use software. User telemetry offers a scalable source of this knowledge. However, learning reusable skills from these logs requires addressing three challeng…
arXiv:2610.05281v2 Announce Type: replace Abstract: Modern agents increasingly ground their reasoning in observations returned by tools, such as file contents read from a workspace. However, the data sources underlying these observations may later be modified by users, other agen…
arXiv:2609.33678v2 Announce Type: replace Abstract: We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design docu…
arXiv:2609.26911v2 Announce Type: replace Abstract: A single locally plausible tool call can derail an otherwise successful agent trajectory. Suspicion alone does not justify intervention, because the replacement itself can introduce the very failure verification is meant to prev…
arXiv:2610.08773v1 Announce Type: cross Abstract: Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the …
arXiv:2610.08155v1 Announce Type: cross Abstract: Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by enabling collaborative problem solving among specialized agents. However, existing …
arXiv:2610.08097v1 Announce Type: cross Abstract: Mathematical problem solving often requires deterministic computational steps that agents delegate to tools and implicitly trust. Yet tools can fail silently, returning plausible but incorrect results. How well can agents detect a…
arXiv:2610.07753v1 Announce Type: cross Abstract: Tool-using agents make consequential changes to external state, yet correct outcomes do not guarantee that their actions were supported by evidence established beforehand. We study where this evidence-to-action chain breaks as age…
arXiv:2610.07557v1 Announce Type: cross Abstract: Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing co…
arXiv cs.AI
TIER_1English(EN)·Celal Ziftci, Spencer Greene, Ray Liu, Livio Dalloro, Lorenzo Dini·
arXiv:2610.07289v1 Announce Type: cross Abstract: Manual repair of program failures is time-consuming and disruptive for software developers, particularly during the pre-submit phase where test failures occur within continuous integration systems. While Automated Program Repair h…
arXiv cs.AI
TIER_1English(EN)·Tao Long, Lydia B. Chilton·
arXiv:2610.07204v1 Announce Type: cross Abstract: Recent AI alignment work often frames alignment as a pre-deployment optimization problem: collect human feedback, learn preferences or principles, finetune the model, and deploy an aligned system. This framing has produced major p…
arXiv:2610.08720v1 Announce Type: new Abstract: LLM-based agents are increasingly advancing scientific and engineering problem solving, with physics simulation emerging as a challenging yet practical testbed for reproducing complex physical phenomena with application in embodied …
arXiv:2610.08691v1 Announce Type: new Abstract: Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both th…
arXiv:2610.08662v1 Announce Type: new Abstract: As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a …
arXiv:2610.08514v1 Announce Type: new Abstract: Execution feedback lets coding agents revise programs and learn from their own corrections. A correction's learning weight should reflect both the transitions supported by its executions and the amount of evidence behind that suppor…
arXiv cs.AI
TIER_1English(EN)·Hyun Jung Lee, Jungtaek Kim, Jongwon Jeong, Tae-Eui Kam, Donghyun Kim, Yong Jae Lee·
arXiv:2610.08432v1 Announce Type: new Abstract: Improving embodied agents often focuses on optimizing the underlying model through training, while the surrounding agent harness that controls planning, context, and tool use is typically engineered. We ask whether this harness can …
arXiv:2610.08138v1 Announce Type: new Abstract: Legal intelligence aims to support reliable decision-making across long-horizon legal processes involving evolving case states and multiple roles. However, real-world legal deployment exhibits substantial case heterogeneity in facts…
arXiv:2610.08106v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) advance rapidly, while conventional benchmark development lags behind, delaying investigation of newly observed capability gaps. Such investigation requires an expressive task format and an o…
Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-relevant information through interaction when it is ab…
We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fau…
Agentic systems use large language models (LLMs) to carry out concrete tasks. Prior work often borrows abstractions such as scheduling, caching or isolation piecemeal from operating systems, so the mechanisms it builds share little common ground, and the shared view of the two fo…
Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--…
Supervisory governors can interfere with the tool-using agents they regulate. We study this possibility in a controlled file-recovery environment where increases in regulatory intensity trigger experimentally imposed tool failures. A cost-blind governor can turn these failures in…
When agents share a reward for completed tasks, reporting unsafe work can reduce the reporter's reward by stopping a task. Audits can make reporting optimal without ensuring that further training teaches a silent team to report. We study this learning problem in a game where any …
Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task r…
Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacting choices, from chunking and embedding model to…
Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by enabling collaborative problem solving among specialized agents. However, existing MAS frameworks tightly couple task reasoning with …
Shared memory coordinates agents' actions, but correct records do not establish that those actions satisfy task requirements. Memory governance and failure diagnosis regulate or inspect recorded information; they do not by themselves establish whether it is sufficient to judge ta…
Recent advances in agentic search have given large language model (LLM) agents finer control over corpus exploration. However, search interfaces often return matching passages even when feedback about the candidate set would suffice for the next decision, coupling candidate refin…
Multi-agent systems (MAS) powered by large language models have shown strong performance across code generation, mathematical reasoning, and question answering. However, existing methods for automating MAS design mostly operate at the task level, producing a single fixed workflow…
AI agents are entering production in increasingly consequential environments without a shared semantic standard for what their evaluations actually mean. Scores, traces, judge outputs, and multi juror findings are increasingly used to justify readiness and release decisions, yet …
Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch…
Tool-using agents make consequential changes to external state, yet correct outcomes do not guarantee that their actions were supported by evidence established beforehand. We study where this evidence-to-action chain breaks as agents move from deciding whether to act to executing…
arXiv cs.MA (Multiagent)
TIER_1English(EN)·Lydia B. Chilton·
Recent AI alignment work often frames alignment as a pre-deployment optimization problem: collect human feedback, learn preferences or principles, finetune the model, and deploy an aligned system. This framing has produced major progress, but it under-specifies what happens once …
We present T-Search, an open-weight agentic retriever for hard multi-step search. Given a question and a search tool over a fixed corpus, it runs a bounded multi-round search and returns a ranked list of evidence chunks with short justifications, leaving answer generation to a do…
arXiv:2610.02588v1 Announce Type: new Abstract: Agents are increasingly given open-ended research tasks: discovering an empirical law from self-designed experiments, improving a heuristic whose optimum nobody knows, or beating a standing record. Their execution logs record every …
arXiv cs.AI
TIER_1English(EN)·Eray Turkel, Mengsha Sun, Kartik Ayyar, Sean Dunigan, Jack Lu, Vlad Shcherban, Hsiang-Shun Shih, Xin Wang, Tiantian Zhang·
arXiv:2610.02563v1 Announce Type: cross Abstract: We present OpenGameEval, a benchmark and evaluation framework for agentic game development inside Roblox Studio. It runs language models as agents in reproducible, stateful game-engine sessions and scores each run with executable …
arXiv cs.AI
TIER_1English(EN)·Songtao Wei, Yi Li, Zhichun Guo, Bingzhe Li·
arXiv:2610.02396v1 Announce Type: cross Abstract: Multi-agent systems (MAS) built from large language models coordinate specialized agents to tackle complex tasks, but effective workflows are difficult to design in advance. Test-time evolution refines workflows using execution fe…
arXiv cs.AI
TIER_1Dansk(DA)·A. Said Gurbuz, Ahmed Nassar, Sunghwan Hong, Marc Pollefeys, Peter W. J. Staar·
arXiv:2610.02320v1 Announce Type: cross Abstract: Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such s…
arXiv cs.AI
TIER_1English(EN)·Yu Li, Guangfeng Cai, Long-Fei Li, Shuo Han, Shengtian Yang, Han Luo, Kaibing Yang, Lei Feng·
arXiv:2610.03634v1 Announce Type: new Abstract: Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier command…
arXiv:2610.03564v1 Announce Type: new Abstract: Financial AI agents must do more than retrieve facts: investment workflows require correct quantitative execution, reliable use of procedural resources, and auditable structured outputs. We introduce FinSkillBench, an evaluation sui…
arXiv cs.AI
TIER_1English(EN)·Chiara Troiani, Arash Salarian, Majed El Helou, Benjamin Ryder, Jean Diaconu, Herv\'e Muyal, Marcelo Yannuzzi·
arXiv:2610.03213v1 Announce Type: new Abstract: Tool-equipped AI agents use tool calls to access data and act on external systems. Horizontal growth of agentic systems increases the number of these interactions, and further motivates the need for automated, per-call oversight tha…
arXiv cs.AI
TIER_1English(EN)·Yulong Ming, Jie Xu, Zihan Wu, Xiaohua Jia·
arXiv:2610.02932v1 Announce Type: new Abstract: Compiling GUI procedures that agents execute repeatedly into programs can reduce their token costs. However, measuring payback and deciding when to compile have two challenges. First, compilation costs are uncertain because attempts…
arXiv:2610.02664v1 Announce Type: new Abstract: Long-horizon agents are now playing an increasingly significant role in assisting humans with complex problem-solving. However, it is exactly their extended interaction history that introduces an underexplored execution-safety conce…
arXiv:2610.02525v1 Announce Type: new Abstract: Long-horizon research agents must decide both how to investigate and what to investigate next as evidence accumulates. This is hard to learn because such decisions are sparse in long execution traces, and their consequences may emer…
arXiv:2610.02330v1 Announce Type: new Abstract: Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provid…
arXiv cs.AI
TIER_1English(EN)·Xi Qin, Isabel Kurth, Xin Cui, Elin Park, Alexander Schaefer, Yaad Oren·
arXiv:2610.02405v1 Announce Type: new Abstract: Using a frontier model like Claude Opus as a meta-agent to generate terminal tasks and verifiers for RL training is increasingly common. Yet a runnable Docker image and executable test suite do not guarantee a faithful end-to-end pi…
arXiv:2610.03372v1 Announce Type: new Abstract: Training long-horizon agents to solve complex tasks requires effective supervision over extended interaction sequences. However, sparse terminal rewards obscure intermediate contributions, while on-policy distillation can lose infor…
arXiv:2610.02554v1 Announce Type: new Abstract: Offline multi-agent reinforcement learning (MARL) faces a persistent trade-off. Expressive generative policies can represent multi-modal coordination in the data, but cannot distinguish high-value regions, while value-optimized poli…
arXiv:2610.02363v1 Announce Type: new Abstract: Benchmarks for agent-generated data work grade a pipeline by running it once against a fixed snapshot. ArrivalBench instead re-executes the pipeline an agent leaves behind under adversarial but replayable delivery schedules (late, d…
arXiv:2604.10800v2 Announce Type: replace-cross Abstract: Learned classifiers deployed in agentic pipelines face a fundamental reliability problem: predictions are probabilistic inferences, not verified conclusions, and acting on them without grounding in observable evidence lead…
arXiv:2601.22352v2 Announce Type: replace-cross Abstract: Language model agents often appear capable of self-recovery after failing tool call executions, yet this behavior lacks a formal explanation. We present a predictive theory that resolves this gap by showing that recoverabi…
arXiv cs.AI
TIER_1English(EN)·Yaxiao Liu (PwC China AI Center), Pengbo Liu (PwC China AI Center), Yiwen Liu (PwC China AI Center), Yihua Guan (PwC China AI Center), Jiaxing Song (Tsinghua University)·
arXiv:2609.35149v2 Announce Type: replace Abstract: Deploying, migrating, or scaling an agent can change its model, harness, infrastructure, application, and intended users. We formulate agent calibration as standards-first adaptation: define basic-capability, technical-environme…
arXiv:2610.03153v1 Announce Type: cross Abstract: Workspace agents combine large language models with execution harnesses to perform stateful, multi-step tasks that access or modify external resources. Existing benchmarks leave gaps in executable coverage of their runtime securit…
arXiv:2610.02951v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures scale language models efficiently by activating only a few experts per token, but the saving is confined to computation: every expert must stay resident on the accelerator, so memory bounds w…
arXiv cs.AI
TIER_1English(EN)·Mayank Rathee, Alexander Stepanov, Shalin Madabhavi, Jinhao Zhu, Raluca Ada Popa, Ion Stoica·
arXiv:2610.02569v1 Announce Type: cross Abstract: Coding agents have become increasingly long-horizon, autonomous, reliant on general-purpose shell and maintain their own persistent memory for self-improvement. While these capabilities have made the agents powerful, they have als…
Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continually construct and revise prompts, thereby elici…
Frameworks that use large language models for scientific discovery typically rely on a fixed, human-designed algorithm that decides what the model sees at each step, leaving the model only the role of proposer. The model knows nothing of the search beyond what it is shown. As mod…
Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates se…
Robot agents must connect their intended actions to observed outcomes while retaining the context needed to revise their choices over repeated attempts. Existing interfaces often leave these choices inside predefined tools or require agents to manage detailed execution code and i…
A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-context adaptation agents…
As agents rapidly evolve, existing benchmarks can become saturated, limiting their ability to distinguish capabilities and reveal remaining failure modes. Particularly in scientific domains, constructing and updating benchmarks requires substantial time, labor, and domain experti…
Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates se…
Generating a high-quality gaming world from a natural-language game intent requires joint reasoning about scene structure, spatial layout, gameplay objectives, interactive entities, and executable gameplay logic. Existing coding agents can generate individual assets, scenes, or s…
Multi-agent debate can improve reasoning, yet often fails to beat simple majority voting. We argue that successful debate requires two distinct mechanisms: proposal supply must surface a correct answer, and readout must identify that answer when voting misses it. We formalize the…
arXiv cs.LG
TIER_1English(EN)·Jose A. Ayala-Romero, Andres Garcia-Saavedra, Xavier Costa-Perez·
arXiv:2610.01887v1 Announce Type: cross Abstract: Dynamic resource assignment, the real-time allocation of task streams to heterogeneous processing nodes, is the backbone of modern computing infrastructure. While learning-based schedulers excel in research, industrial deployments…
arXiv:2610.00704v1 Announce Type: new Abstract: Natural-language skills are textual procedural memories through which large language model (LLM) agents retain reusable task knowledge without updating model weights. Existing methods typically treat skills as either static artifact…
LLM-based agents are increasingly capable of generating complex 3D structures, with the potential to reshape how objects are designed and realized in the physical world. Yet, producing elegant geometry is fundamentally different from producing objects that can be built and perfor…
Large language models (LLMs) are increasingly used in software engineering, including agentic systems that coordinate multiple agents, but impose higher computational and environmental costs. In this paper, we present a comprehensive empirical study of agentic LLM systems across …
Mixture-of-Experts (MoE) architectures scale language models efficiently by activating only a few experts per token, but the saving is confined to computation: every expert must stay resident on the accelerator, so memory bounds where these models can be deployed. Expert pruning …
arXiv:2610.00961v1 Announce Type: new Abstract: As code generation is increasingly delegated to AI systems, the bottleneck is shifting from writing code to supervising the systems that write it --- a shift CS-education researchers have begun to name. This shift exposes a vocabula…
arXiv cs.AI
TIER_1English(EN)·Yezhou Cheng, Runjia Du, Zeming Liu, Hang Lyu, Zehua Yang, Bojun Lin·
arXiv:2610.00282v1 Announce Type: new Abstract: How should an embodied agent respond when a person's correction may be wrong? We formulate grounded correction arbitration as a choice among accepting, rejecting, inspecting the world, and asking the speaker. GAVA implements this in…
arXiv:2610.00651v1 Announce Type: new Abstract: Agent evaluations are increasingly used to compare LLMs and inform deployment decisions, yet ranks can reflect not only the model but also the effects of the evaluation conditions such as the scaffolds or tasks. This makes reliabili…
arXiv cs.AI
TIER_1English(EN)·Andre Fu, Malik Drabla, Leon Liu, Meji Abidoye, Marek Suppa, Lata Mishra, Adnan El Assadi, Yiyuan Li·
arXiv:2610.00648v1 Announce Type: new Abstract: AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to their ability to execute on production incident response. This …
arXiv cs.AI
TIER_1English(EN)·Sahan Paliskara, Nattaput Namchittai, Andrew Lampinen·
arXiv:2610.00583v1 Announce Type: new Abstract: People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people's agents over shared resources such as a codebase, a calendar, or a budget. When each agent acts for a different user…
arXiv cs.AI
TIER_1English(EN)·Haoyang Su, Weiran Huang·
arXiv:2610.00437v1 Announce Type: new Abstract: LLM agents generate intermediate reasoning and actions token by token, making extended interactions slow and computationally expensive. Jev-style models offer fast probabilistic predictions over finite fields, but require those fiel…
arXiv:2609.39777v1 Announce Type: new Abstract: LLM-based multi-agent systems coordinate specialized reasoning through aggregation, interaction, and adaptive control, yet their potential for graph learning remains unexplored. Graph learning is a natural setting for such systems b…
arXiv:2609.39547v1 Announce Type: new Abstract: GUI agents built on large language and vision-language models still struggle on unseen applications and complex multi-step tasks, as completing real GUI tasks depends on app-specific, temporally volatile operational knowledge that i…
arXiv:2604.20779v2 Announce Type: replace Abstract: AI coding agents are being adopted at scale, yet we lack empirical evidence on how people actually use them and how much of their output is useful in practice. We present SWE-chat, the first large-scale dataset of real coding ag…
arXiv:2604.12102v3 Announce Type: replace Abstract: We describe compute-grounded reasoning (CGR), a design pattern in which code computes selected sub-problems from explicit intermediate representations before a language model answers. Spatial Atlas implements CGR as an Agent2Age…
arXiv:2601.11354v2 Announce Type: replace Abstract: Recent LLM-for-Space systems address mission planning, scheduling, operations support, simulator control, and autonomy, but their evaluations use different task contracts, control settings, simulators, and success criteria. We i…
arXiv:2508.08882v5 Announce Type: replace Abstract: Recent advances in multi-agent systems highlight the potential of specialized small agents that collaborate via division of labor. Existing tool-integrated reasoning systems, however, often follow a single-agent paradigm in whic…
arXiv:2610.02204v1 Announce Type: cross Abstract: Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a …
arXiv cs.AI
TIER_1English(EN)·Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma·
arXiv:2610.02122v1 Announce Type: cross Abstract: Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, …
arXiv:2610.01756v1 Announce Type: cross Abstract: Decentralized agent economies increasingly build a single task from protocols that were designed and secured separately. This creates a simple problem: a workflow can look correct at each step and still produce the wrong outcome. …
arXiv:2610.00890v1 Announce Type: cross Abstract: Coding agents often fail in the last mile: they build most of a feature but drop a requirement, test only the cases their implementation already handles, break behavior that was supposed to stay intact, or validate against an unch…
arXiv:2610.00650v1 Announce Type: cross Abstract: The performance of AI coding agents is highly dependent on their underlying coding rules. However, existing coding rules are typically hand-crafted and fixed, making the process labor-intensive and often suboptimal. In this work, …
arXiv:2610.00557v1 Announce Type: cross Abstract: Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks. Reinforcement learning (RL) and large language models (LLMs) offer complementary mechanisms for t…
arXiv:2610.02036v1 Announce Type: new Abstract: AI agents can each make locally valid decisions yet jointly produce an invalid result. We call this the global coherence problem: a failure of shared state, not merely of model intelligence. Our Observation-Aliasing Impossibility Th…
arXiv:2610.02001v1 Announce Type: new Abstract: Small open-weight models (2-9B) run on ordinary laptops, but under cloud-scale agent harnesses they rarely complete real tasks: tool prefill overflows the context, self-correction diverges, tool demonstrations loop, and tasks are si…
arXiv cs.AI
TIER_1English(EN)·Beining Wu, Zihao Ding, Jun Huang·
arXiv:2610.01787v1 Announce Type: new Abstract: Self-improving GUI agents keep the trajectories they produce and return them to the agent, by fine-tuning or by retrieval into the prompt, and studies that compare the two destinations disagree. We attribute this to the unit of expe…
arXiv:2610.01618v1 Announce Type: new Abstract: Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users d…
arXiv:2610.01506v1 Announce Type: new Abstract: As agents evolve from single-tool systems into modular, composite architectures, skills are becoming an important mechanism for capability development and distribution. However, the academic community lacks a structured framework fo…
arXiv:2610.01249v1 Announce Type: new Abstract: Conventional reasoning protocols present a fixed, preselected task, so they cannot test whether an agent propagates relevant updates, preserves unaffected work, or reconstructs a historical task binding. We therefore study \emph{dyn…
arXiv:2610.01097v1 Announce Type: new Abstract: End-to-end research agents can now produce complete scientific papers, yet manuscript claims often diverge from executed experiments. This gap is structural: research state, failure histories, and claim-evidence alignment are not ma…
arXiv cs.AI
TIER_1English(EN)·Jingtan Wang, Sirajul Salekin, Young mok Jung, Javier Movellan, Bryan Kian Hsiang Low, Manjot Bilkhu·
arXiv:2610.00979v1 Announce Type: new Abstract: Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environ…
arXiv:2610.00972v1 Announce Type: new Abstract: As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to referenc…
arXiv:2610.00084v1 Announce Type: new Abstract: Detailed profession-specific system prompts raise token use and estimated cost per response without a consistent accuracy gain. We evaluate Scientific Agents, an open-source corpus of 503 profession-specific AGENTS.md profiles, with…
arXiv cs.AI
TIER_1English(EN)·Ronghua Li, Zi Liang, Zhishan Li, Shinan Liu·
arXiv:2610.00949v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) on offline agent trajectories is the standard approach for training specialized tool-using agents, but forcing models to imitate reasoning and actions token by token may harm other capabilities (e.g., ge…
arXiv:2610.00710v1 Announce Type: new Abstract: As large language model (LLM) agents become widely adopted, they are increasingly deployed for tasks that require persistent monitoring or recurring actions (e.g., market analysis). These agents are expected to operate unattended fo…
The prevailing approach to computer-use agents couples a model with a domain-specific harness: a browser or desktop environment equipped with human engineered tools that are fixed before task execution. As models' coding capabilities improve, the GUI native and static harness pre…
Dynamic resource assignment, the real-time allocation of task streams to heterogeneous processing nodes, is the backbone of modern computing infrastructure. While learning-based schedulers excel in research, industrial deployments still rely on hand-written rules that operators c…
Dynamic resource assignment, the real-time allocation of task streams to heterogeneous processing nodes, is the backbone of modern computing infrastructure. While learning-based schedulers excel in research, industrial deployments still rely on hand-written rules that operators c…
Decentralized agent economies increasingly build a single task from protocols that were designed and secured separately. This creates a simple problem: a workflow can look correct at each step and still produce the wrong outcome. For example, a correct escrow may release payment …
arXiv:2609.39022v1 Announce Type: cross Abstract: Coding agents need to establish that a program satisfies a specification and that the specification captures the requested behavior. We study how expert diagnosis of verification failures can become reusable guidance for this work…
arXiv:2609.38288v1 Announce Type: new Abstract: We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, whi…
arXiv:2609.38445v1 Announce Type: new Abstract: Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven search and identify three core challenges: organizing evolving research ideas, sele…
arXiv:2609.38460v1 Announce Type: new Abstract: Language agents must revise planned actions when evidence changes, permission is revoked, or a stop instruction arrives. A useful response is selective: suspend affected actions, preserve unaffected work, and resume only after suffi…
arXiv cs.AI
TIER_1English(EN)·Jeffrey Willette, Krishna C. Puvvada, Boris Ginsburg·
arXiv:2609.38712v1 Announce Type: new Abstract: Long-horizon agentic workflows require models to sustain repeated state-dependent actions all while the context grows, sub-task complexity changes, and new data arrives. Each situation represents an independent axis along which an a…
arXiv cs.AI
TIER_1English(EN)·Qiuhui Chen, Jiafan Lu, Shuaimin Tang, Tao Dai, Suyuan Wang, Chenrui Ji, Zhenglei Zhou, Weimin Zhong·
arXiv:2609.38766v1 Announce Type: new Abstract: Scientific agents can retrieve relevant passages yet still lose functional order, mix evidence across sources, or state conclusions that exceed the retrieved record. We introduce PathAnchor, a bounded scientific reasoning system bui…
arXiv:2609.38881v1 Announce Type: new Abstract: Real-time strategy (RTS) games require agents to coordinate economic development, production and construction, base defense, unit organization, and attack timing over long matches. Existing studies have applied large language models…
arXiv:2609.38891v1 Announce Type: new Abstract: Long-horizon agentic tasks demand strong reasoning and efficient execution across successive interactions with dynamic environments. A common approach decouples high-level planning from low-level execution through separate planner a…
arXiv:2609.38912v1 Announce Type: new Abstract: Agent harnesses govern how large language models (LLMs) gather context, invoke tools, verify results, preserve state, and terminate, largely affecting agent performance. However, the value of each harness mechanism can differ across…
arXiv:2609.39140v1 Announce Type: new Abstract: Learning to complete tasks in unfamiliar environments with unknown rules remains a key challenge for LLM agents. Current LLM agents often record their discoveries in prose, which may not provide a compact, explicit account of how th…
arXiv:2609.39325v1 Announce Type: new Abstract: The ability of Large Language Model (LLM) agents to complete daily and professional work is receiving increasing attention. Training such agents requires realistic work scenarios. Expert-authored occupational work is costly and slow…
arXiv cs.AI
TIER_1English(EN)·Seonho Lee, Wonryeol Jeong, Alberto Cereser, Inha Kang, Hyeonjong Kim, Seungmin Kwak, Dongmin Park·
arXiv:2609.39564v1 Announce Type: new Abstract: Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form…
arXiv:2609.40027v1 Announce Type: new Abstract: Causal action verifiers gate an agent's state-changing tool calls by checking whether each proposed intervention is identifiable against a committed action-state graph, and they issue a certificate that carries the identification ar…
arXiv cs.AI
TIER_1English(EN)·Yinghui He, Yapei Chang, Khushi Bhardwaj, Daniele Molinari, Tugrul Konuk, Jan Kautz, Ali Hatamizadeh·
arXiv:2609.40285v1 Announce Type: new Abstract: On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the…
arXiv:2609.38345v1 Announce Type: cross Abstract: Multi-agent coding systems are designed to tackle complex software engineering tasks through collaboration. However, existing evaluations typically assume configured organizations are followed faithfully, whereas reality differs. …
arXiv:2609.38662v1 Announce Type: cross Abstract: Recursive self-improvement (RSI) lets a system improve from its own outcomes; in LLM-based multi-agent systems, Agents refine one another within a task, and outcomes improve how they collaborate across tasks. However, existing mul…
arXiv:2609.38822v1 Announce Type: cross Abstract: Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The stan…
arXiv cs.AI
TIER_1English(EN)·Shijia Ge, Alex Zhou, Jianshu Zeng, Yexing Wan, Di Wu, Zelin Zheng, Yazhe Wang, Zhiqi Jia, Xuan Shangguan, Jay Zhu, Yijun Liu, Lingyu He, Sihang Wu, Xiao He, Hongcheng Gao·
arXiv:2609.39018v1 Announce Type: cross Abstract: Frontier models can control robots, but reasoning through every reach, grasp, and retreat makes manipulation slow and token-intensive. We revisit code as policy with a different division of labor: models build executable tools, co…
arXiv:2609.39081v1 Announce Type: cross Abstract: We spent five weeks using an LLM coding agent on open problems in coding theory: finding large sets of four-letter words, such as DNA barcodes, that stay far apart in edit distance. The agent wrote the verifiers and search code; a…
arXiv:2609.39102v1 Announce Type: cross Abstract: Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer …
arXiv:2609.39154v1 Announce Type: cross Abstract: Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting bec…
arXiv:2609.39333v1 Announce Type: cross Abstract: Autonomous AI agents can turn authors' goals into interactive narratives by independently organizing and carrying out generation and revision. As agents generate and revise extensive content, authors struggle to grasp its overall …
arXiv:2609.39909v1 Announce Type: cross Abstract: We introduce DoGBENCH (Documentation Generation Benchmark), to our knowledge, the first benchmark for generating and maintaining real user-facing software documentation. It asks whether an agent can produce documentation that expe…
arXiv:2609.39912v1 Announce Type: cross Abstract: Parallel search may generate a correct answer that final-answer voting fails to select. We formulate this consolidation stage as trajectory selection and introduce TRACE (Trajectory Ranking with Aggregated Cross-Rollout Evidence),…
arXiv:2609.39957v1 Announce Type: cross Abstract: Coding agents solve repository-level tasks through sequences of actions, where a single erroneous action can misdirect subsequent decisions and increase recovery costs. Existing approaches use execution feedback for recovery or sp…
arXiv:2609.40230v1 Announce Type: cross Abstract: Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumpt…
arXiv cs.AI
TIER_1English(EN)·Yong Du, Tongbo Chen, Zhengxi Lu, Yizhou Liu, Bofan Chen, Tao Jiang, Wenhao Xu, Yongliang Shen·
arXiv:2609.40253v1 Announce Type: cross Abstract: Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate acti…
arXiv:2511.22254v5 Announce Type: replace Abstract: Self-evolving agents improve their performance on long-horizon tasks by learning from their own interactions with an environment. A common approach uses the resulting failed trajectories as negatives for preference training. How…
arXiv cs.AI
TIER_1English(EN)·Victor May, Van Khue Nguyen, Aaditya Salgarkar, Yishan Wang, Diganta Misra, Huu Nguyen·
arXiv:2603.01209v3 Announce Type: replace Abstract: In CodeAct, language-model agents write Python that calls tools and use execution feedback to choose actions. Persistent runtimes preserve Python variables between actions; stateless runtimes clear them without resetting task pr…
arXiv cs.AI
TIER_1English(EN)·Daniel Mitropolsky, Riccardo Neumarker, Emanuele Rimoldi, Susan S. Hong, Tomaso Poggio·
arXiv:2605.10851v2 Announce Type: replace Abstract: We initiate the study of the Generalized Turing Test (GTT), a formal generalization of Turing's imitation game from humans to arbitrary interactive agents. For agents $A$ and $B$, $A$ passes the GTT against $B$ if an instance of…
arXiv cs.AI
TIER_1English(EN)·Youngmok Jung, Sirajul Salekin, Henry Tran, Javier Movellan, Zhao Huang, Manjot Bilkhu·
arXiv:2609.32391v2 Announce Type: replace Abstract: Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, c…
arXiv cs.AI
TIER_1English(EN)·Xiao-Wen Yang, Weiyi Xu, Wen Da, Hang Xu, Canwei Li, Hong-Jie You, Pusen Dong, Yucheng Zeng, Zhaokai Luo, Yu-Feng Li, Yao Hu, Mu Chuan·
arXiv:2609.33665v2 Announce Type: replace Abstract: Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require a…
arXiv:2609.33676v2 Announce Type: replace Abstract: LLM agents increasingly take consequential actions through interactions with users, policies, and external tools. Auditing these agents requires automated attribution of realized actions to their historical basis. However, exist…
arXiv:2609.35596v2 Announce Type: replace-cross Abstract: Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills,…
arXiv:2609.38334v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the …
arXiv:2609.38923v1 Announce Type: new Abstract: Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Ex…
As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repea…
As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repea…
Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently…
Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a con…
Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers tok…
Causal action verifiers gate an agent's state-changing tool calls by checking whether each proposed intervention is identifiable against a committed action-state graph, and they issue a certificate that carries the identification argument and a one-sided lower confidence bound. O…
Modern AI coding-agent harnesses (Claude Code, Codex CLI, Cursor) rest their security boundary on a largely unexamined assumption: that the action A a human approves is the same action A' the harness executes, where A is fixed by a stated policy for what a scope grant or session-…
arXiv:2609.34603v2 Announce Type: replace Abstract: Does repairing an episode make its experience a better memory for the next task? We transfer the same failed source before and after accepted repair to a fixed target, alongside independent execution. Our 3,300 runs cover 100 Th…
arXiv:2609.37539v1 Announce Type: new Abstract: Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data …
arXiv:2609.37267v1 Announce Type: new Abstract: Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and e…
arXiv:2609.37236v1 Announce Type: new Abstract: An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on …
arXiv cs.AI
TIER_1English(EN)·Savini Kashmira, Jayanaka L. Dantanarayana, Lingjia Tang, Jason Mars·
arXiv:2609.37743v1 Announce Type: new Abstract: LLM agents performing long-horizon tasks accumulate tool results that later steps may need. Passing the full history to every invocation is costly even when it fits within the context window, while reducing it risks omitting needed …
arXiv:2609.37475v1 Announce Type: new Abstract: Tool-using language-model agents can decide that an action is permissible and execute it only after security-relevant state has changed. We study this proposal-to-commit gap and introduce BSC-R, a deterministic effect-boundary mecha…
arXiv cs.AI
TIER_1English(EN)·Jungwoo Yang, In Jin Kong, Yohan Jo·
arXiv:2609.37968v1 Announce Type: new Abstract: Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstre…
arXiv:2609.38143v1 Announce Type: new Abstract: Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weight…
arXiv cs.AI
TIER_1English(EN)·Paras Dahal, Anton Bakhtin, Taco Cohen, Zhengxing Chen, Carole-Jean Wu, Rob Fergus, Scott Yih, Gabriel Synnaeve, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston, Anirudh Goyal·
arXiv:2609.38147v1 Announce Type: new Abstract: As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to …
arXiv cs.AI
TIER_1English(EN)·Ranuga Disansa, U. S. Samarasinghe, Lasith Gunawardena·
arXiv:2609.35806v1 Announce Type: cross Abstract: Automated skill extraction underpins workforce planning, yet most systems represent skills as flat labels with no notion of the responsibility level at which a skill is practiced. The Skills Framework for the Information Age (SFIA…
arXiv cs.AI
TIER_1English(EN)·Linzhi Peng, Hanting Chen, Heng Chang, Ke Cheng, Bowen Du, Weifeng Lv·
arXiv:2609.35816v1 Announce Type: cross Abstract: Large language model search agents are often trained with synthetic questions whose difficulty is increased through larger evidence graphs, additional hops, and longer trajectories. These global properties, however, are only indir…
arXiv:2609.35912v1 Announce Type: cross Abstract: Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, a…
arXiv cs.AI
TIER_1English(EN)·Xiaoyu Xiong, Tsun-Hsuan Wang, Yi-Ling Qiao, Tao Du, Minchen Li·
arXiv:2609.36593v1 Announce Type: cross Abstract: Creating diverse physical simulations remains labor-intensive because assets, layout, physical parameters, motion, control, and rendering must be designed and debugged jointly. We present Text2Sim, a simulation-specialized agentic…
arXiv:2609.36635v1 Announce Type: cross Abstract: Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings …
arXiv cs.AI
TIER_1English(EN)·Jiexing Qi, Yu He, Jun Liu, Qichen Huang, Shaohua Hu, Zhan Dang, Guohua Chen, Rui Yang, Wen Jiang, Yang Liu, Tao Lyu, Fangming Li·
arXiv:2609.37105v1 Announce Type: cross Abstract: Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates ch…
arXiv cs.AI
TIER_1English(EN)·Soyeong Jeong, Sujay Kumar Jauhar, Sung Ju Hwang, Andrew Joohun Nam·
arXiv:2609.37226v1 Announce Type: cross Abstract: Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its lates…
arXiv:2609.37315v1 Announce Type: cross Abstract: Tool-using agents are entering settings where a wrong action carries real cost, and the benchmarks certifying them grade what each simulated tool call reports having done, assuming the tool did what its interface advertises. The a…
arXiv:2609.37359v1 Announce Type: cross Abstract: Coding agents can now write, run, and debug programs with little human help. Robot tasks, however, are usually specified by a sentence that leaves out how to grasp, in what order to make contact, and what the result should look li…
arXiv:2609.37468v1 Announce Type: cross Abstract: Agentic retrieval-augmented generation (RAG) interleaves reasoning with repeated retrieval, giving the retriever influence over both the evidence an agent observes and its subsequent search decisions. We study retriever backdoors …
arXiv:2609.37810v1 Announce Type: cross Abstract: Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potent…
arXiv cs.AI
TIER_1English(EN)·Yiming Cheng (The University of Chicago), Alfin Wijaya Rahardja (Fudan University), Mengshi Zhang (TensorBlock, Inc), Zihao Chen (TensorBlock, Inc), Zhenpeng Chen (Tsinghua University), Yiling Lou (University of Illinois Urbana-Champaign)·
arXiv:2609.37864v1 Announce Type: cross Abstract: Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is further hindered by existing benchmarks, which contain only a small and fixed number…
arXiv cs.AI
TIER_1English(EN)·Julien Knafou, Luc Mottin, Alexandre Flament, Paul van Rijen, Esteban Gaillac, Patrick Ruch·
arXiv:2609.37993v1 Announce Type: cross Abstract: The BITEM team entered both subtasks of the NTCIR-19 R2C2 task with a single agentic pipeline, in which a model searches, reads and records evidence over a movie corpus while an orchestrator holds the record and rules on what may …
arXiv:2605.09278v2 Announce Type: replace Abstract: Multi-agent debate (MAD) systems increasingly rely on shared memory to support long-horizon reasoning, but this convenience opens a critical vulnerability: a single corrupted entry can contaminate the downstream memory-augmented…
arXiv:2609.32192v2 Announce Type: replace Abstract: Multi-agent workflows require task-relevant information to be shared across agents, while irrelevant, stale, unverified, or incompatible information must remain isolated. We call this task-conditioned scope of information a coll…
arXiv:2609.32490v2 Announce Type: replace Abstract: LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task requirements are sufficiently specified before execution. In practice, user requests are often incomplete, and…
arXiv:2609.32754v2 Announce Type: replace Abstract: Large language model agents can often make reasonable local decisions on short tasks, yet their performance degrades when success requires long sequences of dependent actions and tool calls. During execution, task requirements, …
arXiv cs.AI
TIER_1English(EN)·Zhengkun Di, Bin Shi, Kai Sun, Yiming Xu, Bo Dong·
arXiv:2609.37125v1 Announce Type: new Abstract: Prospective memory allows an agent to retain an intention tied to a future condition, but the stored intention does not reveal whether that condition currently holds. Checking it may require web access, multi-step tool use, and paid…
arXiv:2609.37025v1 Announce Type: new Abstract: As large language models (LLMs) advance, AI agents are increasingly deployed in open-world environments to tackle complex sequential tasks (e.g., document processing, cross-application collaboration), relying heavily on actions rang…
arXiv cs.AI
TIER_1English(EN)·Junjie Yao, Zhangchen Zhou, Zhi-Qin John Xu·
arXiv:2609.37012v1 Announce Type: new Abstract: For a long-horizon agent, context is the bottleneck: the history is resent with every request, the window caps task length, and reasoning degrades as the history grows. Replacing structured objects with compact retrieval Cards short…
arXiv:2609.36892v1 Announce Type: new Abstract: As the capabilities of large language models (LLMs) continue to advance, increasing attention is turning to how to translate their abilities into useful behavior. Personal agents bring this question into everyday settings, where mod…
arXiv:2609.36887v1 Announce Type: new Abstract: Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent ha…
arXiv:2609.36746v1 Announce Type: new Abstract: Agent skills provide a lightweight mechanism for self-evolving agents to accumulate reusable procedural knowledge without updating model parameters. However, existing learned skill curators typically optimize curation without explic…
arXiv cs.AI
TIER_1(CA)·Gabriel Orlanski, Alex L. Zhang, Avi Trost, Vincent Sunn Chen, Frederic Sala, Aws Albarghouthi, Ludwig Schmidt·
arXiv:2609.36730v1 Announce Type: new Abstract: Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDe…
arXiv:2609.36679v1 Announce Type: new Abstract: Machine learning engineering (MLE) agents have made substantial progress, but learning through ML experimentation remains costly in time and computation. Synthetic environments reduce these costs while introducing variations in data…
arXiv:2609.34373v2 Announce Type: replace Abstract: Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol …
arXiv:2609.36488v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated strong performance on complex web navigation tasks, yet they remain brittle in real-world settings where user intentions are underspecified and preferences are heterogeneous. In pr…
arXiv:2609.37143v1 Announce Type: cross Abstract: Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' imple…
arXiv:2609.37017v1 Announce Type: new Abstract: LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. However, directly forwarding all sender latents makes…
arXiv cs.CL
TIER_1English(EN)·Amirhossein Abaskohi, Amirhossein Dabiriaghdam, Lele Wang, Peter West, Giuseppe Carenini·
arXiv:2609.36344v1 Announce Type: new Abstract: Deep-research agents conduct long-horizon investigations through iterative search, evidence evaluation, belief revision, and synthesis. However, they may commit to claims before sufficient evidence is available, causing later reason…
arXiv:2609.33875v2 Announce Type: replace-cross Abstract: Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterfactual Rollout Replay (CRR), a training-ti…
arXiv cs.AI
TIER_1English(EN)·Weiyi Xu, Xiaowen Yang, Wen Da, Hang Xu, Canwei Li, Hongjie You, Pusen Dong, Yucheng Zeng, Zhaokai Luo, Mu Chuan·
arXiv:2609.33772v2 Announce Type: replace Abstract: Executable environments are critical for post-training agents on tasks that require tool use and multi-step interaction, but constructing executable tasks together with their environments remains difficult to scale. Skills provi…
arXiv cs.AI
TIER_1English(EN)·Michael Lee, Zhipeng Wei, Yue Dong, N. Benjamin Erichson·
arXiv:2609.36576v1 Announce Type: new Abstract: Agentic systems are now being widely used to orchestrate tools and reason over long contexts. However, the improving capabilities of the large language models powering these agents also create new attack surfaces for indirect prompt…
arXiv:2609.35875v1 Announce Type: new Abstract: Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity …
arXiv cs.AI
TIER_1English(EN)·Yihao Wang, Linhan Xia, Rui Liu, Zhaofeng Zhang, Hongyu Wu, Yang Yang, Jinglu He, Yu Guo, Kai Lei·
arXiv:2609.36043v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate…
arXiv:2609.36319v1 Announce Type: new Abstract: Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based…
arXiv cs.AI
TIER_1English(EN)·Long Phan, Stephen K. Yang, Jason J. Lim, Mantas Mazeika, Wenyu Zhang, Zheyuan Liu, Richard Ren, Jingxiang Meng, Yaoteng Tan, Weiliang Zhao, Addison Wu, Matei Anghel, Dan Hendrycks·
arXiv:2609.36308v1 Announce Type: new Abstract: Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to…
arXiv cs.AI
TIER_1English(EN)·Thibaud Gloaguen, Niels M\"undler-Sasahara, Mark Niklas M\"uller, Veselin Raychev, Martin Vechev·
arXiv:2602.11988v3 Announce Type: replace-cross Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md. Although this practice is strongly encouraged by agent developers, there is currently no rigo…
arXiv cs.AI
TIER_1English(EN)·Ziluowen Luo, Senzhang Wang, Chaozhuo Li, Jun Yin, Hao Yan, Ming Cheng, Chenxu Wang, Songyang Liu, Litian Zhang, Qiwei Ye, Zheng Liu, Philip S. Yu·
arXiv:2609.36630v1 Announce Type: new Abstract: Modern agents increasingly rely on memories, tools, and execution logic, so their competence extends beyond model parameters. This shift exposes a limitation of conventional knowledge distillation, which asks how a student model imi…
arXiv:2605.12070v3 Announce Type: replace-cross Abstract: Asynchronous reinforcement learning improves rollout throughput for large language model agents by decoupling sample generation from policy optimization, but it also introduces a critical failure mode for PPO-style off-pol…
arXiv cs.AI
TIER_1English(EN)·Ziyu Liu, Jun Chen, Lixu Wang·
arXiv:2609.36626v1 Announce Type: new Abstract: Language-model agents increasingly rely on persistent natural-language skills to adapt beyond their frozen model parameters. When a shared skill is repeatedly revised from a non-stationary, heterogeneous task stream, however, improv…
arXiv:2609.33180v2 Announce Type: replace-cross Abstract: As recursive self-improvement (RSI) rapidly advances, reliable evaluation becomes critical for guiding adaptive search. RSI typically relies on finite evaluation resources, such as fixed benchmarks, to determine which modi…
Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection…
Patch retrieval, the task of finding the commit that fixes a known vulnerability, is the foundation of vulnerability management workflows, yet 60% to 63% of CVEs in the major advisory databases lack a patch link. We present PatchHolmes, a two-phase patch retrieval system that pai…
When a multi-agent system answers correctly, it is tempting to conclude that its agents shared, checked, and used information as intended. Yet a system can break one of its collective mechanisms, the rules that govern how agents route, admit, store, and act on shared information,…
Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Existing pipelines either generate files with mode…
On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound ac…
Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection…
Patch retrieval, the task of finding the commit that fixes a known vulnerability, is the foundation of vulnerability management workflows, yet 60% to 63% of CVEs in the major advisory databases lack a patch link. We present PatchHolmes, a two-phase patch retrieval system that pai…
Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge…
Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form Game Design Documents (GDDs) describe requireme…
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so…
Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting because they support parallel execution and isolate e…
LLM agents generate intermediate reasoning and actions token by token, making extended interactions slow and computationally expensive. Jev-style models offer fast probabilistic predictions over finite fields, but require those fields to be specified in advance. This requirement …
Recursive self-improvement (RSI) lets a system improve from its own outcomes; in LLM-based multi-agent systems, Agents refine one another within a task, and outcomes improve how they collaborate across tasks. However, existing multi-agent collaboration leaves this loop open: coll…
Stronger AI agents do not automatically produce better organizations: teams must also learn which work arrangements to retain and when to reconsider them. We propose a modeling specification for recursive organization improvement and evaluate it through an executable checker, a p…
Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or augment GUI interactions with application specif…
The BITEM team entered both subtasks of the NTCIR-19 R2C2 task with a single agentic pipeline, in which a model searches, reads and records evidence over a movie corpus while an orchestrator holds the record and rules on what may be submitted. A claim is admitted only once an ent…
Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, t…
Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three join…
An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. …
Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach th…
Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' implementation capability to produce correct code edits…
LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. However, directly forwarding all sender latents makes the receiver-side context scale with both the n…
arXiv cs.AI
TIER_1English(EN)·Md Shohel Arman, Igor Molybog·
arXiv:2609.31587v1 Announce Type: cross Abstract: We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether co…
arXiv cs.AI
TIER_1English(EN)·Weida Liang, Shi Qiu, Zhun Wang, Simon Sure, Xiaoyuan Liu, Tianneng Shi, Zhaorun Chen, Wenbo Guo, Dawn Song·
arXiv:2609.31318v1 Announce Type: cross Abstract: AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software co…
arXiv:2609.31301v1 Announce Type: cross Abstract: Large language model agents increasingly act on software systems, no longer merely generating text but also changing databases and online services. However, an approved database update may succeed yet leave an unapproved notificat…
arXiv:2609.30604v1 Announce Type: cross Abstract: Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spre…
arXiv:2609.30558v1 Announce Type: cross Abstract: Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a cogniti…
arXiv cs.AI
TIER_1English(EN)·Mukul Chhabra, Shail Patel, Luigi Medrano·
arXiv:2609.30471v1 Announce Type: cross Abstract: Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically appl…
arXiv cs.AI
TIER_1English(EN)·Leon Goldberg, Gal Engelberg, Eden Yavin, Elad Elouz, Ariel Zadok, Konstantin Koutsyi·
arXiv:2609.30345v2 Announce Type: cross Abstract: Cloud-security investigation is dominated by population tasks: which identities can read a data store, how many resources fail a control, what is reachable from another account. These resolve against a complete inventory, not a na…
arXiv:2609.31563v1 Announce Type: new Abstract: Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analy…
arXiv cs.AI
TIER_1English(EN)·Bart{\l}omiej Cupia{\l}, Jens Tuyls, Maciej Wo{\l}czyk, Davide Paglieri, Martin Klissarov, Benjamin Eysenbach, Piotr Mi{\l}o\'s, Karthik R. Narasimhan·
arXiv:2609.31076v1 Announce Type: new Abstract: Language agents struggle to act and learn in environments that require long sequences of low-level actions. Code-based abstractions can make these agents more productive by letting them invoke reusable skills instead of repeatedly s…
arXiv cs.AI
TIER_1English(EN)·Zhensheng Zou (Peking University), Guoqing Wang (Peking University), Dan Hao (Peking University)·
arXiv:2609.31430v1 Announce Type: new Abstract: Tool observations dominate the context of software-engineering agents, making long interaction histories costly to maintain. Existing context compression methods can discard information needed by later actions, while adapting agents…
arXiv:2609.30971v1 Announce Type: new Abstract: Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely…
arXiv cs.AI
TIER_1Deutsch(DE)·Guanyu Nie, Fangzhou Zhu, Shixiong Kai, Xiongwei Han, Tao Zhong, Mingxuan Yuan·
arXiv:2609.30861v1 Announce Type: new Abstract: Language-model agents increasingly improve by converting execution experience into reusable external skills. Yet repeated skill updates form a learning process of their own: locally useful edits can accumulate into redundant or task…
arXiv:2609.30813v1 Announce Type: new Abstract: Evaluating claim admission in shared agent memory is challenging because repeated claims may be mistaken for independent evidence. An agent may copy or paraphrase a retrieved belief, while admitting a false claim exposes subsequent …
arXiv cs.AI
TIER_1English(EN)·Shane Caldwell, Max Harley, Ads Dawson, Michael Kouremetis, Vincent Abruzzo, Will Pearce·
arXiv:2609.30325v1 Announce Type: new Abstract: Agents are increasingly deployed with real autonomy in web application and network penetration testing, where a single out-of-scope action can breach a client's engagement boundary. Existing offensive-security benchmarks measure raw…
arXiv cs.AI
TIER_1English(EN)·Yiran Hu, Nan Jiang, Shanchao Liang, Anik Dey, Yi Wu, Lin Tan·
arXiv:2609.30725v2 Announce Type: new Abstract: Although effective, coding agents often incur substantial monetary costs. Their recurring cost-inefficient behaviors remain underexplored. We conduct the first study of behavioral cost inefficiencies in coding agents, analyzing 1,20…
arXiv:2609.30383v1 Announce Type: new Abstract: A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable …
arXiv:2609.30328v1 Announce Type: new Abstract: When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for. Multi-agent v…
arXiv cs.AI
TIER_1English(EN)·Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth·
arXiv:2605.06869v3 Announce Type: replace Abstract: AI agent research spans a wide spectrum: from RL agents that learn from scratch to foundation model agents that leverage pre-trained knowledge, yet no unified benchmark enables fair comparison across these approaches. We present…
arXiv:2504.06188v3 Announce Type: replace Abstract: AI agents can extend their capabilities at inference time by loading reusable skills into context, yet equipping an agent with too many skills, particularly irrelevant ones, degrades performance. As community-driven skill reposi…
Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an age…
Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in is…
Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or augment GUI interactions with application specif…
Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three join…
An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. …
Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach th…
Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as policies or controllers, offering limited cove…
Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain und…
Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the cost of additional training and errors that comp…
Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven search and identify three core challenges: organizing evolving research ideas, selecting promising directions, and maintaining alig…
Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed. To make the Builder's experience…
We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, which produces a solution better than the current o…
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We …
Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills, in response to user and environment feedback. However, lo…
Multi-agent systems bring together language model agents with different roles to propose, review, and refine solutions. Each agent's response depends on its model's capabilities, the reasoning strategy defined by its system prompt, and the information in its input context. Existi…
Agents need calibration when deployment conditions change: replacing a driving model, including a foundation-to-post-trained transition; crossing jurisdictions; or scaling across heterogeneous markets and sources. Interface compatibility alone does not establish capability retent…
Memory is critical for AI agents. Many existing agent-memory systems follow an Ahead-of-Time (AOT) design, constructing memory before a specific request arrives. While this reduces online serving cost, such request-agnostic memory construction can discard fine-grained information…
Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer t…
Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer t…
Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many existing credit assign…
Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether rank…
Self-evolving agents can continually improve their behavior, while tools define the executable action space through which they interact with the environment. However, exposing the full tool library to model introduces substantial irrelevant context and can impair tool-use decisio…
LLM-based agents have shown strong capabilities in automated data analysis and are increasingly moving toward long-horizon, multi-stage analytical workflows. However, as the analytical process evolves, constraints, variables, and conclusions remain implicitly embedded in interact…
Recent gains in language model capability have come more from data than from architecture. Frontier labs and data companies produce verifiable agentic tasks, which supervised finetuning and reinforcement learning then turn into capability.This production line still rests on human…
Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized info…
Language model agents serving organizations must coordinate requests from multiple users while using knowledge distributed across their interactions. We identify two complementary capabilities for this setting, namely cross-user interaction and decision-making, as well as cross-u…
Persistent memory lets language-model agents carry information across long-running collaborations, but leaves a lifecycle question open: what may a returning agent still act on once the shared state has changed? A memory can be correctly retrieved, relevant to the current task, a…
As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to speci…
An agent can solve a task today and face the same task under new facts, tools, or governing knowledge tomorrow. Most agent systems can retrieve relevant text or recall prior conversations, but they lack a principled way to decide when a previous solution is still valid, when it m…
Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multipl…
As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to speci…
Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity among agents is the driver, in the setting where…
Agentic retrieval-augmented generation (RAG) interleaves reasoning with repeated retrieval, giving the retriever influence over both the evidence an agent observes and its subsequent search decisions. We study retriever backdoors that exploit this feedback loop and repurpose weak…
Self-improving agents increasingly rely on proxy verifiers to choose policy updates, yet deployment can change the world in which those updates are evaluated. An update that looks better to the verifier can therefore become worse after deployment even when the verifier ranks poli…
Multi-agent coding has emerged as an increasingly active direction in software engineering, where complex development tasks are decomposed across multiple specialized agents working on different parts of the problem. Despite the shift from individual problem solving to distribute…
Reusable skills help LLM-based agents solve complex tasks, but the agent must receive guidance before it commits to an ineffective approach. Existing skill mechanisms often expose only metadata and load full content on demand, leaving useful guidance unavailable until the agent d…
Browser-use agents often carry information in their context as they move between websites. While it may be necessary for task completion, it also creates a privacy risk, especially when the information contains a private fact regarding the user. For example, an agent may learn a …
Learning from experience in LLM agents has become a key paradigm for developing self-evolving agents that continuously learn and expand their capabilities. Within this paradigm, synthesizing the agent skill has emerged as a promising solution for transforming accumulated experien…
Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in impleme…
Multi-step agents are trained on flat action streams: SFT and RLVR weight every token uniformly and ignore the sub-procedures that recur across tasks, the hierarchy that lets humans plan top-down from reusable routines. This structure sits unused, and flat training uses each scar…
Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the an…
arXiv cs.AI
TIER_1English(EN)·Beining Wu, Zihao Ding, Jun Huang·
arXiv:2609.29545v1 Announce Type: new Abstract: Deployed agents run on handed-over knowledge: a frozen policy consults a briefing of consolidated items written before the stream begins. The world then moves while the store stands still: paths close, flags change, price bands move…
arXiv cs.AI
TIER_1English(EN)·Yezhou Cheng, Runjia Du, Zeming Liu, Qibai Chen, Hang Lyu, Yankai Zeng, Yilan Wei, Bojun Lin·
arXiv:2609.29144v1 Announce Type: new Abstract: Persistent memory lets language-model agents improve prompts and skills without updating model weights. We show that matching retrieval scope to certification scope enables these edits to support reliable repeated adaptation across …
arXiv:2609.29875v1 Announce Type: new Abstract: Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. Unlike static Chain of Thought compression, remo…
arXiv:2609.29547v1 Announce Type: cross Abstract: Agentic AI is sold on a simple promise: the system keeps acting when the user stops watching. That promise creates an accountability inversion. As stepwise supervision recedes, verification does not disappear; it moves into the ru…
A multi-agent system can reduce latency on complex tasks by executing work concurrently. Several pioneering harness frameworks support multi-agent systems. However, the scalability of current multi-agent harnesses is often constrained by a central orchestrator's capacity to alloc…
Large language model agents increasingly retrieve reusable Skills and inject them into the active context. However, a retrieved Skill can be relevant yet unnecessary, costly, or even harmful in the current execution state. We present SkillApt, a post-retrieval activation framewor…
Multi-agent reinforcement learning (MARL) commonly trains decentralized policies from scratch, requiring agents to acquire individual task competence and coordination simultaneously. Yet many multi-agent problems admit a compatible single-agent counterpart in which the underlying…
As generative AI agents are deployed at scale, safety will depend not only on technical safeguards and individual model design, but also on collective equilibria that determine how agent populations process information, prioritize actions, and respond to uncertainty. Yet the same…
Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with functional verifiers, as seen in OSWorld. However, such evaluation of end-state performance lacks transparency into how and why agen…
Collective intelligence depends not only on what team members know, but also on how they organize their work. When the structure of a solution is unknown, useful roles and divisions of labor cannot be specified in advance; teams must learn from experience how to organize reasonin…
arXiv cs.CV
TIER_1English(EN)·Liu Renhang, Navonil Majumder, Tej Deep Pala, Soujanya Poria·
arXiv:2610.10388v1 Announce Type: cross Abstract: Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-rele…
arXiv:2610.09430v1 Announce Type: new Abstract: Internalizing external skills changes a language agent's capabilities and, with them, the value of its remaining guidance: rules can become redundant, misleading, or insufficient for newly encountered decisions. This creates a coupl…
arXiv cs.CV
TIER_1Italiano(IT)·Sacha Morin, Kumaraditya Gupta, Mahtab Sandhu, Charlie Gauthier, Francesco Argenziano, Kirsty Ellis, Liam Paull·
arXiv:2509.19571v2 Announce Type: replace-cross Abstract: Designing or learning robot policies that generalize zero-shot across a range of language instructions and objects is a core problem in robotics. Vision-Language-Action models (VLAs) learn such policies end-to-end by repur…
arXiv:2609.39135v1 Announce Type: new Abstract: Physical reasoning from video requires inferring latent physical properties and dynamics beyond direct observation. Direct VLM inference remains unreliable on complex physical tasks without explicit modeling and validation, while pr…
arXiv:2609.38008v1 Announce Type: new Abstract: Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or …
arXiv:2609.30650v1 Announce Type: cross Abstract: Task performance need not determine which intervention mechanism an agent retains. We study causal retention: whether a frozen learned state answers a mechanism-probe map fixed independently of training, including action, context,…
AWS Machine Learning Blog
TIER_1English(EN)·Manish Ballal·
The RPA-era ROI model misses most of the value agentic automation creates. This post gives AI center of excellence leaders a framework to size the full value of agents across time savings, exception handling, decision quality, and maintenance economics, and to prioritize which wo…
AWS Machine Learning Blog
TIER_1English(EN)·Thiago Verney·
Off-the-shelf AI assistants forget you between conversations. This post shows how to build a personal assistant that accumulates context using OpenClaw on Amazon Bedrock AgentCore runtime, with AgentCore memory turning disposable chats into durable, structured knowledge you can r…
AWS Machine Learning Blog
TIER_1English(EN)·Mona Mona·
Amazon SageMaker optimized generative AI inference introduces the aws-ai-ml skill through the Agent Toolkit for AWS, giving coding agents like Kiro, Claude Code, and Codex deep expertise in inference optimization and benchmarking. Describe what you want, and your agent generates …
AWS Machine Learning Blog
TIER_1English(EN)·Manideep Reddy Gillela·
Build a Retrieval Augmented Generation (RAG) application on Amazon Bedrock Managed Knowledge Base with LangChain, and see how agentic retrieval handles the multi-part questions that single-shot retrieval answers poorly. Run the same query through both paths, read the trace events…
AWS Machine Learning Blog
TIER_1English(EN)·Kanishk Mahajan·
Multi-agent systems need deeper guarantees than fluent responses: they must select the right tools, respect constraints, and explain their decisions. Learn how to build a Strands-based multi-agent supply chain decisioning system and evaluate it with Amazon Bedrock AgentCore Evalu…
<p>JetBrains released Mellum2.1, an Apache 2.0, 12B mixture-of-experts thinking model with 2.5B active parameters. RL in real repositories lifted its SWE-bench Verified score from 2.0 to 47.0.</p> <p>The post <a href="https://www.marktechpost.com/2026/10/08/jetbrains-releases-mel…
dev.to — Claude Code tag
TIER_1English(EN)·Steven Gonsalvez·
<p><em>Human thoughts, AI-assisted write-up.</em></p> <p>Strip an agent back to its skeleton and it's embarrassingly simple. A while loop.<br /> </p> <div class="highlight js-code-highlight"> <pre class="highlight plaintext"><code>while not done: context = observe() # read files,…
<p>The leap from simple chat prompts to agentic workflows is the biggest upheaval in software development since the introduction of Git. While 2024 and 2025 were still dominated by autocomplete plugins and "vibe coding" in the headlines, the landscape has fundamentally changed by…
<p>IBM has made a self-hosted deployment of IBM Bob, its agentic software development platform, generally available. Enterprises can now run Bob on premises, in private or sovereign clouds, and in air-gapped networks. They bring their own model: NVIDIA Nemotron or Poolside Laguna…
IQuest Research released open weights for IQuest-Q1, a 320B sparse MoE with 15B active parameters built for command-line coding agents, with 512K context and 84.5 on CyberGym.
dev.to — MCP tag
TIER_1English(EN)·Fernando Azevedo·
<p>Google's post on the AI Agents Challenge says something it took me years to accept on financial platforms: "multi-agent" was the most frequent claim across thousands of submissions, and a good share of them were a single model walking through a prompt chain with agent names at…
Medium — Claude tag
TIER_1English(EN)·Fellows Monika·
Email — Every
TIER_1English(EN)·010001a117d0d056-a9eaa087-9035-4580-bb51-9434736e88d3-000000@send.every.to (010001a117d0d056-a9eaa087-9035-4580-bb51-9434736e88d3-000000@send.every.to)·
<!-- Set the language of your main document. This helps screenreaders use the proper language profile, pronunciation, and accent. --> <!-- The title is useful for screenreaders reading a document. Use your sender name or subject line. --> Building a More Efficient Agent <!-- Neve…
dev.to — Anthropic tag
TIER_1English(EN)·Jorge Peraza·
<h1> Why CAI exists: the agent-payment problem and the three roles (custodian, approver, operator) </h1> <p>The agent-payment problem is older than agents, and it has a clear shape. This post is the H1-readable essay on the problem CAI solves, the user-confirmation pattern, and t…
<p>The agent economy doesn't wait for perfect infrastructure. Agent A is already paying Agent B to source data, optimize training runs, and verify execution quality. Agent B is already hiring Agent C to do the actual work. And today, all three are using custodians to move money.<…
Medium — MLOps tag
TIER_1English(EN)·Swapnil Surushe·
<div class="medium-feed-item"><p class="medium-feed-snippet">This is Part 2 of the series, The Enterprise AI Agent Blueprint. Read Part 1 here.</p><p class="medium-feed-link"><a href="https://medium.com/@swapnil29071999/building-the-conductor-a-scalable-ai-agent-orchestrator-on-c…
<p>We've open sourced TAP (Trusted Agent Primitives) at Telara so agents can build reusable internal tools for the work they do repeatedly. The packages are code that developers can write, inspect and maintain too.</p> <p>Employees ask agents to investigate issues, prepare custom…
<p>We've open sourced TAP (Trusted Agent Primitives) at Telara so agents can build reusable internal tools for the work they do repeatedly. The packages are code that developers can write, inspect and maintain too.</p> <p>Employees ask agents to investigate issues, prepare custom…
<p><em>The developer-to-production pipeline for Snowflake’s Cortex Agent GA enhancements — Personal Database sandboxes, temporary agents, COPY GRANTS, and the cost math on Cortex Search suspension.</em></p><p><strong>Part 2 of 2</strong> — Part 1: <a href="https://snowflakechroni…
<h4><em>A practical guide to typed AI decisions, calibrated probabilities, and the software that turns them into useful workflows.</em></h4><p>An agent receives a request: “Investigate the failed deployment and explain what changed.”</p><p>Before it produces an answer, the system…
Medium — Claude tag
TIER_1English(EN)·Rathish Poovadan·
<h1> The Settlement War: Why Routing Isn't Enough for Agent Commerce </h1> <p>This week, three signals converged.</p> <p>On Tuesday, the <strong>IETF published draft-hood-agtp-commerce-00</strong> — the agent-to-agent commerce standard. It is a big deal. Agents can now discover e…
<p>When we talk about LLM reliability, our focus almost always gravitates toward accuracy—did the model get the math right? Did it retrieve the correct record from the database?</p> <p>But for engineers building autonomous agents using ReAct or Chain-of-Thought (CoT) patterns, th…
dev.to — MCP tag
TIER_1English(EN)·Shakar Bisetty·
<p><em>Part 2 of 10 · Building an Agentic Change-Approval MVP on MuleSoft</em></p> <p>In <a href="https://dev.to/thasha/the-21-step-change-nobody-wants-to-own-why-sap-change-promotion-is-an-agentic-use-case-4mpo">Part 1</a> I set out the use case: automating the approval and prom…
<p>The Model Context Protocol repository sits at 90,950 stars with 16 language SDKs and a collection of reference servers that expose how agent-tool boundaries actually work. These are not production systems. They are educational implementations that reveal transport layer choice…
dev.to — MCP tag
TIER_1English(EN)·Praveen Raj Thulasi S·
<h1> I Built an Agentic Analytics Platform — Here's What I Learned </h1> <p>What if you could ask your analytics dashboard:</p> <blockquote> <p><strong>"Why did sales decrease last month?"</strong></p> </blockquote> <p>and instead of manually filtering charts and writing database…
Medium — Claude tag
TIER_1Nederlands(NL)·Ivan Yanishevskyi·
<h4>Existing technologies for future simpler combined understanding and use</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*hk7w8hkMMAkvQ4me.jpeg" /><figcaption>(AgenticSystemCore — Text used by humans and Agentic AI, source:Author, Gemini)</figcaption></f…
<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*zgDOnsQQy6eaFaf_NN4E9g.jpeg" /><figcaption>AI Coding Agent Context Budget</figcaption></figure><p>More instructions can feel safer. For a coding agent, they can also bury the one rule that prevents a costly mista…
<p><strong>Short answer:</strong> An agent's search tool and a RAG index are both retrieval, and they differ in<br /> four places that decide everything downstream: what they read (the live web against a corpus you<br /> ingested), who owns the ranking (an engine you do not opera…
<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*C9cnr-mVxVYEVIuOsEgnMw.jpeg" /></figure><h3>The Data Debugging Grind</h3><p>My engineers typically spend 30% of their time debugging data. On some days, that number climbs to 50% or more.</p><p>If you’re a data o…
<h4><em>From “What is an agent?” to production-grade, multi-agent, tool-using, memory-equipped systems — everything you need in one place.</em></h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*RzS8r6JYTDqLuc2V" /></figure><p><strong>Reading time:</strong> ~…
SPONSORED POST: How agentic AI, real-time visibility, and stronger governance can help enterprises protect critical services and manage increasingly complex IT environments.
dev.to — MCP tag
TIER_1English(EN)·MANI BHUSHANAM K·
<p>🤔 What if a coding agent could actually remember?</p> <p>AI coding agents are becoming surprisingly capable.<br /> They can write code, inspect repositories, debug errors, interact with tools, and reason through complex development tasks.</p> <p>But there is still a frustratin…
Medium — MLOps tag
TIER_1English(EN)·Pavan Gosangi·
<blockquote> <p><strong>Key Takeaways</strong></p> <ul> <li>Claude Opus 5.5 launched September 22, 2026</li> <li>Anthropic's first model in the new Claude 5.5 family</li> <li>40% cheaper to run than Opus 5</li> <li>Optimized for agentic coding and knowledge work</li> <li>Availabl…
<p>Running one AI agent is easy. Running thirteen that actually cooperate is where memory becomes the real bottleneck.</p> <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A…
<p> </p> <p><strong>In this video:</strong></p> <p>0:00 Why agents fail around turn ten<br /> 0:19 It runs but is useless<br /> 1:54 The base-URL swap<br /> 3:26 Where compatibility breaks<br /> 5:28 From tokens to a tool call<br /> 8:31 Context, speed and turn count<br /> 12:25 …
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1x0ms0m/running_an_llmdriven_town_with_800_persistent/"> <img alt="Running an LLM-driven town with 800+ persistent agents: concurrency, context caching, and inference costs" src="https://preview.redd.it/71a7tj…
<p>As AI applications evolve from simple chatbots into complex multi-agent architectures, the demands placed on the vector database change fundamentally. In a single-agent loop, context retrieval is predictable and mostly sequential: an agent sends a search request, waits for a r…
<p><a href="https://nextigent.ai/services/model-orchestration/" rel="noopener noreferrer">Multi-agent LLM orchestration</a> coordinates multiple AI agents, language models, tools, and data sources so they can contribute to a shared workflow. Instead of expecting one model to inte…
<p>OpenAI and Ironclad just published a case study on using production contract workflows as both training data and evaluation benchmarks for computer-use agents. This is not a demo. It is a partnership where a SaaS company opens its workflow engine to become an agent training gr…
<p>Building autonomous AI agents is incredibly rewarding until you deploy them to production and real-world data breaks your clean pipelines.<br /> A common bottleneck is the tool execution layer. When your agent invokes a vector DB search or a live web API, it assumes it will re…
<p>Sarvam just released Arya, an orchestration stack designed to move agent systems from prototype to production. The announcement centers on a concrete problem: frontier models can write compilers and rebuild rendering pipelines in demos, but they fail silently and inconsistentl…
dev.to — LLM tag
TIER_1English(EN)·Abdulmuiz Adebayo·
<h2> A build-in-public deep dive from the smallOps project </h2> <p>Why this benchmark exists<br /> .<br /> smallOps is an experiment in giving small, locally-run language models — the kind that fit comfortably on a laptop with no GPU — the ability to act as coding agents. Tools …
<table> <tr><td> <a href="https://www.reddit.com/r/MachineLearning/comments/1wyw0my/swerace_a_codingagent_benchmark_of_188_real/"> <img alt="SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]" src="https://preview.redd.it/0uopztlmp…
<p>Every agent trace tool shows you a waterfall of steps. You spot that step six produced garbage. Now what? You re-run the entire pipeline and hope it lands in the same place. With a non-deterministic model, it doesn't. You can never separate your change from model jitter.</p> <…
dev.to — LLM tag
TIER_1English(EN)·Prabhakar Chaudhary·
<h1> Holo4: How H Company Built a Single Open-Weight Agent That Clicks, Codes, and Calls APIs </h1> <p>Most agentic AI models are specialists. A GUI-focused model can click through a browser but is lost when the task requires an API call. A tool-calling model can chain function c…
<p>Large Language Models (LLMs) can answer many questions, but answering questions over a large and complex corpus becomes difficult when the required information is spread across multiple documents or depends on relationships between different entities.<br /> To address this cha…
<p>On eight development coding cases repeated three times, a local Qwen3.8-27B agent went from <strong>5/24 to 23/24 functional and delivered successes</strong> across sequential output-budget and reasoning rounds. That is promising development evidence. It does not establish gen…
dev.to — LLM tag
TIER_1English(EN)·Rashid Mahmood·
<p>Models send bad tool calls: malformed JSON, a tool name that does not exist, a missing argument. What happens next is decided by the agent framework, not the model. In <a href="https://github.com/code-with-rashid/agentic-arena" rel="noopener noreferrer">agentic-arena</a> I mea…
<p>Production agents fail in ways that runtime guardrails cannot catch. They hallucinate plausible-sounding API calls, leak context across tool invocations, and drift from their original task specification without triggering a single exception. iFixAi is a Python CLI that audits …
dev.to — LLM tag
TIER_1English(EN)·Aleksei Romanov·
<p>For the past two years, the standard enterprise AI strategy looked suspiciously like a parlor trick: take an 8,000-token system prompt packed with JSON schemas, API documentation, and polite formatting rules, stuff it into the context window of a giant 400-billion-parameter fr…
dev.to — LLM tag
TIER_1English(EN)·Aleksei Romanov·
<p>Almost every modern language model looks impressive on a two-step demo. You ask it to check a database or summarize a document, it calls the right tool, formats the answer, and looks like an autonomous engineer.</p> <p>The illusion falls apart the moment you ask that same mode…
<blockquote> <p>Originally published at <a href="https://smartgate.network/industry/agentic-retrieval?utm_source=devto&utm_medium=syndication" rel="noopener noreferrer">Agentic Retrieval: The Loop That Decides When to Search</a> on smartgate.network.</p> </blockquote> <p>A sh…
<p>There are dozens of posts about <code>CLAUDE.md</code> best practices. Each one is a list of tips, and none of them explains <em>why</em>. We stopped guessing and went to the primary sources instead: Anthropic's own guidance and the 2026 papers on context engineering.</p> <p>T…
<p>Training LLM agents with reinforcement learning hits a hard wall: you need environments that provide verifiable rewards, support long-horizon interaction, and scale without burning budget. Human-curated data is expensive. LLM-generated environments hallucinate and leak benchma…
<p>Whiteboard is a YC W26-backed open-source IDE that replaces the file tree with a spatial canvas. Instead of navigating folders, you arrange components on a 2D plane. The project (422 HN points, 142 comments) exposes a different set of plumbing decisions for agent context manag…
<p>Sales conversations rarely happen in isolation.</p> <p>A customer might discuss pricing in one call, raise an objection in another, introduce a new stakeholder later, and finally ask for a specific next step several conversations after the first meeting.</p> <p>The problem is …
<p>When teams deploy large language models into enterprise workflows, the default architectural pattern connects a model to a user interface with high-level system instructions. In brief, deterministic tasks like code refactoring or single-turn customer support, this pattern work…
<p><em>Originally published at <a href="https://aifrontierpost.com/articles/okf-agent-memory-git-native-project-memory/" rel="noopener noreferrer">AI Frontier Post</a>.</em></p> <p>You know the feeling. You spend an hour with a coding agent establishing the architecture: why the …
dev.to — LLM tag
TIER_1English(EN)·Manidhar Bheempadu·
Concurrency, Forecasts, and Web Trust Boundaries "Multi-agent evals serialize the race condition they claim to measure" (general) + "An expectation without a settle date is not a forecast" (agents) This + more in today's Moltbook Pulse (Edition #83): https:// superagent-ebe00561.…
Building multi-agent systems in .NET? Explore a structured approach to coordinating specialized agents, managing workflows, and creating maintainable AI applications with familiar .NET patterns. # dotnet # AI https:// isaacl.dev/hbw
Docker Agent: a Tool for building, running and sharing AI Agents using a simple YAML configuration, rich tool Ecosystem and multi-agent Orchestration # AI # Agent https:// github.com/docker/docker-agent
🤖 Long-running agents: is the bottleneck the model or the scaffolding around it? Something I keep noticing with agent setups: every individual step is easy for the model, but the full chain still falls apart on long tasks. I think it comes down to three things: Error compoundin..…
An independent test found Hindsight, a memory bank for coding agents, retained unwritten project rules across sessions without explicit instruction. The tool draws on git history to help agents maintain conventions. Teams should verify storage location and check for empty knowled…
<table> <tr><td> <a href="https://www.reddit.com/r/ClaudeAI/comments/1wzt7ei/repos_dungeons_watch_your_agents_fight_their_way/"> <img alt="Repos & Dungeons: watch your agents fight their way through your codebase" src="https://external-preview.redd.it/aWZqZHE3YTM4MHVoMdDq__bK…
<table> <tr><td> <a href="https://www.reddit.com/r/ClaudeAI/comments/1wzkh5s/agent_communication_orchestration_in_velaterm/"> <img alt="Agent Communication & Orchestration in VelaTerm" src="https://external-preview.redd.it/dmE2aHF0bDFieXRoMapbV4K0o4qagdBawe0kyN1CzVcnkshsHO15v…