PulseAugur
中
实时 19:42:50
English(EN) Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

新的 AI 智能体研究聚焦于训练、可靠性和多智能体协调

多项研究工作正致力于提升 AI 智能体的能力和可靠性。Hugging Face 推出了 AutoSynthData,用于为企业智能体生成训练数据,以及 Holo4,一系列专为多样化接口设计的智能体模型。Apple 的 SCLATE 为持续学习智能体训练和评估提供了基础,而 xAI 的 Team Bots 则允许 AI 同事之间共享上下文和学习。此外,新的研究论文探讨了智能体技能演化(SkillSpec)、资源分配(TRACE)、多智能体团队中的可靠性(Worse Together),以及在复杂环境(如事件响应(Incident-Arena))和用户纠正仲裁(GAVA)中提升智能体性能。 AI

影响 智能体训练、评估和多智能体协调方面的进步可能加速企业采用,并提高 AI 在复杂任务中的可靠性。

排序理由 多篇与 AI 智能体开发、训练和评估相关的研究论文和平台公告。

在 Microsoft Research 阅读 →

AI 生成摘要 · Google Gemini · 来自 420 个来源。 我们如何撰写摘要 →

新的 AI 智能体研究聚焦于训练、可靠性和多智能体协调

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇与 AI 智能体开发、训练和评估相关的研究论文和平台公告。
Source corroboration
420 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
product, paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
19 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+164 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [420]

  1. Microsoft Research TIER_1 English(EN) · Zhiyuan He, Yuqing Yang ·

    Agent Lightning v1.0:一个3500行代码的轻量级Agentic RL框架,用于使用真实线束训练Agent

    <p>Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them.</p> <p>The …

  2. Apple Machine Learning Research TIER_1 English(EN) ·

    RISED:用于智能体多环境选择和自蒸馏的评分标准

    Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based sign…

  3. Hugging Face Blog TIER_1 English(EN) ·

    AutoSynthData:为企业智能体生成训练数据

  4. Apple Machine Learning Research TIER_1 English(EN) ·

    SCLATE:持续学习代理训练与评估的基石

    Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, crons, and memory consolidation. Yet existing benchma…

  5. Hugging Face Blog TIER_1 English(EN) ·

    Holo4:赋能通用计算机使用代理

  6. xAI news TIER_1 English(EN) ·

    团队机器人:能从你的团队学习的 AI 同事

    Give a Grok Bot the files, apps, and expertise it needs, then share it so your whole team can work from the same context.

  7. arXiv cs.LG TIER_1 English(EN) · Sangmin Lee, Youngju Na, Chanmi Lee, Sung-eui Yoon ·

    通过支持保持蒸馏实现多智能体协调

    arXiv:2610.10087v1 Announce Type: new Abstract: Offline MARL increasingly relies on generative policies to model multimodal joint behavior, typically by distilling a centralized teacher into decentralized one-step actors under the CTDE. We identify a failure mode at the teacher t…

  8. arXiv cs.LG TIER_1 English(EN) · Tianruo Rose Xu, Jiawei Ren, Yichi Yang, Zhaoxu Zheng, Lianhui Qin ·

    RT-Safe:实时具身环境中的智能体安全基准测试

    arXiv:2610.09294v1 Announce Type: cross Abstract: Rapid progress in AI agents has brought growing attention to agent safety, with extensive evaluation focused on digital environments. As agents move into the physical world, embodied safety becomes increasingly important: failures…

  9. arXiv cs.LG TIER_1 English(EN) · Yu Cheng, Dehai Zhao, Zhongxin Liu, Qing Huang, Zhenchang Xing, Xiaoxue Ren ·

    Agent Skills' Downstream Utility 的实证研究

    arXiv:2610.08875v1 Announce Type: cross Abstract: Agent Skills package procedural guidance and resources for reuse, but a relevant Skill does not necessarily improve task performance. Existing studies characterize Skill content and evaluate downstream performance, yet provide lim…

  10. arXiv cs.LG TIER_1 Deutsch(DE) · Prakhar Ganesh, Kyra Wilson, Luca Zappella, Barry-John Theobald, Nicholas Apostoloff, Lucas Monteiro Paes, Nivedha Sivakumar ·

    多智能体系统中的同质化

    arXiv:2610.09824v1 Announce Type: new Abstract: Multi-agent systems (MAS) leverage interactions between agents to perform complex tasks. Despite their success, we show that these interactions can also lead to homogenization, i.e., agents converging to similar behaviors. Homogeniz…

  11. arXiv cs.LG TIER_1 English(EN) · Yuanzhe Li, Pengxin Wang, Yuxin Ren, Jianing Deng, Jingtong Hu, Song Wang, Jingdi Chen, Huanrui Yang ·

    TAP:通过轨迹锚定恢复实现高效长视域代理剪枝

    arXiv:2610.09074v1 Announce Type: new Abstract: Emerging long-horizon agentic tasks require repeated model calls, worsening the inference cost of already-costly language models. While narrow agentic tasks suggest potential for aggressive model pruning without performance drop, em…

  12. arXiv cs.CL TIER_1 English(EN) · Igor Slinko, Yaroslav Golubev, Sergey Titov ·

    编码代理基准测试应匹配其用户的任务流程

    arXiv:2610.09633v1 Announce Type: cross Abstract: The evaluation of coding agents generally strives to be as realistic as possible. In our study, we collect 4,782 agent sessions of real software engineers in JetBrains IDEs, which we call Production Sessions. Since our subject is …

  13. arXiv cs.CL TIER_1 English(EN) · Xinglin Wang, Zishen Liu, Tong Zheng, Shaoxiong Feng, Peiwen Yuan, Yiwei Li, Jiayi Shi, Yueqi Zhang, Chuyi Tan, Ji Zhang, Boyuan Pan, Kan Li ·

    从帕累托到偏好:通过摊销代理策略发现实现个性化测试时缩放

    arXiv:2610.09684v1 Announce Type: new Abstract: Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dime…

  14. arXiv cs.CL TIER_1 English(EN) · Giordano De Marzo, Andres L. Marin, David Garcia ·

    AI 代理的集体行为:Moltbook 的案例

    arXiv:2602.09270v2 Announce Type: replace-cross Abstract: We present a large scale data analysis of Moltbook, a Reddit-style social media platform exclusively populated by AI agents. Analyzing over 4 million posts and 19 million comments from approximately 185,000 active agents, …

  15. Hugging Face Daily Papers TIER_1 English(EN) ·

    OOM-RL II:现实是神谕而非调试器——持续演进的代理工程系统中的来源约束诊断

    Reality may establish that an outcome occurred without identifying which evolving procedure produced it or why. This distinction matters in production ML systems whose code, configuration, and artifacts change while external feedback accumulates. We examine it in a human-directed…

  16. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Sung-eui Yoon ·

    通过支持保留蒸馏实现多智能体协调

    Offline MARL increasingly relies on generative policies to model multimodal joint behavior, typically by distilling a centralized teacher into decentralized one-step actors under the CTDE. We identify a failure mode at the teacher training stage: standard flow-based teachers pair…

  17. arXiv cs.AI TIER_1 English(EN) · Wenxuan Wang, Zekai Liu, Weinan Zhang, Yu Cheng, Yang Yang ·

    通过 On-Policy Context Distillation 将 Agent 体验内化到 Diffusion Model 权重中

    arXiv:2610.07250v1 Announce Type: new Abstract: Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continu…

  18. arXiv cs.AI TIER_1 English(EN) · Bowen Ye, Yongchao Xu, Junkai Ma, Xiang Yin, Wenzhao Li ·

    指导原则,启发行动:通过知识抽象实现智能体演进

    arXiv:2610.06964v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated strong capabilities in interactive environments, yet their ability to continually evolve from experience remains limited. Although fine-tuning enables adaptation, its dependence on…

  19. arXiv cs.AI TIER_1 English(EN) · Zhe Yu, Zixuan Wang, Peidong Wang, Hehai Lin, Ruochen Zhao, Chengwei Qin ·

    超越纠正记忆:多智能体系统中的执行一致性

    arXiv:2610.08101v1 Announce Type: new Abstract: Shared memory coordinates agents' actions, but correct records do not establish that those actions satisfy task requirements. Memory governance and failure diagnosis regulate or inspect recorded information; they do not by themselve…

  20. arXiv cs.AI TIER_1 English(EN) · Qianhan Feng, Zhongzhen Huang, Yakun Zhu, Xiaofan Zhang, Qi Dou ·

    从修订后果中学习:事后元经验蒸馏用于自改进代理

    arXiv:2610.07979v1 Announce Type: new Abstract: As agents continuously improve by generating and revising Skills, the process that discovers and refines those Skills becomes a learnable object in its own right. Task-Skills directly act on task execution, whereas Meta-Skills gover…

  21. arXiv cs.AI TIER_1 English(EN) · Yeji Park, Jaeyun Shim, Taesik Gong ·

    智能体能为所有人工作吗?个性化用户界面中跨用户移动GUI智能体的可靠性

    arXiv:2610.07972v1 Announce Type: new Abstract: Mobile GUI agents increasingly operate on interfaces influenced by users' histories and preferences, but their reliability across different users remains underexplored. We introduce PAIR (Personalized Application-state Instantiation…

  22. arXiv cs.AI TIER_1 English(EN) · Qi Cheng, Shengyu Chen, Wei Cheng, Zhengzhang Chen, Xiaowei Jia, Haoyu Wang, Haifeng Chen ·

    WorkflowOps: 学习代理协作先验知识以实现多代理工作流编排

    arXiv:2610.07860v1 Announce Type: new Abstract: Multi-agent systems are increasingly deployed for complex knowledge work, yet their orchestration layers remain largely memoryless: each new task is decomposed, assigned, and executed from scratch with no benefit from prior successf…

  23. arXiv cs.AI TIER_1 English(EN) · Qi Cheng, Shengyu Chen, Wei Cheng, Yiqun Xie, Xiaowei Jia, Haoyu Wang, Haifeng Chen ·

    OOPMAS:面向对象的查询级工作流生成多智能体系统

    arXiv:2610.07787v1 Announce Type: new Abstract: Multi-agent systems (MAS) powered by large language models have shown strong performance across code generation, mathematical reasoning, and question answering. However, existing methods for automating MAS design mostly operate at t…

  24. arXiv cs.AI TIER_1 English(EN) · Qi Cheng, Rongchao Dong, Shengyu Chen, Licheng Liu, Dan Lu, Zhengzhang Chen, Wei Cheng, Yiqun Xie, Haifeng Chen, Xiaowei Jia, Haoyu Wang ·

    ST-Bench:用于科学研究任务的多智能体系统生成空间时间基准

    arXiv:2610.07763v1 Announce Type: new Abstract: The rapid progress of LLM-based multi-agent systems (MAS) has shown that they largely outperform single agents on coding, math, and QA tasks, where executable tests provide a binary success signal. Whether this advantage transfers t…

  25. arXiv cs.AI TIER_1 English(EN) · Fouad Bousetouane ·

    EIO-Agents:AI代理评估缺失的语义层

    arXiv:2610.07675v1 Announce Type: new Abstract: AI agents are entering production in increasingly consequential environments without a shared semantic standard for what their evaluations actually mean. Scores, traces, judge outputs, and multi juror findings are increasingly used …

  26. arXiv cs.AI TIER_1 English(EN) · Jianglin Qiao, Siyi Hu, Thien Hoang Nguyen, Zehong Cao, Salah Sukkarieh ·

    与未来合作者协作:交错参与下的多智能体强化学习

    arXiv:2610.07578v1 Announce Type: new Abstract: In cooperative Multi-Agent Reinforcement Learning (MARL), agents are often trained under concurrent participation, while in many tasks some agents act earlier and leave task-relevant information that becomes useful to agents partici…

  27. arXiv cs.AI TIER_1 English(EN) · Xinle Wu, Yao Lu ·

    解耦多智能体编排

    arXiv:2610.07556v1 Announce Type: new Abstract: Learned orchestration can automatically construct effective language-model multi-agent systems, but existing approaches couple planning to fixed worker pools and train decomposition and collaboration from the same terminal outcome, …

  28. arXiv cs.AI TIER_1 English(EN) · Mohammadreza Sediqin, Shivali Dalmia, Srinivasa Karthikeya Reddy Kovvuri, Abhishek Mukherji ·

    用于 Agent 评估的信任层

    arXiv:2610.07274v1 Announce Type: new Abstract: Deterministic benchmark scores show that an agent received credit, but not whether that credit was earned, reported honestly, or would hold on a second run. We introduce a Trust Layer for Agent Evaluation, an additive post-hoc frame…

  29. arXiv cs.AI TIER_1 English(EN) · Sumanyu Muku ·

    验证并行编码代理中的协调:NP-Bench 和调度规划器

    arXiv:2610.07261v1 Announce Type: new Abstract: A team of coding agents can look fine agent by agent yet fail as a team: each passes its own tests while the merged result is broken, and single-agent evaluation never catches it. As teams run several LLM coding agents in parallel o…

  30. arXiv cs.AI TIER_1 English(EN) · Sumanyu Muku ·

    MemMux:并行编码代理集群的运行时验证和诚实资源归属

    arXiv:2610.07257v1 Announce Type: new Abstract: Developers increasingly run a fleet of coding agents side by side on one workstation. The tools they reach for, terminal multiplexers like tmux and a new generation of agent managers, were built to arrange windows, not to govern mem…

  31. arXiv cs.AI TIER_1 English(EN) · Sen Zhao, Jia Tang, Ruiqi Kong, Zuyu Zhang, Lifeng Shen, Ding Zou, Xinyu He, Xu Zhang, Junwei Han ·

    面向基于LLM的智能体的拓扑一致性任务规划在蜂窝工作流复合物上的应用

    arXiv:2610.07004v1 Announce Type: new Abstract: Task planning for LLM agents requires workflows that satisfy both user intent and complex sub-task dependencies. While existing planners work well for sequential or directed acyclic graph (DAG)-like structures, they struggle with wo…

  32. arXiv cs.LG TIER_1 English(EN) · Hyesung Jeon, Hyeongju Ha, Seoyoung Lee, Beomseok Kang, Jae-Joon Kim ·

    KVCMAS:多智能体系统中共享上下文的高效KV缓存校正

    arXiv:2609.34060v2 Announce Type: replace Abstract: Prompt-specialized multi-agent systems enable multiple agents to share a model while performing complementary roles to solve complex tasks. However, agent-specific prefixes change the KV cache generated for the same shared conte…

  33. arXiv cs.LG TIER_1 English(EN) · Jia Liufu, Bin Hu, Linglin Jing, Terry Kong, Yuki Huang, Ashwath Aithal, Wenming Yang, Jun Yang ·

    FC-SWE:面向长时域软件工程智能体的故障条件强化学习

    arXiv:2610.07898v1 Announce Type: new Abstract: Repository-level software engineering (SWE) is a challenging long-horizon setting: agents must reason over extended interactions, use tools, and adapt to stateful environments. Recent work trains SWE agents with reinforcement learni…

  34. arXiv cs.LG TIER_1 English(EN) · Lan Shi, Daigo Shishika, Xuan Wang ·

    通过有限深度策略敏感性适应智能体行为变化

    arXiv:2610.07475v1 Announce Type: new Abstract: Adapting a reinforcement learning policy to changes in another agent's behavior typically requires a large amount of new interaction data. Policy sensitivity provides a first-order prediction of how a locally optimal policy changes …

  35. arXiv cs.LG TIER_1 English(EN) · Ziyang Cai, Christos Ziakas, Vasilis Kontonis, Tim Pearce, Siddhartha Sen, Akshay Krishnamurthy, Shivam Garg, Dimitris Papailiopoulos ·

    Fork-and-Flush:在自主研究代理中逃离思想盆地

    arXiv:2610.07447v1 Announce Type: new Abstract: Autoresearch agents tackle open-ended problems by repeatedly proposing candidate solutions, evaluating them, and using feedback to guide subsequent experiments. We show that independent runs of the same agent on the same task often …

  36. arXiv cs.AI TIER_1 English(EN) · Haoyue Yang, Jingyao Li, Zhengfan Wu, Jing Liu, Xuanle Zhao, Kang Liu ·

    GAMEGO:利用锚定于真实世界资产的合成轨迹训练游戏开发代理

    arXiv:2610.06910v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have demonstrated remarkable capabilities in web front-end execution, with browser-based game generation emerging as a particularly prominent frontier. While previous efforts frequentl…

  37. arXiv cs.CL TIER_1 English(EN) · Shivani Kumar, Adarsh Bharathwaj, David Jurgens ·

    合作画像预测多智能体LLM团队在AI for Science工作流中的表现

    arXiv:2604.20658v2 Announce Type: replace Abstract: Multi-agent systems built from teams of large language models (LLMs) are increasingly deployed for collaborative scientific reasoning and problem-solving. These systems require agents to coordinate under shared constraints, such…

  38. arXiv cs.CL TIER_1 English(EN) · Naoki Wake, Justin Wagle ·

    SharedKV-BT:行为树代理的节点本地类型化决策

    arXiv:2610.07327v1 Announce Type: cross Abstract: Agent tasks require sequences of interdependent decisions. Autoregressive models support more flexible decision interfaces than conventional classifiers but incur the latency of token-by-token generation. Recent shared-prefix meth…

  39. arXiv cs.CL TIER_1 English(EN) · Lasse B. Strand, Robert Jakob, Kevin O'Sullivan, Markus Kreft ·

    Agentic AutoRAG:通过推理驱动的代理优化 RAG 管道

    arXiv:2610.08452v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacti…

  40. arXiv cs.AI TIER_1 Norsk(NO) · Justin Chih-Yao Chen, Elias Stengel-Eskin, Yan Chen, Pol Llado, Scott Counts, Mohit Bansal, Benjamin Van Durme, Harsh Jhamtani, Gaurav Verma ·

    TeleTune:从离线遥测数据中演进智能体技能

    arXiv:2610.05437v2 Announce Type: replace Abstract: Computer-use agents need to capture procedural knowledge of how people use software. User telemetry offers a scalable source of this knowledge. However, learning reusable skills from these logs requires addressing three challeng…

  41. arXiv cs.AI TIER_1 English(EN) · Yingying Liu, Junzhou Fang, Chenxiong Qian ·

    当代理上下文过时:易变代理上下文中的不连贯性

    arXiv:2610.05281v2 Announce Type: replace Abstract: Modern agents increasingly ground their reasoning in observations returned by tools, such as file contents read from a workspace. However, the data sources underlying these observations may later be modified by users, other agen…

  42. arXiv cs.AI TIER_1 English(EN) · Xiaoyu Chen, Lai Wei, Jin Wang, Xiangyu Zou, Ruochen Fan, Enze Luo, Mingzhe Yao, Jiahui Zhu, Yuhua Wen, Linghe Kong, Weiran Huang ·

    SWE-Game:编码代理能否构建我们想要的游戏?

    arXiv:2609.33678v2 Announce Type: replace Abstract: We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design docu…

  43. arXiv cs.AI TIER_1 English(EN) · Jiaxuan Dai, Tianyi Huang ·

    TwinCheck:基于证据的负面双重验证,用于有状态工具代理

    arXiv:2609.26911v2 Announce Type: replace Abstract: A single locally plausible tool call can derail an otherwise successful agent trajectory. Suspicion alone does not justify intervention, because the replacement itself can introduce the very failure verification is meant to prev…

  44. arXiv cs.AI TIER_1 English(EN) · Sarim Hashmi, Mukul Ranjan, Kshitij Mishra, Mikhail Kuznetsov, Praneeth Vepakomma, Nils Lukas ·

    AdvSim2Real:在 Web 世界模型中针对自适应提示注入训练 Web 代理

    arXiv:2610.08773v1 Announce Type: cross Abstract: Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the …

  45. arXiv cs.AI TIER_1 English(EN) · Zihan Zhou, Xinzhe Hu, Hanxu Yang, Liangjian Wen, Zhao Kang ·

    通过 System One 指导的计算分工实现高效率的多智能体协作

    arXiv:2610.08155v1 Announce Type: cross Abstract: Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by enabling collaborative problem solving among specialized agents. However, existing …

  46. arXiv cs.AI TIER_1 English(EN) · Kavienan Jegatheesan, Gayathri Lihinikaduarachchi ·

    当工具撒谎时:数学代理在工具反馈损坏下的可靠性

    arXiv:2610.08097v1 Announce Type: cross Abstract: Mathematical problem solving often requires deterministic computational steps that agents delegate to tools and implicitly trust. Yet tools can fail silently, returning plausible but incorrect results. How well can agents detect a…

  47. arXiv cs.AI TIER_1 English(EN) · Hongzhan Lin, Shidong Cao, Ziyang Luo, Wenhao Chai, Mong-Li Lee, Wynne Hsu ·

    从证据到行动:工具使用型智能体为何会失败

    arXiv:2610.07753v1 Announce Type: cross Abstract: Tool-using agents make consequential changes to external state, yet correct outcomes do not guarantee that their actions were supported by evidence established beforehand. We study where this evidence-to-action chain breaks as age…

  48. arXiv cs.AI TIER_1 English(EN) · Hang He, Li Wang, Hao Chen, Yuchen Shao, Yuling Shi, Lisheng Wang, Peiyang Liu, Goose Lin, Zaiyuan Wang, Haiying Sun, Ting Su, Chengcheng Wan ·

    CheckerBench:长视界代理能否合成静态分析检查器?

    arXiv:2610.07557v1 Announce Type: cross Abstract: Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing co…

  49. arXiv cs.AI TIER_1 English(EN) · Celal Ziftci, Spencer Greene, Ray Liu, Livio Dalloro, Lorenzo Dini ·

    在谷歌规模下,低延迟的智能体程序修复,捕捉开发者的工作流

    arXiv:2610.07289v1 Announce Type: cross Abstract: Manual repair of program failures is time-consuming and disruptive for software developers, particularly during the pre-submit phase where test failures occur within continuous integration systems. While Automated Program Repair h…

  50. arXiv cs.AI TIER_1 English(EN) · Tao Long, Lydia B. Chilton ·

    SPEAR:人机对齐的五项原则

    arXiv:2610.07204v1 Announce Type: cross Abstract: Recent AI alignment work often frames alignment as a pre-deployment optimization problem: collect human feedback, learn preferences or principles, finetune the model, and deploy an aligned system. This framing has produced major p…

  51. arXiv cs.AI TIER_1 English(EN) · Siru Jiang, Yongzhe Lyu, Shuo Lu, Yubin Wang, Yuxiang Zhang, Yue Liao, Bin Wang, Jian Liang, Tieniu Tan ·

    WorldSolver:LLM智能体能否通过生成求解器来模拟物理动力学?

    arXiv:2610.08720v1 Announce Type: new Abstract: LLM-based agents are increasingly advancing scientific and engineering problem solving, with physics simulation emerging as a challenging yet practical testbed for reproducing complex physical phenomena with application in embodied …

  52. arXiv cs.AI TIER_1 English(EN) · Mingda Zhang, Wenjin Liu, Tiesunlong Shen, Zikai Xiao, Zhenghong Lin, Qing Xu, Erik Cambria, Xiaoying Tang, Haoran Luo ·

    ScienceClaw:跨越自然科学与社会科学的科学人工智能代理的持续自我演化基准测试

    arXiv:2610.08691v1 Announce Type: new Abstract: Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both th…

  53. arXiv cs.AI TIER_1 English(EN) · Hanjun Luo, Xiucheng Zhang, Zhuoning Xu, Zhimu Huang, Yingbin Jin, Xinfeng Li, Hanan Salam ·

    ParanoiaEval:评估代理编码中不必要的防御性工作

    arXiv:2610.08662v1 Announce Type: new Abstract: As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a …

  54. arXiv cs.AI TIER_1 English(EN) · Yunbo Long, Guangya Hao, Yuhan Liu, Yiting Duan, Longyan Tan, Yunchen Long, Hao Wu ·

    Coding Agent的自我纠错应有多少证据?自蒸馏的自适应狄利克雷证据

    arXiv:2610.08514v1 Announce Type: new Abstract: Execution feedback lets coding agents revise programs and learn from their own corrections. A correction's learning weight should reflect both the transitions supported by its executions and the amount of evidence behind that suppor…

  55. arXiv cs.AI TIER_1 English(EN) · Hyun Jung Lee, Jungtaek Kim, Jongwon Jeong, Tae-Eui Kam, Donghyun Kim, Yong Jae Lee ·

    EMHO:通过经验轨迹实现具身智能体强化优化

    arXiv:2610.08432v1 Announce Type: new Abstract: Improving embodied agents often focuses on optimizing the underlying model through training, while the surrounding agent harness that controls planning, context, and tool use is typically engineered. We ask whether this harness can …

  56. arXiv cs.AI TIER_1 English(EN) · Haotian Chen, Shuaicheng Niu, Haocong Rao, Kaisong Song, Jun Lin, Lizhen Cui, Zhiqi Shen, Yonghui Xu ·

    面向长时域法律推理的测试时智能体演化

    arXiv:2610.08138v1 Announce Type: new Abstract: Legal intelligence aims to support reliable decision-making across long-horizon legal processes involving evolving case states and multiple roles. However, real-world legal deployment exhibits substantial case heterogeneity in facts…

  57. arXiv cs.AI TIER_1 English(EN) · Langxi Huang, Pingping Zhang, Lanyun Zhu, Chunyang Jiang, Jiawei Shao, Haocheng Yuan, Peilin Chen ·

    ChartBmkAgent:基于稀疏错误分类规范的、由Harness治理的多智能体图表问答基准构建

    arXiv:2610.08106v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) advance rapidly, while conventional benchmark development lags behind, delaying investigation of newly observed capability gaps. Such investigation requires an expressive task format and an o…

  58. Hugging Face Daily Papers TIER_1 English(EN) ·

    RoboQuest:通用物理代理,用于搜索、检查和测试

    Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-relevant information through interaction when it is ab…

  59. Hugging Face Daily Papers TIER_1 English(EN) ·

    SWE-Game:编码代理能否构建我们想要的游戏?

    We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fau…

  60. Hugging Face Daily Papers TIER_1 English(EN) ·

    我们查询,故我们计算:论机器之外的Oracle计算及其在智能体上的应用

    Agentic systems use large language models (LLMs) to carry out concrete tasks. Prior work often borrows abstractions such as scheduling, caching or isolation piecemeal from operating systems, so the mechanisms it builds share little common ground, and the shared view of the two fo…

  61. Hugging Face Daily Papers TIER_1 English(EN) ·

    从帕累托到偏好:通过摊销代理策略发现实现个性化测试时缩放

    Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--…

  62. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Veronique Ziegler ·

    当州长成为干扰:受控生成的干扰与有成本意识的退避在受治理的工具使用代理中

    Supervisory governors can interfere with the tool-using agents they regulate. We study this possibility in a controlled file-recovery environment where increases in regulatory intensity trigger experimentally imposed tool failures. A cost-blind governor can turn these failures in…

  63. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Hariganesh Tangirala ·

    在多智能体游戏中学习报告不安全任务

    When agents share a reward for completed tasks, reporting unsafe work can reduce the reporter's reward by stopping a task. Audits can make reporting optimal without ensuring that further training teaches a silent team to report. We study this learning problem in a game where any …

  64. Hugging Face Daily Papers TIER_1 English(EN) ·

    AdvSim2Real:在 Web 世界模型中针对自适应提示注入训练 Web 代理

    Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task r…

  65. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Markus Kreft ·

    Agentic AutoRAG:通过推理驱动的代理优化 RAG 管道

    Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacting choices, from chunking and embedding model to…

  66. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Zhao Kang ·

    通过 System One 指导的计算分工实现的高效 Token 多智能体协作

    Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by enabling collaborative problem solving among specialized agents. However, existing MAS frameworks tightly couple task reasoning with …

  67. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越纠正记忆:多智能体系统中的执行一致性

    Shared memory coordinates agents' actions, but correct records do not establish that those actions satisfy task requirements. Memory governance and failure diagnosis regulate or inspect recorded information; they do not by themselves establish whether it is sufficient to judge ta…

  68. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Dongha Lee ·

    从交付到状态化探索:重新思考 Agentic 搜索的索引

    Recent advances in agentic search have given large language model (LLM) agents finer control over corpus exploration. However, search interfaces often return matching passages even when feedback about the candidate set would suffice for the next decision, coupling candidate refin…

  69. Hugging Face Daily Papers TIER_1 English(EN) ·

    OOPMAS:面向对象的查询级工作流生成多智能体系统

    Multi-agent systems (MAS) powered by large language models have shown strong performance across code generation, mathematical reasoning, and question answering. However, existing methods for automating MAS design mostly operate at the task level, producing a single fixed workflow…

  70. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Fouad Bousetouane ·

    EIO-Agents:AI Agent 评估缺失的语义层

    AI agents are entering production in increasingly consequential environments without a shared semantic standard for what their evaluations actually mean. Scores, traces, judge outputs, and multi juror findings are increasingly used to justify readiness and release decisions, yet …

  71. Hugging Face Daily Papers TIER_1 English(EN) ·

    CheckerBench:长视界代理能否合成静态分析检查器?

    Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch…

  72. Hugging Face Daily Papers TIER_1 English(EN) ·

    从证据到行动:工具使用型智能体为何会失败

    Tool-using agents make consequential changes to external state, yet correct outcomes do not guarantee that their actions were supported by evidence established beforehand. We study where this evidence-to-action chain breaks as agents move from deciding whether to act to executing…

  73. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Lydia B. Chilton ·

    SPEAR:人机对齐的五项原则

    Recent AI alignment work often frames alignment as a pre-deployment optimization problem: collect human feedback, learn preferences or principles, finetune the model, and deploy an aligned system. This framing has produced major progress, but it under-specifies what happens once …

  74. Hugging Face Daily Papers TIER_1 English(EN) ·

    T-Search:一个开放的代理检索器和硬多步搜索的游乐场

    We present T-Search, an open-weight agentic retriever for hard multi-step search. Given a question and a search tool over a fixed corpus, it runs a bounded multi-round search and returns a ranked list of evidence chunks with short justifications, leaving answer generation to a do…

  75. arXiv cs.AI TIER_1 English(EN) · Chengyang Shi, Xianglin Ji, Jintao Huang, Jicheng Wang, Yifeng He, Jiachen Liu ·

    Open-Endedness Bench: 衡量来自代理记录的认知过程

    arXiv:2610.02588v1 Announce Type: new Abstract: Agents are increasingly given open-ended research tasks: discovering an empirical law from self-designed experiments, improving a heuristic whose optimum nobody knows, or beating a standing record. Their execution logs record every …

  76. arXiv cs.AI TIER_1 English(EN) · Eray Turkel, Mengsha Sun, Kartik Ayyar, Sean Dunigan, Jack Lu, Vlad Shcherban, Hsiang-Shun Shih, Xin Wang, Tiantian Zhang ·

    OpenGameEval:在有状态游戏引擎中对代理式编程和探索进行基准测试

    arXiv:2610.02563v1 Announce Type: cross Abstract: We present OpenGameEval, a benchmark and evaluation framework for agentic game development inside Roblox Studio. It runs language models as agents in reproducible, stateful game-engine sessions and scores each run with executable …

  77. arXiv cs.AI TIER_1 English(EN) · Songtao Wei, Yi Li, Zhichun Guo, Bingzhe Li ·

    Inherit-MAS:通过工作流和执行继承实现多智能体系统的测试时演化

    arXiv:2610.02396v1 Announce Type: cross Abstract: Multi-agent systems (MAS) built from large language models coordinate specialized agents to tackle complex tasks, but effective workflows are difficult to design in advance. Test-time evolution refines workflows using execution fe…

  78. arXiv cs.AI TIER_1 Dansk(DA) · A. Said Gurbuz, Ahmed Nassar, Sunghwan Hong, Marc Pollefeys, Peter W. J. Staar ·

    DeskForge:来自桌面环境的密集监督用于计算机使用代理

    arXiv:2610.02320v1 Announce Type: cross Abstract: Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such s…

  79. arXiv cs.AI TIER_1 English(EN) · Yu Li, Guangfeng Cai, Long-Fei Li, Shuo Han, Shengtian Yang, Han Luo, Kaibing Yang, Lei Feng ·

    功归所依:面向终端代理的依赖感知策略优化

    arXiv:2610.03634v1 Announce Type: new Abstract: Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier command…

  80. arXiv cs.AI TIER_1 English(EN) · Jermyn Zhen Yong Bek, Zhuang Qiang Bok, Zhongtian Sun ·

    知识还是计算器?分解可验证金融代理工作流中的技能溢价

    arXiv:2610.03564v1 Announce Type: new Abstract: Financial AI agents must do more than retrieve facts: investment workflows require correct quantitative execution, reliable use of procedural resources, and auditable structured outputs. We introduce FinSkillBench, an evaluation sui…

  81. arXiv cs.AI TIER_1 English(EN) · Chiara Troiani, Arash Salarian, Majed El Helou, Benjamin Ryder, Jean Diaconu, Herv\'e Muyal, Marcelo Yannuzzi ·

    迈向基于SLM的代理任务工具意图匹配

    arXiv:2610.03213v1 Announce Type: new Abstract: Tool-equipped AI agents use tool calls to access data and act on external systems. Horizontal growth of agentic systems increases the number of these interactions, and further motivates the need for automated, per-call oversight tha…

  82. arXiv cs.AI TIER_1 English(EN) · Yulong Ming, Jie Xu, Zihan Wu, Xiaohua Jia ·

    何时编译计算机使用代理?衡量投资回报并为代币效率做出编译决策

    arXiv:2610.02932v1 Announce Type: new Abstract: Compiling GUI procedures that agents execute repeatedly into programs can reduce their token costs. However, measuring payback and deciding when to compile have two challenges. First, compilation costs are uncertain because attempts…

  83. arXiv cs.AI TIER_1 English(EN) · XinPeng Shen, Lan Zhang, Yixiao Huang, Haoran Cheng, Jiewei Lai, Leilei Chen, Haoxiang Deng ·

    长时域Agent中的“幽灵”:跨回合被忽视的安全约束带来的治理风险

    arXiv:2610.02664v1 Announce Type: new Abstract: Long-horizon agents are now playing an increasingly significant role in assisting humans with complex problem-solving. However, it is exactly their extended interaction history that introduces an underexplored execution-safety conce…

  84. arXiv cs.AI TIER_1 English(EN) · Ankur Samanta, Yonathan Efroni, Paul Sajda, Kaveh Hassani, Anirudh Goyal ·

    学习下一步调查什么:用于长周期研究代理的元推理

    arXiv:2610.02525v1 Announce Type: new Abstract: Long-horizon research agents must decide both how to investigate and what to investigate next as evidence accumulates. This is hard to learn because such decisions are sparse in long execution traces, and their consequences may emer…

  85. arXiv cs.AI TIER_1 English(EN) · Yu Li, Zheng Zhang, Xin Liu, Shengtian Yang, Guangfeng Cai, Lei Feng ·

    行动前的选择:面向长时域工具使用智能体的比较价值估计

    arXiv:2610.02330v1 Announce Type: new Abstract: Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provid…

  86. arXiv cs.AI TIER_1 English(EN) · Xi Qin, Isabel Kurth, Xin Cui, Elin Park, Alexander Schaefer, Yaad Oren ·

    当终端代理训练停滞时:揭秘数据生成与验证挑战

    arXiv:2610.02405v1 Announce Type: new Abstract: Using a frontier model like Claude Opus as a meta-agent to generate terminal tasks and verifiers for RL training is increasingly common. Yet a runnable Docker image and executable test suite do not guarantee a faithful end-to-end pi…

  87. arXiv cs.LG TIER_1 English(EN) · Shangyang Wu, Shuai Zhao, Ziyue Zhu, Jinyang Wu, Anh Tuan Luu, Haoran Luo ·

    SCAD:面向长时域智能体的结构化信用分配与蒸馏

    arXiv:2610.03372v1 Announce Type: new Abstract: Training long-horizon agents to solve complex tasks requires effective supervision over extended interaction sequences. However, sparse terminal rewards obscure intermediate contributions, while on-policy distillation can lose infor…

  88. arXiv cs.LG TIER_1 English(EN) · Dongsu Lee, Haoran Xu, Amy Zhang ·

    测试时多智能体分解价值梯度流协调

    arXiv:2610.02554v1 Announce Type: new Abstract: Offline multi-agent reinforcement learning (MARL) faces a persistent trade-off. Expressive generative policies can represent multi-modal coordination in the data, but cannot distinguish high-value regions, while value-optimized poli…

  89. arXiv cs.LG TIER_1 English(EN) · Pranay Kothari ·

    ArrivalBench:由代理生成的数据管道一次性正确,但随时间推移而错误

    arXiv:2610.02363v1 Announce Type: new Abstract: Benchmarks for agent-generated data work grade a pipeline by running it once against a fixed snapshot. ArrivalBench instead re-executes the pipeline an agent leaves behind under adversarial but replayable delivery schedules (late, d…

  90. arXiv cs.AI TIER_1 English(EN) · Jugal Gajjar ·

    修复前验证:用于可信跨语言代码分析的智能执行基础

    arXiv:2604.10800v2 Announce Type: replace-cross Abstract: Learned classifiers deployed in agentic pipelines face a fundamental reliability problem: predictions are probabilistic inferences, not verified conclusions, and acting on them without grounding in observable evidence lead…

  91. arXiv cs.AI TIER_1 English(EN) · Sri Vatsa Vuddanti, Satwik Kumar Chittiprolu ·

    可恢复性有法可循:面向工具增强型Agent的ERR度量

    arXiv:2601.22352v2 Announce Type: replace-cross Abstract: Language model agents often appear capable of self-recovery after failing tool call executions, yet this behavior lacks a formal explanation. We present a predictive theory that resolves this gap by showing that recoverabi…

  92. arXiv cs.AI TIER_1 English(EN) · Yaxiao Liu (PwC China AI Center), Pengbo Liu (PwC China AI Center), Yiwen Liu (PwC China AI Center), Yihua Guan (PwC China AI Center), Jiaxing Song (Tsinghua University) ·

    从迁移到校准:跨模型、司法管辖区和规模保留代理能力

    arXiv:2609.35149v2 Announce Type: replace Abstract: Deploying, migrating, or scaling an agent can change its model, harness, infrastructure, application, and intended users. We formulate agent calibration as standards-first adaptation: define basic-capability, technical-environme…

  93. arXiv cs.AI TIER_1 English(EN) · Shiyi Kuang, Xuemei Luo, Kun Liu, Junhai Li, Rui Tian, Feng Shi, Bo Shen, Nianyu Li, Dehui Li, Ping Chen ·

    EvoRiskBench:工作空间代理运行时安全风险的演进基准

    arXiv:2610.03153v1 Announce Type: cross Abstract: Workspace agents combine large language models with execution harnesses to perform stateful, multi-step tasks that access or modify external resources. Existing benchmarks leave gaps in executable coverage of their runtime securit…

  94. arXiv cs.AI TIER_1 English(EN) · Jabin Koo, Soheil Abbasloo, Sungjae Lee, Jungseul Ok ·

    面向多智能体系统的动态专家剪枝

    arXiv:2610.02951v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures scale language models efficiently by activating only a few experts per token, but the saving is confined to computation: every expert must stay resident on the accelerator, so memory bounds w…

  95. arXiv cs.AI TIER_1 English(EN) · Mayank Rathee, Alexander Stepanov, Shalin Madabhavi, Jinhao Zhu, Raluca Ada Popa, Ion Stoica ·

    Pincer:使用数字孪生进行代理的资源授权

    arXiv:2610.02569v1 Announce Type: cross Abstract: Coding agents have become increasingly long-horizon, autonomous, reliant on general-purpose shell and maintain their own persistent memory for self-improvement. While these capabilities have made the agents powerful, they have als…

  96. Hugging Face Daily Papers TIER_1 English(EN) ·

    通过 On-Policy Context Distillation 将 Agent 体验内化到 Diffusion Model 权重中

    Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continually construct and revise prompts, thereby elici…

  97. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Dileep Kalathil ·

    AgentDiscover:极简搜索脚手架下的自主发现

    Frameworks that use large language models for scientific discovery typically rely on a fixed, human-designed algorithm that decides what the model sees at each step, leaving the model only the role of proposer. The model knows nothing of the search beyond what it is shown. As mod…

  98. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Xiaohui Yan ·

    SearchJev:搜索代理的快速且经过校准的系统1模型

    Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates se…

  99. Hugging Face Daily Papers TIER_1 English(EN) ·

    RobotUse:分配计算、上下文和决策

    Robot agents must connect their intended actions to observed outcomes while retaining the context needed to revise their choices over repeated attempts. Existing interfaces often leave these choices inside predefined tools or require agents to manage detailed execution code and i…

  100. Hugging Face Daily Papers TIER_1 English(EN) ·

    ASCENT:通过经验验证的自蒸馏实现长时域智能体的在线测试时训练

    A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-context adaptation agents…

  101. Hugging Face Daily Papers TIER_1 English(EN) ·

    AutoSciBench:为评估科学智能体生成自主基准

    As agents rapidly evolve, existing benchmarks can become saturated, limiting their ability to distinguish capabilities and reveal remaining failure modes. Particularly in scientific domains, constructing and updating benchmarks requires substantial time, labor, and domain experti…

  102. Hugging Face Daily Papers TIER_1 English(EN) ·

    SearchJev:搜索代理的快速且经过校准的系统1模型

    Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates se…

  103. Hugging Face Daily Papers TIER_1 English(EN) ·

    Code2Games:赋能游戏世界生成的编码代理

    Generating a high-quality gaming world from a natural-language game intent requires joint reasoning about scene structure, spatial layout, gameplay objectives, interactive entities, and executable gameplay logic. Existing coding agents can generate individual assets, scenes, or s…

  104. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Hao Wang ·

    当辩论有益时:多智能体推理中的提议供给与验证感知读出

    Multi-agent debate can improve reasoning, yet often fails to beat simple majority voting. We argue that successful debate requires two distinct mechanisms: proposal supply must surface a correct answer, and readout must identify that answer when voting misses it. We formalize the…

  105. arXiv cs.LG TIER_1 English(EN) · Jose A. Ayala-Romero, Andres Garcia-Saavedra, Xavier Costa-Perez ·

    TRACE:通过代理启发式设计解决现实世界资源分配问题

    arXiv:2610.01887v1 Announce Type: cross Abstract: Dynamic resource assignment, the real-time allocation of task streams to heterogeneous processing nodes, is the backbone of modern computing infrastructure. While learning-based schedulers excel in research, industrial deployments…

  106. arXiv cs.LG TIER_1 English(EN) · Huancheng Chen, Xiaodi Sun, Zhaoqiong Huang, Shenyang Huang Shreya Singhal, Jingwen Lu ·

    SkillSpec:通过表征专业化实现共识门控的智能体技能演化

    arXiv:2610.00704v1 Announce Type: new Abstract: Natural-language skills are textual procedural memories through which large language model (LLM) agents retain reusable task knowledge without updating model weights. Existing methods typically treat skills as either static artifact…

  107. Hugging Face Daily Papers TIER_1 English(EN) ·

    LMBuild:评估用于生成可构建和功能性结构的 LLM Agent

    LLM-based agents are increasingly capable of generating complex 3D structures, with the potential to reshape how objects are designed and realized in the physical world. Yet, producing elegant geometry is fundamentally different from producing objects that can be built and perfor…

  108. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Hui Song ·

    工程化可持续智能体:面向开发者工作流的智能体大模型系统性比较

    Large language models (LLMs) are increasingly used in software engineering, including agentic systems that coordinate multiple agents, but impose higher computational and environmental costs. In this paper, we present a comprehensive empirical study of agentic LLM systems across …

  109. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Jungseul Ok ·

    面向多智能体系统的动态专家剪枝

    Mixture-of-Experts (MoE) architectures scale language models efficiently by activating only a few experts per token, but the saving is confined to computation: every expert must stay resident on the accelerator, so memory bounds where these models can be deployed. Expert pruning …

  110. arXiv cs.AI TIER_1 English(EN) · J\'er\'emie Lumbroso ·

    控制论与认识论:可信代理委托的缺失词汇

    arXiv:2610.00961v1 Announce Type: new Abstract: As code generation is increasingly delegated to AI systems, the bottleneck is shifting from writing code to supervising the systems that write it --- a shift CS-education researchers have begun to name. This shift exposes a vocabula…

  111. arXiv cs.AI TIER_1 English(EN) · Yezhou Cheng, Runjia Du, Zeming Liu, Hang Lyu, Zehua Yang, Bojun Lin ·

    何时让步:基于文本的具身智能体的用户纠正的地面仲裁

    arXiv:2610.00282v1 Announce Type: new Abstract: How should an embodied agent respond when a person's correction may be wrong? We formulate grounded correction arbitration as a choice among accepting, rejecting, inspecting the world, and asking the speaker. GAVA implements this in…

  112. arXiv cs.AI TIER_1 English(EN) · Michael Hardy, Ruhana Azam, Anka Reuel, Mykel Kochenderfer, Sanmi Koyejo ·

    Agent 评估可靠性:更多任务不(总是)能修复 Agent 排行榜

    arXiv:2610.00651v1 Announce Type: new Abstract: Agent evaluations are increasingly used to compare LLMs and inform deployment decisions, yet ranks can reflect not only the model but also the effects of the evaluation conditions such as the scaffolds or tasks. This makes reliabili…

  113. arXiv cs.AI TIER_1 English(EN) · Andre Fu, Malik Drabla, Leon Liu, Meji Abidoye, Marek Suppa, Lata Mishra, Adnan El Assadi, Yiyuan Li ·

    Incident-Arena:让智能体达到可靠性的最后九强

    arXiv:2610.00648v1 Announce Type: new Abstract: AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to their ability to execute on production incident response. This …

  114. arXiv cs.AI TIER_1 English(EN) · Sahan Paliskara, Nattaput Namchittai, Andrew Lampinen ·

    同室操戈:多用户多智能体团队中的性能如何下降

    arXiv:2610.00583v1 Announce Type: new Abstract: People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people's agents over shared resources such as a codebase, a calendar, or a budget. When each agent acts for a different user…

  115. arXiv cs.AI TIER_1 English(EN) · Haoyang Su, Weiran Huang ·

    JevSpawn:通过组合式动作空间实现自适应代理推理

    arXiv:2610.00437v1 Announce Type: new Abstract: LLM agents generate intermediate reasoning and actions token by token, making extended interactions slow and computationally expensive. Jev-style models offer fast probabilistic predictions over finite fields, but require those fiel…

  116. arXiv cs.LG TIER_1 English(EN) · Jiayi Yang, Yifang Chen, Yuanfu Sun, Xinyan Ge, Qiaoyu Tan ·

    GraphMAS:图学习多智能体协调的系统性基准测试

    arXiv:2609.39777v1 Announce Type: new Abstract: LLM-based multi-agent systems coordinate specialized reasoning through aggregation, interaction, and adaptive control, yet their potential for graph learning remains unexplored. Graph learning is a natural setting for such systems b…

  117. arXiv cs.LG TIER_1 English(EN) · Bo Han, Qianyi Wang, Shuai Liu, Xiong Zifan, Changqiao Wu, Yuanfa Li, Pengzhi Gao, Wei Liu, Jian Luan, Heng Qu, Yunpeng Song, Zhongmin Cai ·

    在不完美先验下学习可靠的GUI代理

    arXiv:2609.39547v1 Announce Type: new Abstract: GUI agents built on large language and vision-language models still struggle on unseen applications and complex multi-step tasks, as completing real GUI tasks depends on app-specific, temporally volatile operational knowledge that i…

  118. arXiv cs.AI TIER_1 English(EN) · Joachim Baumann, Vishakh Padmakumar, Xiang Li, John Yang, Diyi Yang, Sanmi Koyejo ·

    SWE-chat:来自真实用户的真实编码代理交互

    arXiv:2604.20779v2 Announce Type: replace Abstract: AI coding agents are being adopted at scale, yet we lack empirical evidence on how people actually use them and how much of their output is useful in practice. We present SWE-chat, the first large-scale dataset of real coding ag…

  119. arXiv cs.AI TIER_1 English(EN) · Arun Sharma ·

    Spatial Atlas:面向空间感知研究代理基准的计算基础推理

    arXiv:2604.12102v3 Announce Type: replace Abstract: We describe compute-grounded reasoning (CGR), a design pattern in which code computes selected sub-problems from explicit intermediate representations before a language model answers. Spatial Atlas implements CGR as an Agent2Age…

  120. arXiv cs.AI TIER_1 English(EN) · Weiyi Wang, Xinchi Chen, Jingjing Gong, Xuanjing Huang, Xipeng Qiu ·

    AstroAgentBench:在太空任务规划任务上评估Agentic规划

    arXiv:2601.11354v2 Announce Type: replace Abstract: Recent LLM-for-Space systems address mission planning, scheduling, operations support, simulator control, and autonomy, but their evaluations use different task contracts, control settings, simulators, and success criteria. We i…

  121. arXiv cs.AI TIER_1 English(EN) · Dayu Wang, Yutong Liu, Jiaye Yang, Weikang Li, Jiahui Liang, Yang Li ·

    通过多小型智能体强化学习降低工具使用的认知开销

    arXiv:2508.08882v5 Announce Type: replace Abstract: Recent advances in multi-agent systems highlight the potential of specialized small agents that collaborate via division of labor. Existing tool-integrated reasoning systems, however, often follow a single-agent paradigm in whic…

  122. arXiv cs.AI TIER_1 English(EN) · Yen-Jen Wang, Haozhe Jiang, Shuying Deng, Haoru Xue, Weirui Ye, Rocky Duan, Nika Haghtalab, S. Shankar Sastry, Pieter Abbeel, Haozhi Qi ·

    重建、练习、实操:具身智能体的引导式自我提升

    arXiv:2610.02204v1 Announce Type: cross Abstract: Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a …

  123. arXiv cs.AI TIER_1 English(EN) · Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma ·

    Argo-Bench:在企业级工作流上评估数据代理

    arXiv:2610.02122v1 Announce Type: cross Abstract: Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, …

  124. arXiv cs.AI TIER_1 Română(RO) · Rui Sun, Xihan Xiong, Qin Wang, Fei Gao, Zelin Li, Zehua Cheng, Jiahao Sun, Zhipeng Wang ·

    SoK:去中心化代理经济基础设施

    arXiv:2610.01756v1 Announce Type: cross Abstract: Decentralized agent economies increasingly build a single task from protocols that were designed and secured separately. This creates a simple problem: a workflow can look correct at each step and still produce the wrong outcome. …

  125. arXiv cs.AI TIER_1 English(EN) · Sushant Mehta, Logan Ritchie, Edwin Chen ·

    从 RL 在 Agentic 编码任务上的跨基准迁移

    arXiv:2610.00890v1 Announce Type: cross Abstract: Coding agents often fail in the last mile: they build most of a feature but drop a requirement, test only the cases their implementation already handles, break behavior that was supposed to stay intact, or validate against an unch…

  126. arXiv cs.AI TIER_1 English(EN) · Zhengyuan Jiang, Reachal Wang, Yuepeng Hu, Yupu Wang, Yuqi Jia, Neil Zhenqiang Gong ·

    AI 编码代理的自进化编码规则

    arXiv:2610.00650v1 Announce Type: cross Abstract: The performance of AI coding agents is highly dependent on their underlying coding rules. However, existing coding rules are typically hand-crafted and fixed, making the process labor-intensive and often suboptimal. In this work, …

  127. arXiv cs.AI TIER_1 English(EN) · Ayan Javeed Shaikh, Arunesh Sinha, Nathaniel D. Bastian, Ankit Shah ·

    没有一种架构适合所有情况:分层红队代理的跨环境评估

    arXiv:2610.00557v1 Announce Type: cross Abstract: Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks. Reinforcement learning (RL) and large language models (LLMs) offer complementary mechanisms for t…

  128. arXiv cs.AI TIER_1 English(EN) · Xin Heng ·

    全球一致性:当每个智能体都正确,但团队仍然出错——多智能体协作的局部到全局语义基础

    arXiv:2610.02036v1 Announce Type: new Abstract: AI agents can each make locally valid decisions yet jointly produce an invalid result. We call this the global coherence problem: a failure of shared state, not merely of model intelligence. Our Observation-Aliasing Impossibility Th…

  129. arXiv cs.AI TIER_1 English(EN) · Hao Wang, Ting Huang ·

    Mingbird:一个支持小型开放模型完成实际任务的本地优先代理框架

    arXiv:2610.02001v1 Announce Type: new Abstract: Small open-weight models (2-9B) run on ordinary laptops, but under cloud-scale agent harnesses they rarely complete real tasks: tool prefill overflows the context, self-correction diverges, tool demonstrations loop, and tasks are si…

  130. arXiv cs.AI TIER_1 English(EN) · Beining Wu, Zihao Ding, Jun Huang ·

    并非所有经验都属于权重:用于自改进 GUI 代理的组件路由

    arXiv:2610.01787v1 Announce Type: new Abstract: Self-improving GUI agents keep the trajectories they produce and return them to the agent, by fine-tuning or by retrieval into the prompt, and studies that compare the two destinations disagree. We attribute this to the unit of expe…

  131. arXiv cs.AI TIER_1 English(EN) · Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata ·

    智能体是系统,而非模型:重新思考智能体评估

    arXiv:2610.01618v1 Announce Type: new Abstract: Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users d…

  132. arXiv cs.AI TIER_1 English(EN) · Zongrui Yang, Li Xintong, Runchen Xu, Zhongsheng Wang, Zhedong Lin, Haoyuan Li, Jiamou Liu ·

    MCRI:一种分析和评估智能体技能的四维框架

    arXiv:2610.01506v1 Announce Type: new Abstract: As agents evolve from single-tool systems into modular, composite architectures, skills are becoming an important mechanism for capability development and distribution. However, the academic community lacks a structured framework fo…

  133. arXiv cs.AI TIER_1 English(EN) · Yan Luo, Selim-Antoine Lali, Jeremy Moebel, Iliass Khoutaibi, Ahmadou Aidara, Mengyu Wang ·

    面向动态推理的感知感知独立代理图

    arXiv:2610.01249v1 Announce Type: new Abstract: Conventional reasoning protocols present a fixed, preselected task, so they cannot test whether an agent propagates relevant updates, preserves unaffected work, or reconstructs a historical task binding. We therefore study \emph{dyn…

  134. arXiv cs.AI TIER_1 English(EN) · Yoonkyu Woo, Woojin Lee, Jin-Xia Huang ·

    YouRA:一种用于证据可追溯自主研究代理的持久化状态架构

    arXiv:2610.01097v1 Announce Type: new Abstract: End-to-end research agents can now produce complete scientific papers, yet manuscript claims often diverge from executed experiments. This gap is structural: research state, failure histories, and claim-evidence alignment are not ma…

  135. arXiv cs.AI TIER_1 English(EN) · Jingtan Wang, Sirajul Salekin, Young mok Jung, Javier Movellan, Bryan Kian Hsiang Low, Manjot Bilkhu ·

    RISED:用于智能体多环境选择和自蒸馏的RubrIcs

    arXiv:2610.00979v1 Announce Type: new Abstract: Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environ…

  136. arXiv cs.AI TIER_1 English(EN) · Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu, Nigel Collier, Tomas Pfister, Chen-Yu Lee ·

    VeriHarness:为长时任务扩展代理验证

    arXiv:2610.00972v1 Announce Type: new Abstract: As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to referenc…

  137. arXiv cs.AI TIER_1 English(EN) · Timothy Kassis ·

    科学智能体:在科学任务上评估特定职业的系统提示

    arXiv:2610.00084v1 Announce Type: new Abstract: Detailed profession-specific system prompts raise token use and estimated cost per response without a consistent accuracy gain. We evaluate Scientific Agents, an open-source corpus of 503 profession-specific AGENTS.md profiles, with…

  138. arXiv cs.AI TIER_1 English(EN) · Ronghua Li, Zi Liang, Zhishan Li, Shinan Liu ·

    PG-SFT:在离线代理微调中平衡能力获取与保留

    arXiv:2610.00949v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) on offline agent trajectories is the standard approach for training specialized tool-using agents, but forcing models to imitate reasoning and actions token by token may harm other capabilities (e.g., ge…

  139. arXiv cs.AI TIER_1 English(EN) · Xisen Jin, Jingheng Li, Zhenglun Chen, Junyi Du, Xiang Ren ·

    ReLiveGym: 在数周重放现实中评估长期代理

    arXiv:2610.00710v1 Announce Type: new Abstract: As large language model (LLM) agents become widely adopted, they are increasingly deployed for tasks that require persistent monitoring or recurring actions (e.g., market analysis). These agents are expected to operate unattended fo…

  140. Hugging Face Daily Papers TIER_1 English(EN) ·

    CUAWright:数字代理的最小统一接口

    The prevailing approach to computer-use agents couples a model with a domain-specific harness: a browser or desktop environment equipped with human engineered tools that are fixed before task execution. As models' coding capabilities improve, the GUI native and static harness pre…

  141. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Xavier Costa-Perez ·

    TRACE:通过代理启发式设计解决现实世界资源分配问题

    Dynamic resource assignment, the real-time allocation of task streams to heterogeneous processing nodes, is the backbone of modern computing infrastructure. While learning-based schedulers excel in research, industrial deployments still rely on hand-written rules that operators c…

  142. Hugging Face Daily Papers TIER_1 English(EN) ·

    TRACE:通过代理启发式设计解决现实世界资源分配问题

    Dynamic resource assignment, the real-time allocation of task streams to heterogeneous processing nodes, is the backbone of modern computing infrastructure. While learning-based schedulers excel in research, industrial deployments still rely on hand-written rules that operators c…

  143. arXiv cs.MA (Multiagent) TIER_1 Română(RO) · Zhipeng Wang ·

    SoK:去中心化代理经济基础设施

    Decentralized agent economies increasingly build a single task from protocols that were designed and secured separately. This creates a simple problem: a workflow can look correct at each step and still produce the wrong outcome. For example, a correct escrow may release payment …

  144. arXiv cs.AI TIER_1 English(EN) · Yuqing Zhai, Xiaohong Chen, Lingming Zhang, Sriram Vishwanath, Grigore Rosu ·

    从验证失败到可复用的代码代理指南

    arXiv:2609.39022v1 Announce Type: cross Abstract: Coding agents need to establish that a program satisfies a specification and that the specification captures the requested behavior. We study how expert diagnosis of verification failures can become reusable guidance for this work…

  145. arXiv cs.AI TIER_1 English(EN) · Hongjin Qian, Chaofan Li, Kun Luo, Wenqing Wei, Jianlyu Chen, Shuqi Lu, Yuyang Hu, Hongwang Xiao, Hui Wang, Chaozhuo Li, Qiwei Ye, Zhicheng Dou, Defu Lian, Zheng Liu ·

    AREX-2:通过长时程反思性任务推进自改进代理

    arXiv:2609.38288v1 Announce Type: new Abstract: We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, whi…

  146. arXiv cs.AI TIER_1 English(EN) · Hyeong Kyu Choi, Bhavana Dalvi Mishra, Jiefeng Chen, Mihir Parmar, Rui Meng, Chun-Liang Li, Xiangru Tang, Sharon Li, Jinsung Yoon, Tomas Pfister ·

    AIM: 智能体式想法管理,实现自动化研究

    arXiv:2609.38445v1 Announce Type: new Abstract: Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven search and identify three core challenges: organizing evolving research ideas, sele…

  147. arXiv cs.AI TIER_1 English(EN) · Mohamed Abouzahra ·

    NAQD Env:语言代理选择性撤回的基准测试

    arXiv:2609.38460v1 Announce Type: new Abstract: Language agents must revise planned actions when evidence changes, permission is revoked, or a stop instruction arrives. A useful response is selective: suspend affected actions, preserve unaffected work, and resume only after suffi…

  148. arXiv cs.AI TIER_1 English(EN) · Jeffrey Willette, Krishna C. Puvvada, Boris Ginsburg ·

    保持专注:测试长时程代理可靠性的基础

    arXiv:2609.38712v1 Announce Type: new Abstract: Long-horizon agentic workflows require models to sustain repeated state-dependent actions all while the context grows, sub-task complexity changes, and new data arrives. Each situation represents an independent axis along which an a…

  149. arXiv cs.AI TIER_1 English(EN) · Qiuhui Chen, Jiafan Lu, Shuaimin Tang, Tao Dai, Suyuan Wang, Chenrui Ji, Zhenglei Zhou, Weimin Zhong ·

    PathAnchor:为科学智能体提供路径结构化证据

    arXiv:2609.38766v1 Announce Type: new Abstract: Scientific agents can retrieve relevant passages yet still lose functional order, mix evidence across sources, or state conclusions that exceed the retrieved record. We introduce PathAnchor, a bounded scientific reasoning system bui…

  150. arXiv cs.AI TIER_1 English(EN) · Xinhe Tian, Xiaoyue Zhang, Ziyou Zhang, Jiacheng Li, Xiaoqiang Jin, Qianchuan Zhao, Gaochen Cui ·

    STRATA:通过角色对齐的分层代理进行实时策略游戏的自学习

    arXiv:2609.38881v1 Announce Type: new Abstract: Real-time strategy (RTS) games require agents to coordinate economic development, production and construction, base defense, unit organization, and attack timing over long matches. Existing studies have applied large language models…

  151. arXiv cs.AI TIER_1 English(EN) · Heng-Zhuang Li, Yi-Kai Zhang, Yu Wang, Yueqing Sun, Jiayuan Zhang, Qi Gu, Han-Jia Ye ·

    面向长时域Agent任务的一致性计划-行动

    arXiv:2609.38891v1 Announce Type: new Abstract: Long-horizon agentic tasks demand strong reasoning and efficient execution across successive interactions with dynamic environments. A common approach decouples high-level planning from low-level execution through separate planner a…

  152. arXiv cs.AI TIER_1 English(EN) · Peng Kuang, Haibo Jin, Dehao Wu, Feiyang Deng, Xiaopeng Yuan, Jerry Wang, Haohan Wang ·

    在测试时使用可重用原语组合特定任务的代理

    arXiv:2609.38912v1 Announce Type: new Abstract: Agent harnesses govern how large language models (LLMs) gather context, invoke tools, verify results, preserve state, and terminate, largely affecting agent performance. However, the value of each harness mechanism can differ across…

  153. arXiv cs.AI TIER_1 English(EN) · Guanning Zeng, Jiani Wang, Wenjie Ma, Shaofeng Yin, Chenyang Wang, Shichen Liu, Angjoo Kanazawa, Wode Ni, Xiuyu Li, Andrea Zanette, Haiwen Feng ·

    Schema:通过代理程序归纳发现未知环境

    arXiv:2609.39140v1 Announce Type: new Abstract: Learning to complete tasks in unfamiliar environments with unknown rules remains a key challenge for LLM agents. Current LLM agents often record their discoveries in prose, which may not provide a compact, explicit account of how th…

  154. arXiv cs.AI TIER_1 English(EN) · Xinyu Zhu, Fenyi Liu, Yuzhu Cai, Shuo Tang, Rui Ye, Linfeng Zhang, Siheng Chen ·

    WorkGenesis:构建训练智能体工作的世界

    arXiv:2609.39325v1 Announce Type: new Abstract: The ability of Large Language Model (LLM) agents to complete daily and professional work is receiving increasing attention. Training such agents requires realistic work scenarios. Expert-authored occupational work is costly and slow…

  155. arXiv cs.AI TIER_1 English(EN) · Seonho Lee, Wonryeol Jeong, Alberto Cereser, Inha Kang, Hyeonjong Kim, Seungmin Kwak, Dongmin Park ·

    A2Z GameSpec-Bench:编码代理能多忠实地根据游戏设计规范生成游戏?

    arXiv:2609.39564v1 Announce Type: new Abstract: Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form…

  156. arXiv cs.AI TIER_1 English(EN) · Fabio Rovai ·

    谁来验证图?语言代理因果行动验证中的错误指定攻击

    arXiv:2609.40027v1 Announce Type: new Abstract: Causal action verifiers gate an agent's state-changing tool calls by checking whether each proposed intervention is identifiable against a committed action-state graph, and they issue a certificate that carries the identification ar…

  157. arXiv cs.AI TIER_1 English(EN) · Yinghui He, Yapei Chang, Khushi Bhardwaj, Daniele Molinari, Tugrul Konuk, Jan Kautz, Ali Hatamizadeh ·

    PivotOPD:学习从多轮智能体中的关键性错误中恢复

    arXiv:2609.40285v1 Announce Type: new Abstract: On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the…

  158. arXiv cs.AI TIER_1 English(EN) · Chun-Wah Hsu, Kai Gong, Yu Wu, Xianhe Chen, Mengyang Liu, Jie Li, Hanyu Li, Zhixuan Liu, Naisheng Tang, Jiaying Chi, Ziheng Fan, Xuning He, Xiaokang Yang, Xue Jiang, Yihong Dong ·

    OpenCollab:一个具有可编程协作和可控运行时的多智能体编码框架

    arXiv:2609.38345v1 Announce Type: cross Abstract: Multi-agent coding systems are designed to tackle complex software engineering tasks through collaboration. However, existing evaluations typically assume configured organizations are followed faithfully, whereas reality differs. …

  159. arXiv cs.AI TIER_1 English(EN) · Xiao Huang, Mingda Zhang, Junming Zhang, Qiang Huang, Hanwen Zhang, Yue Dai, Zijia Wang, Xiaoying Tang ·

    CollabFlow:Agent协作的递归自我改进

    arXiv:2609.38662v1 Announce Type: cross Abstract: Recursive self-improvement (RSI) lets a system improve from its own outcomes; in LLM-based multi-agent systems, Agents refine one another within a task, and outcomes improve how they collaborate across tasks. However, existing mul…

  160. arXiv cs.AI TIER_1 Dansk(DA) · Guanqun Yang, Wenlong Zhang, Tian Shi, Ping Wang ·

    SkillSeek:在市场规模上重新审视智能体技能检索

    arXiv:2609.38822v1 Announce Type: cross Abstract: Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The stan…

  161. arXiv cs.AI TIER_1 English(EN) · Shijia Ge, Alex Zhou, Jianshu Zeng, Yexing Wan, Di Wu, Zelin Zheng, Yazhe Wang, Zhiqi Jia, Xuan Shangguan, Jay Zhu, Yijun Liu, Lingyu He, Sihang Wu, Xiao He, Hongcheng Gao ·

    让代码即政策再创辉煌:前沿智能体编写、调用并进化机器人工具

    arXiv:2609.39018v1 Announce Type: cross Abstract: Frontier models can control robots, but reasoning through every reach, grasp, and retreat makes manipulation slow and token-intensive. We revisit code as policy with a different division of labor: models build executable tools, co…

  162. arXiv cs.AI TIER_1 English(EN) · Abraham Yeung ·

    用于编码理论的编码代理

    arXiv:2609.39081v1 Announce Type: cross Abstract: We spent five weeks using an LLM coding agent on open problems in coding theory: finding large sets of four-letter words, such as DNA barcodes, that stay far apart in edit distance. The agent wrote the verifiers and search code; a…

  163. arXiv cs.AI TIER_1 English(EN) · Meijia Chen, Hao Li, Zheng Lu, Hongshan Lin, Junbai Tian, Yichen Liu, Zijun Tian, Yufan Zou, Shuhan Sun, Hanxin Chen, Zeyu Zhang, Weizhi Du, Yueting Li, Tianyu Shi, Alaa Khamis ·

    虚假前沿:诊断与缓解自进化搜索代理中的协同作弊

    arXiv:2609.39102v1 Announce Type: cross Abstract: Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer …

  164. arXiv cs.AI TIER_1 English(EN) · Hanwen Liu, Yuanfu Sun, Qiaoyu Tan ·

    DAGent:深度研究代理的评估后增长规划

    arXiv:2609.39154v1 Announce Type: cross Abstract: Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting bec…

  165. arXiv cs.AI TIER_1 English(EN) · Wenjin Wang, Jiazhen Lei, Yuxin Sha, Nuwa Xi, Meng Zhao, Xingxi Yin, Qi Liu, Yuliang Shen, Zixun Sun ·

    NarrativeSteward:在代理辅助交互式叙事创作中协调委托、指导和验证

    arXiv:2609.39333v1 Announce Type: cross Abstract: Autonomous AI agents can turn authors' goals into interactive narratives by independently organizing and carrying out generation and revision. As agents generate and revise extensive content, authors struggle to grasp its overall …

  166. arXiv cs.AI TIER_1 English(EN) · Frances Liu, Manny Silva, Paige Calvert, Ayu Adiati, Sarah Sanders ·

    DoGBench:代理能否达到用户文档专家的标准?

    arXiv:2609.39909v1 Announce Type: cross Abstract: We introduce DoGBENCH (Documentation Generation Benchmark), to our knowledge, the first benchmark for generating and maintaining real user-facing software documentation. It asks whether an agent can produce documentation that expe…

  167. arXiv cs.AI TIER_1 English(EN) · Qisheng Zhou, Zhen Xiong, Qiaoyu Tan ·

    TRACE:并行搜索代理的轨迹选择

    arXiv:2609.39912v1 Announce Type: cross Abstract: Parallel search may generate a correct answer that final-answer voting fails to select. We formulate this consolidation stage as trajectory selection and introduce TRACE (Trajectory Ranking with Aggregated Cross-Rollout Evidence),…

  168. arXiv cs.AI TIER_1 English(EN) · Jiangrui Zhao, Chenglong Li, Meng Zhang, Xiaoting Du ·

    学习何时以及如何干预:用于编码代理的后见之明蒸馏哨兵

    arXiv:2609.39957v1 Announce Type: cross Abstract: Coding agents solve repository-level tasks through sequences of actions, where a single erroneous action can misdirect subsequent decisions and increase recovery costs. Existing approaches use execution feedback for recovery or sp…

  169. arXiv cs.AI TIER_1 English(EN) · Kaixuan Fan, Kaituo Feng, Tianshuo Peng, Yilei Jiang, Manyuan Zhang, Junke Wang, Xiangyu Yue ·

    EviRover: 强化代理感知,超越一瞥

    arXiv:2609.40230v1 Announce Type: cross Abstract: Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumpt…

  170. arXiv cs.AI TIER_1 English(EN) · Yong Du, Tongbo Chen, Zhengxi Lu, Yizhou Liu, Bofan Chen, Tao Jiang, Wenhao Xu, Yongliang Shen ·

    ComputerSD:来自实时反馈的在线自蒸馏用于计算机使用代理

    arXiv:2609.40253v1 Announce Type: cross Abstract: Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate acti…

  171. arXiv cs.AI TIER_1 English(EN) · Yeonsung Jung, Trilok Padhi, Sina Shaham, Dipika Khullar, Joonhyun Jeong, Ninareh Mehrabi, Eunho Yang ·

    协同进化代理:从失败中学习作为硬负例

    arXiv:2511.22254v5 Announce Type: replace Abstract: Self-evolving agents improve their performance on long-horizon tasks by learning from their own interactions with an environment. A common approach uses the resulting failed trajectories as negatives for preference training. How…

  172. arXiv cs.AI TIER_1 English(EN) · Victor May, Van Khue Nguyen, Aaditya Salgarkar, Yishan Wang, Diganta Misra, Huu Nguyen ·

    评估运行时合约中的智能体:当不匹配会牺牲效率或质量时

    arXiv:2603.01209v3 Announce Type: replace Abstract: In CodeAct, language-model agents write Python that calls tools and use execution feedback to choose actions. Persistent runtimes preserve Python variables between actions; stateless runtimes clear them without resetting task pr…

  173. arXiv cs.AI TIER_1 English(EN) · Daniel Mitropolsky, Riccardo Neumarker, Emanuele Rimoldi, Susan S. Hong, Tomaso Poggio ·

    将图灵测试推广至交互式代理

    arXiv:2605.10851v2 Announce Type: replace Abstract: We initiate the study of the Generalized Turing Test (GTT), a formal generalization of Turing's imitation game from humans to arbitrary interactive agents. For agents $A$ and $B$, $A$ passes the GTT against $B$ if an instance of…

  174. arXiv cs.AI TIER_1 English(EN) · Youngmok Jung, Sirajul Salekin, Henry Tran, Javier Movellan, Zhao Huang, Manjot Bilkhu ·

    SCLATE:持续学习代理训练与评估的基石

    arXiv:2609.32391v2 Announce Type: replace Abstract: Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, c…

  175. arXiv cs.AI TIER_1 English(EN) · Xiao-Wen Yang, Weiyi Xu, Wen Da, Hang Xu, Canwei Li, Hong-Jie You, Pusen Dong, Yucheng Zeng, Zhaokai Luo, Yu-Feng Li, Yao Hu, Mu Chuan ·

    CompoWorld: 通用智能体的组合式环境扩展

    arXiv:2609.33665v2 Announce Type: replace Abstract: Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require a…

  176. arXiv cs.AI TIER_1 English(EN) · Yifan Liu, Praveen Venkateswaran, Abdulhamid Adebayo, Dong Wang ·

    通过查询条件归因审计代理行为

    arXiv:2609.33676v2 Announce Type: replace Abstract: LLM agents increasingly take consequential actions through interactions with users, policies, and external tools. Auditing these agents requires automated attribution of realized actions to their historical basis. However, exist…

  177. arXiv cs.AI TIER_1 Deutsch(DE) · Saswat Das, Parvati Viswanathan, Daniel Donnelly, Chang Huang, Sahar Abdelnabi, Ferdinando Fioretto ·

    SEABench:对自进化智能体中内源性错位进行基准测试

    arXiv:2609.35596v2 Announce Type: replace-cross Abstract: Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills,…

  178. arXiv cs.CL TIER_1 English(EN) · Yuhan Guo, Jinming Liu, Liang Xu, Ziqiang Li, Jianguo Huang, Zhicheng Wang, Hu Zhu, Qiuyu Chen, Yuntao Wei, Xin Jin, Wenjun Zeng ·

    EVOKE:在代理中引发世界知识以实现可转移决策

    arXiv:2609.38334v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the …

  179. arXiv cs.CL TIER_1 English(EN) · Qisheng Su, Hanchen Wang, Guanru Zhu, Huicheng Jiang, Qiuyinzhe Zhang, Kou Shi, Zhen Fang, Ziao Zhang, Qingnan Ren, Zehui Chen, Tao Gui, Feng Zhao ·

    GraphForge:通过图锚工作区合成训练可工作的代理

    arXiv:2609.38923v1 Announce Type: new Abstract: Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Ex…

  180. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Chen-Yu Lee ·

    VeriHarness:为长时任务扩展代理验证

    As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repea…

  181. Hugging Face Daily Papers TIER_1 English(EN) ·

    VeriHarness:为长时任务扩展代理验证

    As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repea…

  182. Hugging Face Daily Papers TIER_1 English(EN) ·

    Argo-Bench:在企业级工作流上评估数据代理

    Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently…

  183. Hugging Face Daily Papers TIER_1 Dansk(DA) ·

    DeskForge:来自桌面环境的密集监督用于计算机使用代理

    Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a con…

  184. Hugging Face Daily Papers TIER_1 English(EN) ·

    ComputerSD:来自实时反馈的在线自蒸馏,用于计算机使用代理

    Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers tok…

  185. Hugging Face Daily Papers TIER_1 English(EN) ·

    谁来验证图?语言代理因果行动验证中的错误指定攻击

    Causal action verifiers gate an agent's state-changing tool calls by checking whether each proposed intervention is identifiable against a committed action-state graph, and they issue a certificate that carries the identification argument and a one-sided lower confidence bound. O…

  186. Hugging Face Daily Papers TIER_1 English(EN) ·

    批准洗白:AI编码代理工具中批准-执行绑定失败的系统化

    Modern AI coding-agent harnesses (Claude Code, Codex CLI, Cursor) rest their security boundary on a largely unexamined assumption: that the action A a human approves is the same action A' the harness executes, where A is fixed by a stated policy for what a scope grant or session-…

  187. arXiv cs.AI TIER_1 English(EN) · Yanfei Zhang, Xu Lin ·

    修复之后:已更正代理体验的转移

    arXiv:2609.34603v2 Announce Type: replace Abstract: Does repairing an episode make its experience a better memory for the next task? We transfer the same failed source before and after accepted repair to a fixed target, alongside independent execution. Our 3,300 runs cover 100 Th…

  188. arXiv cs.AI TIER_1 English(EN) · Renxi Wang, Mingshan Hee, Fajri Koto, Timothy Baldwin, Haonan Li ·

    SkillGym:利用自动可验证环境生成训练技能使用代理

    arXiv:2609.37539v1 Announce Type: new Abstract: Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data …

  189. arXiv cs.AI TIER_1 English(EN) · Jio Oh, Seunghyun Do, Young-Jun Lee, Steven Euijong Whang, Dongyeop Kang ·

    主动代理的基础:原则、技术层与 Proactivity-Gym

    arXiv:2609.37267v1 Announce Type: new Abstract: Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and e…

  190. arXiv cs.AI TIER_1 English(EN) · Ido Levy, Asaf Yehudai, Segev Shlomov, Asaf Adi, Leshem Choshen ·

    提出从未被要求的要求:智能体中的水平和垂直主动性

    arXiv:2609.37236v1 Announce Type: new Abstract: An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on …

  191. arXiv cs.AI TIER_1 English(EN) · Savini Kashmira, Jayanaka L. Dantanarayana, Lingjia Tang, Jason Mars ·

    ContextRender:从执行依赖到代理上下文

    arXiv:2609.37743v1 Announce Type: new Abstract: LLM agents performing long-horizon tasks accumulate tool results that later steps may need. Passing the full history to every invocation is costly even when it fits within the context window, while reducing it risks omitting needed …

  192. arXiv cs.AI TIER_1 English(EN) · Wesley Shu ·

    Boundary-State Control for Tool-Using Language-Model Agents: Commit-Time Consistency under State Drift

    arXiv:2609.37475v1 Announce Type: new Abstract: Tool-using language-model agents can decide that an action is permissible and execute it only after security-relevant state has changed. We study this proposal-to-commit gap and introduce BSC-R, a deterministic effect-boundary mecha…

  193. arXiv cs.AI TIER_1 English(EN) · Jungwoo Yang, In Jin Kong, Yohan Jo ·

    SelfSearch:无奖励的自我改进代理搜索

    arXiv:2609.37968v1 Announce Type: new Abstract: Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstre…

  194. arXiv cs.AI TIER_1 English(EN) · Cheng Qian, Kunlun Zhu, Beibin Li, Zhenhailong Wang, Heng Ji ·

    为测试时AI4AI的Agent Harness设计学习元技能

    arXiv:2609.38143v1 Announce Type: new Abstract: Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weight…

  195. arXiv cs.AI TIER_1 English(EN) · Paras Dahal, Anton Bakhtin, Taco Cohen, Zhengxing Chen, Carole-Jean Wu, Rob Fergus, Scott Yih, Gabriel Synnaeve, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston, Anirudh Goyal ·

    思考前的思考:通过元推理扩展代理推理

    arXiv:2609.38147v1 Announce Type: new Abstract: As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to …

  196. arXiv cs.AI TIER_1 English(EN) · Ranuga Disansa, U. S. Samarasinghe, Lasith Gunawardena ·

    从词汇基线到代理检索增强生成:使用 SFIA 框架进行结构化技能和责任级别提取

    arXiv:2609.35806v1 Announce Type: cross Abstract: Automated skill extraction underpins workforce planning, yet most systems represent skills as flat labels with no notion of the responsibility level at which a skill is practiced. The Skills Framework for the Information Age (SFIA…

  197. arXiv cs.AI TIER_1 English(EN) · Linzhi Peng, Hanting Chen, Heng Chang, Ke Cheng, Bowen Du, Weifeng Lv ·

    PrimeSeeker:面向能力的深度搜索代理监督

    arXiv:2609.35816v1 Announce Type: cross Abstract: Large language model search agents are often trained with synthetic questions whose difficulty is increased through larger evidence graphs, additional hops, and longer trajectories. These global properties, however, are only indir…

  198. arXiv cs.AI TIER_1 English(EN) · Lingqi Jiang, Jialuo Chen, Jianan Ma, Xinhao Deng, Xiaohu Du, Sibo Yi, Yuqi Qing, Zhenguang Liu, Qinming He, Shiwen Cui, Changhua Men ·

    MMSkillRisk:当多模态技能变成陷阱时,代理能否保持安全?

    arXiv:2609.35912v1 Announce Type: cross Abstract: Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, a…

  199. arXiv cs.AI TIER_1 English(EN) · Xiaoyu Xiong, Tsun-Hsuan Wang, Yi-Ling Qiao, Tao Du, Minchen Li ·

    Text2Sim:基于蒸馏专业知识的代理物理模拟生成

    arXiv:2609.36593v1 Announce Type: cross Abstract: Creating diverse physical simulations remains labor-intensive because assets, layout, physical parameters, motion, control, and rendering must be designed and debugged jointly. We present Text2Sim, a simulation-specialized agentic…

  200. arXiv cs.AI TIER_1 English(EN) · Haomin Qi, Xiangzhe Xu, Yiming Huang, Jingbo Shang, Chengpeng Wang ·

    WitnessGym:对构建 Bug Witness 的编码代理进行基准测试

    arXiv:2609.36635v1 Announce Type: cross Abstract: Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings …

  201. arXiv cs.AI TIER_1 English(EN) · Jiexing Qi, Yu He, Jun Liu, Qichen Huang, Shaohua Hu, Zhan Dang, Guohua Chen, Rui Yang, Wen Jiang, Yang Liu, Tao Lyu, Fangming Li ·

    VACE:代理模型和工具的验证门控交替协同进化

    arXiv:2609.37105v1 Announce Type: cross Abstract: Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates ch…

  202. arXiv cs.AI TIER_1 English(EN) · Soyeong Jeong, Sujay Kumar Jauhar, Sung Ju Hwang, Andrew Joohun Nam ·

    跟随实体:为代理搜索构建语料库地图

    arXiv:2609.37226v1 Announce Type: cross Abstract: Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its lates…

  203. arXiv cs.AI TIER_1 English(EN) · Rohith Reddy Bellibatlu, Zichong Wang, Wenbin Zhang ·

    Agent基准测试是否名副其实?对使用工具的Agent环境的可执行合约审计

    arXiv:2609.37315v1 Announce Type: cross Abstract: Tool-using agents are entering settings where a wrong action carries real cost, and the benchmarks certifying them grade what each simulated tool call reports having done, assuming the tool did what its interface advertises. The a…

  204. arXiv cs.AI TIER_1 English(EN) · Yifan Kang, Zihan Wang, Zhiwen Fan, Bangya Liu ·

    Encore:操纵策略的少样本代理发现

    arXiv:2609.37359v1 Announce Type: cross Abstract: Coding agents can now write, run, and debug programs with little human help. Robot tasks, however, are usually specified by a sentence that leaves out how to grasp, in what order to make contact, and what the result should look li…

  205. arXiv cs.AI TIER_1 English(EN) · Beining Xu, Peichun Hua, Yunming Xiao ·

    循环中的后门:通过恶意检索器破坏代理搜索

    arXiv:2609.37468v1 Announce Type: cross Abstract: Agentic retrieval-augmented generation (RAG) interleaves reasoning with repeated retrieval, giving the retriever influence over both the evidence an agent observes and its subsequent search decisions. We study retriever backdoors …

  206. arXiv cs.AI TIER_1 English(EN) · Sicheng Xie, Yitong Chen, Haidong Cao, Shunlin Lu, Zuxuan Wu, Yu-Gang Jiang ·

    探索、执行、进化:具身智能体的技能获取与复用循环

    arXiv:2609.37810v1 Announce Type: cross Abstract: Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potent…

  207. arXiv cs.AI TIER_1 English(EN) · Yiming Cheng (The University of Chicago), Alfin Wijaya Rahardja (Fudan University), Mengshi Zhang (TensorBlock, Inc), Zihao Chen (TensorBlock, Inc), Zhenpeng Chen (Tsinghua University), Yiling Lou (University of Illinois Urbana-Champaign) ·

    AgentBug-Smith:在 Agentic 系统中自动复现真实世界的 Harness Bug

    arXiv:2609.37864v1 Announce Type: cross Abstract: Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is further hindered by existing benchmarks, which contain only a small and fixed number…

  208. arXiv cs.AI TIER_1 English(EN) · Julien Knafou, Luc Mottin, Alexandre Flament, Paul van Rijen, Esteban Gaillac, Patrick Ruch ·

    BITEM 在 NTCIR-19 R2C2 任务中:从 Agentic RAG Pipeline 信号预测置信度

    arXiv:2609.37993v1 Announce Type: cross Abstract: The BITEM team entered both subtasks of the NTCIR-19 R2C2 task with a single agentic pipeline, in which a model searches, reads and records evidence over a movie corpus while an orchestrator holds the record and rules on what may …

  209. arXiv cs.AI TIER_1 English(EN) · Yuqiao Meng, Luoxi Tang, Sakshi Sunil Narvekar, Rupali Rajendra Vaje, Yingxue Zhang, Muchao Ye, Zhaohan Xi ·

    EquiMem:通过博弈论均衡校准多智能体辩论中的共享内存

    arXiv:2605.09278v2 Announce Type: replace Abstract: Multi-agent debate (MAD) systems increasingly rely on shared memory to support long-horizon reasoning, but this convenience opens a critical vulnerability: a single corrupted entry can contaminate the downstream memory-augmented…

  210. arXiv cs.AI TIER_1 English(EN) · Sen Zhao, Ruiqi Kong, Zuyu Zhang, Lifeng Shen, Xinyu He, Ding Zou, Xu Zhang, Qinghua Zhang ·

    CoMemBench:跨多智能体工作流拓扑的协作记忆边界基准测试

    arXiv:2609.32192v2 Announce Type: replace Abstract: Multi-agent workflows require task-relevant information to be shared across agents, while irrelevant, stale, unverified, or incompatible information must remain isolated. We call this task-conditioned scope of information a coll…

  211. arXiv cs.AI TIER_1 English(EN) · Yuchen Song, Andong Chen, Wenxin Zhu, Muyun Yang, Tiejun Zhao ·

    RepoMAS:使用问题驱动的多智能体系统解决渐进式指定任务

    arXiv:2609.32490v2 Announce Type: replace Abstract: LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task requirements are sufficiently specified before execution. In practice, user requests are often incomplete, and…

  212. arXiv cs.AI TIER_1 English(EN) · Jiecong Wang, Hao Peng, Zhanyi Wang ·

    用于长时域智能体的自适应一致性图

    arXiv:2609.32754v2 Announce Type: replace Abstract: Large language model agents can often make reasonable local decisions on short tasks, yet their performance degrades when success requires long sequences of dependent actions and tool calls. During execution, task requirements, …

  213. arXiv cs.AI TIER_1 English(EN) · Zhengkun Di, Bin Shi, Kai Sun, Yiming Xu, Bo Dong ·

    代理何时应检查外部状态?为存储的意图进行预算观察

    arXiv:2609.37125v1 Announce Type: new Abstract: Prospective memory allows an agent to retain an intention tied to a future condition, but the stored intention does not reveal whether that condition currently holds. Checking it may require web access, multi-step tool use, and paid…

  214. arXiv cs.AI TIER_1 English(EN) · Lingrui Xu, Yangqin Jiang, Jiachang Zhang, Xubin Ren, Chao Huang ·

    AnyAct:通用动作,用于自进化智能体

    arXiv:2609.37025v1 Announce Type: new Abstract: As large language models (LLMs) advance, AI agents are increasingly deployed in open-world environments to tackle complex sequential tasks (e.g., document processing, cross-application collaboration), relying heavily on actions rang…

  215. arXiv cs.AI TIER_1 English(EN) · Junjie Yao, Zhangchen Zhou, Zhi-Qin John Xu ·

    CADOC:面向长时域智能体的缓存感知动态对象上下文

    arXiv:2609.37012v1 Announce Type: new Abstract: For a long-horizon agent, context is the bottleneck: the history is resent with every request, the window caps task length, and reasoning degrades as the history grows. Replacing structured objects with compact retrieval Cards short…

  216. arXiv cs.AI TIER_1 English(EN) · Zeyu Gan, Zixuan Gong, Yong Liu ·

    将进化作为学习:自改进个人代理的近似、泛化和优化极限

    arXiv:2609.36892v1 Announce Type: new Abstract: As the capabilities of large language models (LLMs) continue to advance, increasing attention is turning to how to translate their abilities into useful behavior. Personal agents bring this question into everyday settings, where mod…

  217. arXiv cs.AI TIER_1 English(EN) · Bo Mao, Hang He, Linting Wang, Lizhi Lin, Maosen Zhou, Guanming Liu, Jinxiu Liu, Tianyu Huai, Chaoyun Zhang, Bingxuan Li, Kepeng Lei, Guanting Dong, Zhou Shao, Rui Zheng, Hang Yan, Jie Zhou, Chengcheng Wan, Tao Gui, Liang He, Xipeng Qiu ·

    WEFT:通用智能体训练后工具使用扩展

    arXiv:2609.36887v1 Announce Type: new Abstract: Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent ha…

  218. arXiv cs.AI TIER_1 English(EN) · Zhen Xiong, Qiaoyu Tan ·

    EASE:面向自进化智能体的行为自适应技能策展

    arXiv:2609.36746v1 Announce Type: new Abstract: Agent skills provide a lightweight mechanism for self-evolving agents to accumulate reusable procedural knowledge without updating model parameters. However, existing learned skill curators typically optimize curation without explic…

  219. arXiv cs.AI TIER_1 (CA) · Gabriel Orlanski, Alex L. Zhang, Avi Trost, Vincent Sunn Chen, Frederic Sala, Aws Albarghouthi, Ludwig Schmidt ·

    代理可以为代理设计库吗?

    arXiv:2609.36730v1 Announce Type: new Abstract: Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDe…

  220. arXiv cs.AI TIER_1 English(EN) · Xin Yu, Lizhu Zhang, Jiamu Bai, Yanhong Wu, Zellux Wang, Serena Li, Weiwei Li, Lingzhou Xue, Xiangjun Fan, Bo Peng ·

    MLToolBench:为机器学习开发学习增强工具的智能体

    arXiv:2609.36679v1 Announce Type: new Abstract: Machine learning engineering (MLE) agents have made substantial progress, but learning through ML experimentation remains costly in time and computation. Synthetic environments reduce these costs while introducing variations in data…

  221. arXiv cs.LG TIER_1 English(EN) · Mihir Chauhan, Aniket Bera ·

    涌现而非带宽:物理耦合与学习型多智能体通信的局限性

    arXiv:2609.34373v2 Announce Type: replace Abstract: Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol …

  222. arXiv cs.LG TIER_1 English(EN) · Dongchan Shin, Xing Han L\`u, Jiaqi Deng, Jay Gala, Tom\'as Vergara Browne, Jaewon Moon, Fengyuan Liu, Alexandre Drouin, Siva Reddy, Alexandre Lacoste ·

    AdaptArena:评估网络代理的测试时个性化

    arXiv:2609.36488v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated strong performance on complex web navigation tasks, yet they remain brittle in real-world settings where user intentions are underspecified and preferences are heterogeneous. In pr…

  223. arXiv cs.CL TIER_1 English(EN) · Yun Peng, Zihan Wu, Zeyang Zhuang, Xin Zhou, Rui Shu, Xu Han, Chun Yong Chong, Yuan Wang, Jiakun Liu ·

    LoLBench:使用长时序提案在大型软件系统上评估编码代理

    arXiv:2609.37143v1 Announce Type: cross Abstract: Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' imple…

  224. arXiv cs.CL TIER_1 English(EN) · Shinan Zhang, Tao Zhang, Qihui Zhu, Mengjie Zhang, Dong Jin, Yunpeng Hou, Shuangwu Chen, Xiaobin Tan, Quan Zheng, Jian Yang ·

    LatCom:跨代理潜在压缩以实现高效多代理协作

    arXiv:2609.37017v1 Announce Type: new Abstract: LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. However, directly forwarding all sender latents makes…

  225. arXiv cs.CL TIER_1 English(EN) · Amirhossein Abaskohi, Amirhossein Dabiriaghdam, Lele Wang, Peter West, Giuseppe Carenini ·

    DeepRewind:预测和修复深度研究代理中的过早承诺

    arXiv:2609.36344v1 Announce Type: new Abstract: Deep-research agents conduct long-horizon investigations through iterative search, evidence evaluation, belief revision, and synthesis. However, they may commit to claims before sufficient evidence is available, causing later reason…

  226. arXiv cs.AI TIER_1 English(EN) · Yuanhao Li, Hongbo Wang, Xuhong Chen, Yiming Cao, Xunzhu Tang ·

    Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents

    arXiv:2609.33875v2 Announce Type: replace-cross Abstract: Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterfactual Rollout Replay (CRR), a training-ti…

  227. arXiv cs.AI TIER_1 English(EN) · Weiyi Xu, Xiaowen Yang, Wen Da, Hang Xu, Canwei Li, Hongjie You, Pusen Dong, Yucheng Zeng, Zhaokai Luo, Mu Chuan ·

    Skill2Env: 面向通用智能体的面向能力的技能环境合成

    arXiv:2609.33772v2 Announce Type: replace Abstract: Executable environments are critical for post-training agents on tasks that require tool use and multi-step interaction, but constructing executable tasks together with their environments remains difficult to scale. Skills provi…

  228. arXiv cs.AI TIER_1 English(EN) · Michael Lee, Zhipeng Wei, Yue Dong, N. Benjamin Erichson ·

    分割与注入:代理能否从碎片中重建间接提示注入?

    arXiv:2609.36576v1 Announce Type: new Abstract: Agentic systems are now being widely used to orchestrate tools and reason over long contexts. However, the improving capabilities of the large language models powering these agents also create new attack surfaces for indirect prompt…

  229. arXiv cs.AI TIER_1 English(EN) · Leonardo Ferreira, Gardenia Liu, Kaden Zheng ·

    超越对称代理:小型语言模型中的认知多样性与多代理辩论

    arXiv:2609.35875v1 Announce Type: new Abstract: Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity …

  230. arXiv cs.AI TIER_1 English(EN) · Yihao Wang, Linhan Xia, Rui Liu, Zhaofeng Zhang, Hongyu Wu, Yang Yang, Jinglu He, Yu Guo, Kai Lei ·

    SAGE:一种用于自演化代理的统计接受门

    arXiv:2609.36043v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate…

  231. arXiv cs.AI TIER_1 English(EN) · Ziyang Yu, Liang Zhao, Bowen Zhu, Hasibul Haque ·

    StateTape:面向长时域编码代理的动作条件证据生命周期建模

    arXiv:2609.36319v1 Announce Type: new Abstract: Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based…

  232. arXiv cs.AI TIER_1 English(EN) · Long Phan, Stephen K. Yang, Jason J. Lim, Mantas Mazeika, Wenyu Zhang, Zheyuan Liu, Richard Ren, Jingxiang Meng, Yaoteng Tan, Weiliang Zhao, Addison Wu, Matei Anghel, Dan Hendrycks ·

    CheatBench:衡量AI代理中的奖励博弈

    arXiv:2609.36308v1 Announce Type: new Abstract: Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to…

  233. arXiv cs.AI TIER_1 English(EN) · Thibaud Gloaguen, Niels M\"undler-Sasahara, Mark Niklas M\"uller, Veselin Raychev, Martin Vechev ·

    评估 AGENTS.md:仓库级上下文文件对编码代理有帮助吗?

    arXiv:2602.11988v3 Announce Type: replace-cross Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md. Although this practice is strongly encouraged by agent developers, there is currently no rigo…

  234. arXiv cs.AI TIER_1 English(EN) · Ziluowen Luo, Senzhang Wang, Chaozhuo Li, Jun Yin, Hao Yan, Ming Cheng, Chenxu Wang, Songyang Liu, Litian Zhang, Qiwei Ye, Zheng Liu, Philip S. Yu ·

    提炼智能体系统:跨越模型、构件和约束的路线图

    arXiv:2609.36630v1 Announce Type: new Abstract: Modern agents increasingly rely on memories, tools, and execution logic, so their competence extends beyond model parameters. This shift exposes a limitation of conventional knowledge distillation, which asks how a student model imi…

  235. arXiv cs.AI TIER_1 English(EN) · Zhong Guan, Yongjian Guo, Haoran Sun, Wen Huang, Shuai Di, Likang Wu, Xiong Jun Wu, Hongke Zhao ·

    异步智能体强化学习中缺失的旧logit:语义不匹配与离策略校正的修复方法

    arXiv:2605.12070v3 Announce Type: replace-cross Abstract: Asynchronous reinforcement learning improves rollout throughput for large language model agents by decoupling sample generation from policy optimization, but it also introduces a critical failure mode for PPO-style off-pol…

  236. arXiv cs.AI TIER_1 English(EN) · Ziyu Liu, Jun Chen, Lixu Wang ·

    语言智能体持续自进化的语义投影

    arXiv:2609.36626v1 Announce Type: new Abstract: Language-model agents increasingly rely on persistent natural-language skills to adapt beyond their frozen model parameters. When a shared skill is repeatedly revised from a non-stationary, heterogeneous task stream, however, improv…

  237. arXiv cs.AI TIER_1 English(EN) · Xiaojing Sun, Yuhan Zeng, Zihua She, Xiao Wang ·

    哪些自我改进值得信赖?当智能体重复使用其基准测试时的可靠自我改进

    arXiv:2609.33180v2 Announce Type: replace-cross Abstract: As recursive self-improvement (RSI) rapidly advances, reliable evaluation becomes critical for guiding adaptive search. RSI typically relies on finite evaluation resources, such as fixed benchmarks, to determine which modi…

  238. arXiv cs.MA (Multiagent) TIER_1 Dansk(DA) · Ping Wang ·

    SkillSeek:在市场规模上重新审视代理技能检索

    Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection…

  239. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Xueqing Liu ·

    PatchHolmes:通过列表式选择进行智能补丁检索

    Patch retrieval, the task of finding the commit that fixes a known vulnerability, is the foundation of vulnerability management workflows, yet 60% to 63% of CVEs in the major advisory databases lack a patch link. We present PatchHolmes, a two-phase patch retrieval system that pai…

  240. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Zhengye Han ·

    多智能体系统在哪里失效?基于证据的集体机制诊断

    When a multi-agent system answers correctly, it is tempting to conclude that its agents shared, checked, and used information as intended. Yet a system can break one of its collective mechanisms, the rules that govern how agents route, admit, store, and act on shared information,…

  241. Hugging Face Daily Papers TIER_1 English(EN) ·

    GraphForge:通过图锚定工作空间合成训练可工作的代理

    Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Existing pipelines either generate files with mode…

  242. Hugging Face Daily Papers TIER_1 English(EN) ·

    PivotOPD:学习从多轮智能体中的关键性错误中恢复

    On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound ac…

  243. Hugging Face Daily Papers TIER_1 Dansk(DA) ·

    SkillSeek:在市场规模上重新审视代理技能检索

    Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection…

  244. Hugging Face Daily Papers TIER_1 English(EN) ·

    PatchHolmes:通过列表式选择进行智能补丁检索

    Patch retrieval, the task of finding the commit that fixes a known vulnerability, is the foundation of vulnerability management workflows, yet 60% to 63% of CVEs in the major advisory databases lack a patch link. We present PatchHolmes, a two-phase patch retrieval system that pai…

  245. Hugging Face Daily Papers TIER_1 English(EN) ·

    EviRover: 强化代理感知,超越一瞥

    Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge…

  246. Hugging Face Daily Papers TIER_1 English(EN) ·

    A2Z GameSpec-Bench:编码代理能多忠实地根据游戏设计规范生成游戏?

    Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form Game Design Documents (GDDs) describe requireme…

  247. Hugging Face Daily Papers TIER_1 English(EN) ·

    虚假边界:诊断与缓解自进化搜索代理中的协同作弊

    Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so…

  248. Hugging Face Daily Papers TIER_1 English(EN) ·

    DAGent:深度研究代理的评估后增长规划

    Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting because they support parallel execution and isolate e…

  249. Hugging Face Daily Papers TIER_1 English(EN) ·

    JevSpawn:通过组合式动作空间实现自适应代理推理

    LLM agents generate intermediate reasoning and actions token by token, making extended interactions slow and computationally expensive. Jev-style models offer fast probabilistic predictions over finite fields, but require those fields to be specified in advance. This requirement …

  250. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Xiaoying Tang ·

    CollabFlow:Agent协作的递归自我改进

    Recursive self-improvement (RSI) lets a system improve from its own outcomes; in LLM-based multi-agent systems, Agents refine one another within a task, and outcomes improve how they collaborate across tasks. However, existing multi-agent collaboration leaves this loop open: coll…

  251. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Zilong Wang ·

    递归组织改进:人类-代理组织建模规范

    Stronger AI agents do not automatically produce better organizations: teams must also learn which work arrangements to retain and when to reconsider them. We propose a modeling specification for recursive organization improvement and evaluate it through an executable checker, a p…

  252. Hugging Face Daily Papers TIER_1 English(EN) ·

    HybridCUA:学习编排 GUI 和 CLI 以用于计算机使用代理

    Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or augment GUI interactions with application specif…

  253. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Patrick Ruch ·

    BITEM 在 NTCIR-19 R2C2 任务中:从 Agentic RAG Pipeline 信号预测置信度

    The BITEM team entered both subtasks of the NTCIR-19 R2C2 task with a single agentic pipeline, in which a model searches, reads and records evidence over a movie corpus while an orchestrator holds the record and rules on what may be submitted. A claim is admitted only once an ent…

  254. Hugging Face Daily Papers TIER_1 English(EN) ·

    探索、执行、进化:具身智能体的技能获取与复用循环

    Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, t…

  255. Hugging Face Daily Papers TIER_1 English(EN) ·

    主动代理的基础:原则、技术层和 Proactivity-Gym

    Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three join…

  256. Hugging Face Daily Papers TIER_1 English(EN) ·

    提出从未被要求的要求:智能体中的水平和垂直主动性

    An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. …

  257. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Andrew Joohun Nam ·

    跟随实体:为代理搜索构建语料库地图

    Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach th…

  258. Hugging Face Daily Papers TIER_1 English(EN) ·

    LoLBench:使用长时程提案评估大型软件系统上的编码代理

    Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' implementation capability to produce correct code edits…

  259. Hugging Face Daily Papers TIER_1 English(EN) ·

    LatCom:高效多智能体协作的跨智能体潜在压缩

    LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. However, directly forwarding all sender latents makes the receiver-side context scale with both the n…

  260. arXiv cs.AI TIER_1 English(EN) · Md Shohel Arman, Igor Molybog ·

    面向编码代理的紧凑型文档:一个基准测试、一个优化器以及它为何不具有迁移性

    arXiv:2609.31587v1 Announce Type: cross Abstract: We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether co…

  261. arXiv cs.AI TIER_1 English(EN) · Weida Liang, Shi Qiu, Zhun Wang, Simon Sure, Xiaoyuan Liu, Tianneng Shi, Zhaorun Chen, Wenbo Guo, Dawn Song ·

    AgentXploit:AI Agent 的自主式从仓库到运行时的红队测试

    arXiv:2609.31318v1 Announce Type: cross Abstract: AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software co…

  262. arXiv cs.AI TIER_1 English(EN) · Haoran Zhang, Hengtong Zhang, Zhiyu Liang, Yu Yan, Decheng Zuo, Hongzhi Wang ·

    超越已批准操作:Agent工作流中持久化结果的运行时验证

    arXiv:2609.31301v1 Announce Type: cross Abstract: Large language model agents increasingly act on software systems, no longer merely generating text but also changing databases and online services. However, an approved database update may succeed yet leave an unapproved notificat…

  263. arXiv cs.AI TIER_1 English(EN) · Alexander Gill, Md Farhan Ishmam, Xuyen Nguyen, Neha Bhat, Parker Henry DeYoung, Fateme Hashemi Chaleshtori, Nathan Stringham, Kenneth Marino, Ana Marasovi\'c ·

    搜索之后才是难点:在知识合成、组织和展示方面对网络代理进行基准测试

    arXiv:2609.30604v1 Announce Type: cross Abstract: Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spre…

  264. arXiv cs.AI TIER_1 English(EN) · Jiaqi Ding, Guorong Wu ·

    通过认知实验范式探究智能体记忆中的稳定性-可塑性权衡

    arXiv:2609.30558v1 Announce Type: cross Abstract: Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a cogniti…

  265. arXiv cs.AI TIER_1 English(EN) · Mukul Chhabra, Shail Patel, Luigi Medrano ·

    CARGO:生产环境中智能体AI的上下文感知检索门控评估

    arXiv:2609.30471v1 Announce Type: cross Abstract: Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically appl…

  266. arXiv cs.AI TIER_1 English(EN) · Leon Goldberg, Gal Engelberg, Eden Yavin, Elad Elouz, Ariel Zadok, Konstantin Koutsyi ·

    代码智能体不足以应对!评估企业安全大脑以实现智能云调查

    arXiv:2609.30345v2 Announce Type: cross Abstract: Cloud-security investigation is dominated by population tasks: which identities can read a data store, how many resources fail a control, what is reachable from another account. These resolve against a complete inventory, not a na…

  267. arXiv cs.AI TIER_1 English(EN) · Carolina Fortuna, Blaz Bertalanic ·

    多智能体跨越析取式和补偿式任务的扩展

    arXiv:2609.31563v1 Announce Type: new Abstract: Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analy…

  268. arXiv cs.AI TIER_1 English(EN) · Bart{\l}omiej Cupia{\l}, Jens Tuyls, Maciej Wo{\l}czyk, Davide Paglieri, Martin Klissarov, Benjamin Eysenbach, Piotr Mi{\l}o\'s, Karthik R. Narasimhan ·

    抽象阶梯的上下:语言代理的基于代码的技能

    arXiv:2609.31076v1 Announce Type: new Abstract: Language agents struggle to act and learn in environments that require long sequences of low-level actions. Code-based abstractions can make these agents more productive by letting them invoke reusable skills instead of repeatedly s…

  269. arXiv cs.AI TIER_1 English(EN) · Zhensheng Zou (Peking University), Guoqing Wang (Peking University), Dan Hao (Peking University) ·

    压缩你所见的,而非你所说的:基于锚定上下文蒸馏的潜在观察式软件工程代理

    arXiv:2609.31430v1 Announce Type: new Abstract: Tool observations dominate the context of software-engineering agents, making long interaction histories costly to maintain. Existing context compression methods can discard information needed by later actions, while adapting agents…

  270. arXiv cs.AI TIER_1 English(EN) · Maokai Qin, Chuan Qin, Qi Zhang, Dianyu Liu, Zirui Liu, Hongting Niu, Yuanchun Zhou, Hengshu Zhu ·

    SciHorizon-eLab:一个用于可扩展科学具身智能体基准测试的智能体协议到任务编译器

    arXiv:2609.30971v1 Announce Type: new Abstract: Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely…

  271. arXiv cs.AI TIER_1 Deutsch(DE) · Guanyu Nie, Fangzhou Zhu, Shixiong Kai, Xiongwei Han, Tao Zhong, Mingxuan Yuan ·

    SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting

    arXiv:2609.30861v1 Announce Type: new Abstract: Language-model agents increasingly improve by converting execution experience into reusable external skills. Yet repeated skill updates form a learning process of their own: locally useful edits can accumulate into redundant or task…

  272. arXiv cs.AI TIER_1 English(EN) · Xiaoyang Li, Yiqi Wang, Chencheng Zhu, KE XU, Wencheng Yang, Zequn Sun, Pingan Song, Yiqun Duan, Taotao Cai ·

    共享代理记忆中认识论承认的基准和诊断研究

    arXiv:2609.30813v1 Announce Type: new Abstract: Evaluating claim admission in shared agent memory is challenging because repeated claims may be mistaken for independent evidence. An agent may copy or paraphrase a retrieved belief, while admitting a false claim exposes subsequent …

  273. arXiv cs.AI TIER_1 English(EN) · Shane Caldwell, Max Harley, Ads Dawson, Michael Kouremetis, Vincent Abruzzo, Will Pearce ·

    ScopeBench:在目标压力下,智能体是否会保持参与边界?

    arXiv:2609.30325v1 Announce Type: new Abstract: Agents are increasingly deployed with real autonomy in web application and network penetration testing, where a single out-of-scope action can breach a client's engagement boundary. Existing offensive-security benchmarks measure raw…

  274. arXiv cs.AI TIER_1 English(EN) · Yiran Hu, Nan Jiang, Shanchao Liang, Anik Dey, Yi Wu, Lin Tan ·

    分析和缓解编码代理中的成本低效行为

    arXiv:2609.30725v2 Announce Type: new Abstract: Although effective, coding agents often incur substantial monetary costs. Their recurring cost-inefficient behaviors remain underexplored. We conduct the first study of behavioral cost inefficiencies in coding agents, analyzing 1,20…

  275. arXiv cs.AI TIER_1 English(EN) · Zihao Zhu, Siwei Lyu, Adel Bibi, Baoyuan Wu ·

    隐蔽性各异,危害共存:基于技能的智能体系统的技能级联攻击

    arXiv:2609.30383v1 Announce Type: new Abstract: A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable …

  276. arXiv cs.AI TIER_1 English(EN) · Salma Roshdy Aly, Hussein Assaf, Ziad Kobti ·

    多智能体代码评判何时真正落地?两种无标签测量方法,以及一个拒绝猜测的评判器

    arXiv:2609.30328v1 Announce Type: new Abstract: When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for. Multi-agent v…

  277. arXiv cs.AI TIER_1 English(EN) · Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth ·

    Agentick:通用序贯决策智能体的统一基准

    arXiv:2605.06869v3 Announce Type: replace Abstract: AI agent research spans a wide spectrum: from RL agents that learn from scratch to foundation model agents that leverage pre-trained knowledge, yet no unified benchmark enables fair comparison across these approaches. We present…

  278. arXiv cs.AI TIER_1 English(EN) · Fangzhou Li, Pagkratios Tagkopoulos, Ilias Tagkopoulos ·

    SkillFlow:可扩展且高效的代理技能检索系统

    arXiv:2504.06188v3 Announce Type: replace Abstract: AI agents can extend their capabilities at inference time by loading reusable skills into context, yet equipping an agent with too many skills, particularly irrelevant ones, degrades performance. As community-driven skill reposi…

  279. Hugging Face Daily Papers TIER_1 (CA) ·

    代理可以为代理设计库吗?

    Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an age…

  280. Hugging Face Daily Papers TIER_1 English(EN) ·

    WEFT:通用智能体训练后工具使用扩展

    Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in is…

  281. Hugging Face Daily Papers TIER_1 English(EN) ·

    HybridCUA:学习编排 GUI 和 CLI 以用于计算机使用代理

    Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or augment GUI interactions with application specif…

  282. Hugging Face Daily Papers TIER_1 English(EN) ·

    主动代理的基础:原则、技术层和 Proactivity-Gym

    Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three join…

  283. Hugging Face Daily Papers TIER_1 English(EN) ·

    提出从未被要求的要求:智能体中的水平和垂直主动性

    An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. …

  284. Hugging Face Daily Papers TIER_1 English(EN) ·

    追踪实体:为代理搜索构建语料库地图

    Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach th…

  285. Hugging Face Daily Papers TIER_1 English(EN) ·

    RLE-Bench:为机器人学习工程师设计的编码代理资格考试

    Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as policies or controllers, offering limited cove…

  286. Hugging Face Daily Papers TIER_1 English(EN) ·

    SkillGym:利用自动可验证环境生成训练技能使用代理

    Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain und…

  287. Hugging Face Daily Papers TIER_1 English(EN) ·

    EVOKE:在代理中引发世界知识以实现可转移决策

    Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the cost of additional training and errors that comp…

  288. Hugging Face Daily Papers TIER_1 English(EN) ·

    AIM:用于自动化研究的代理式思想管理

    Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven search and identify three core challenges: organizing evolving research ideas, selecting promising directions, and maintaining alig…

  289. Hugging Face Daily Papers TIER_1 English(EN) ·

    为测试时AI4AI的Agent Harness设计学习元技能

    Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed. To make the Builder's experience…

  290. Hugging Face Daily Papers TIER_1 English(EN) ·

    AREX-2:通过长时程反思性任务推进自改进代理

    We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, which produces a solution better than the current o…

  291. Hugging Face Daily Papers TIER_1 English(EN) ·

    令人震惊的简单自我反思无需RL即可改进Agentic模型

    People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We …

  292. Hugging Face Daily Papers TIER_1 Deutsch(DE) ·

    SEABench:对自进化智能体中内源性错位进行基准测试

    Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills, in response to user and environment feedback. However, lo…

  293. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Junpei Komiyama ·

    用于多智能体推理的自适应专家组

    Multi-agent systems bring together language model agents with different roles to propose, review, and refine solutions. Each agent's response depends on its model's capabilities, the reasoning strategy defined by its system prompt, and the information in its input context. Existi…

  294. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Jiaxing Song ·

    从迁移到校准:跨模型、司法管辖区和规模保留代理能力

    Agents need calibration when deployment conditions change: replacing a driving model, including a foundation-to-post-trained transition; crossing jurisdictions; or scaling across heterogeneous markets and sources. Interface compatibility alone does not establish capability retent…

  295. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Zheng Liu ·

    运行时主体研究的即时代理记忆

    Memory is critical for AI agents. Many existing agent-memory systems follow an Ahead-of-Time (AOT) design, constructing memory before a specific request arrives. While this reduces online serving cost, such request-agnostic memory construction can discard fine-grained information…

  296. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Aniket Bera ·

    涌现而非带宽:物理耦合与学习型多智能体通信的局限性

    Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer t…

  297. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Aniket Bera ·

    涌现而非带宽:物理耦合与学习型多智能体通信的局限性

    Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer t…

  298. Hugging Face Daily Papers TIER_1 English(EN) ·

    Dr.Credit:基于规则的深度研究代理过程信用分配

    Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many existing credit assign…

  299. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Rishika Lall ·

    何时选择取代提取?基于类型化决策模型的代理记忆预注册测试

    Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether rank…

  300. Hugging Face Daily Papers TIER_1 English(EN) ·

    通过似然引导的工具空间优化实现自演化智能体

    Self-evolving agents can continually improve their behavior, while tools define the executable action space through which they interact with the environment. However, exposing the full tool library to model introduces substantial irrelevant context and can impair tool-use decisio…

  301. Hugging Face Daily Papers TIER_1 English(EN) ·

    StateGuard:具有感知有效性的干预措施,用于长时序数据代理的分析状态管理

    LLM-based agents have shown strong capabilities in automated data analysis and are increasingly moving toward long-horizon, multi-stage analytical workflows. However, as the analytical process evolves, constraints, variables, and conclusions remain implicitly embedded in interact…

  302. Hugging Face Daily Papers TIER_1 English(EN) ·

    AutoDataBench:代理能否编写用于自我改进循环的数据?

    Recent gains in language model capability have come more from data than from architecture. Frontier labs and data companies produce verifiable agentic tasks, which supervised finetuning and reinforcement learning then turn into capability.This production line still rests on human…

  303. Hugging Face Daily Papers TIER_1 English(EN) ·

    CheatBench:衡量AI代理中的奖励博弈

    Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized info…

  304. Hugging Face Daily Papers TIER_1 English(EN) ·

    Org-Agent:超越个人助理,迈向组织化智能体

    Language model agents serving organizations must coordinate requests from multiple users while using knowledge distributed across their interactions. We identify two complementary capabilities for this setting, namely cross-user interaction and decision-making, as well as cross-u…

  305. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Muning Wen ·

    TRACE:在演进式多智能体系统中管理记忆有效性

    Persistent memory lets language-model agents carry information across long-running collaborations, but leaves a lifecycle question open: what may a returning agent still act on once the shared state has changed? A memory can be correctly retrieved, relevant to the current task, a…

  306. arXiv cs.MA (Multiagent) TIER_1 English(EN) · EverMind AI ·

    Raven:可组合智能体智能的“万能马具”

    As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to speci…

  307. arXiv cs.MA (Multiagent) TIER_1 Deutsch(DE) · Manik Gupta ·

    CORTEX: 通用智能体的已验证体验层

    An agent can solve a task today and face the same task under new facts, tools, or governing knowledge tomorrow. Most agent systems can retrieve relevant text or recall prior conversations, but they lack a principled way to decide when a previous solution is still valid, when it m…

  308. Hugging Face Daily Papers TIER_1 English(EN) ·

    WideSWE:编码代理能否协调跨存储库的更改?

    Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multipl…

  309. Hugging Face Daily Papers TIER_1 English(EN) ·

    Raven:可组合智能体智能的“装备之装备”

    As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to speci…

  310. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Kaden Zheng ·

    超越对称智能体:小型语言模型中的认知多样性与多智能体辩论

    Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity among agents is the driver, in the setting where…

  311. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Yunming Xiao ·

    后门即搜索:通过恶意检索器破坏智能体搜索

    Agentic retrieval-augmented generation (RAG) interleaves reasoning with repeated retrieval, giving the retriever influence over both the evidence an agent observes and its subsequent search decisions. We study retriever backdoors that exploit this feedback loop and repurpose weak…

  312. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Changlun Li ·

    当更好变成更糟:自适应世界中自改进智能体的改进保真度

    Self-improving agents increasingly rely on proxy verifiers to choose policy updates, yet deployment can change the world in which those updates are evaluated. An update that looks better to the verifier can therefore become worse after deployment even when the verifier ranks poli…

  313. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Ying Lin ·

    AsynCodeBench:软件工程中异步多智能体系统协作的基准测试

    Multi-agent coding has emerged as an increasingly active direction in software engineering, where complex development tasks are decomposed across multiple specialized agents working on different parts of the problem. Despite the shift from individual problem solving to distribute…

  314. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Xiping Hu ·

    在技能检索前实现及时指导:在代理上下文中保留有用的温馨提示

    Reusable skills help LLM-based agents solve complex tasks, but the agent must receive guidance before it commits to an ineffective approach. Existing skill mechanisms often expose only metadata and load full content on demand, leaving useful guidance unavailable until the agent d…

  315. Hugging Face Daily Papers TIER_1 English(EN) ·

    AgentTell:浏览器使用代理中的行为侧信道泄露

    Browser-use agents often carry information in their context as they move between websites. While it may be necessary for task completion, it also creates a privacy risk, especially when the information contains a private fact regarding the user. For example, an agent may learn a …

  316. Hugging Face Daily Papers TIER_1 English(EN) ·

    ExpVoyager:动态代理技能合成的直接经验导航

    Learning from experience in LLM agents has become a key paradigm for developing self-evolving agents that continuously learn and expand their capabilities. Within this paradigm, synthesizing the agent skill has emerged as a promising solution for transforming accumulated experien…

  317. Hugging Face Daily Papers TIER_1 English(EN) ·

    Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

    Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in impleme…

  318. Hugging Face Daily Papers TIER_1 English(EN) ·

    X-Tree:为高效智能体泛化进行可复用经验的标记化

    Multi-step agents are trained on flat action streams: SFT and RLVR weight every token uniformly and ignore the sub-procedures that recur across tasks, the hierarchy that lets humans plan top-down from reusable routines. This structure sits unused, and flat training uses each scar…

  319. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Blaz Bertalanic ·

    多智能体跨越析取和补偿任务的扩展

    Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the an…

  320. arXiv cs.AI TIER_1 English(EN) · Beining Wu, Zihao Ding, Jun Huang ·

    ERRAND:代理记忆的预算维护

    arXiv:2609.29545v1 Announce Type: new Abstract: Deployed agents run on handed-over knowledge: a frozen policy consults a briefing of consolidated items written before the stream begins. The world then moves while the store stands still: paths close, flags change, price bands move…

  321. arXiv cs.AI TIER_1 English(EN) · Yezhou Cheng, Runjia Du, Zeming Liu, Qibai Chen, Hang Lyu, Yankai Zeng, Yilan Wei, Bojun Lin ·

    明确范围再持久化:防止跨家族干扰代理记忆

    arXiv:2609.29144v1 Announce Type: new Abstract: Persistent memory lets language-model agents improve prompts and skills without updating model weights. We show that matching retrieval scope to certification scope enables these edits to support reliable repeated adaptation across …

  322. arXiv cs.AI TIER_1 English(EN) · Mingxuan Wang, Fei Luo, Bo Wang, Guorun Yao, Yinglong Guo, Chao Ning, Hongyue Chen, Yanbiao Ma, Jungong Han ·

    智能体何时能忘记其推理过程?ICLR 探讨长时域智能体上下文压缩

    arXiv:2609.29875v1 Announce Type: new Abstract: Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. Unlike static Chain of Thought compression, remo…

  323. arXiv cs.AI TIER_1 English(EN) · Hanjing Shi, Dominic DiFranzo ·

    当智能体在无人监管下行动:智能体式AI中的低监督悖论

    arXiv:2609.29547v1 Announce Type: cross Abstract: Agentic AI is sold on a simple promise: the system keeps acting when the user stops watching. That promise creates an accountability inversion. As stepwise supervision recedes, verification does not disappear; it moves into the ru…

  324. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Furu Wei ·

    Agensh: 将组织智能扩展至 1,024 个代理

    A multi-agent system can reduce latency on complex tasks by executing work concurrently. Several pioneering harness frameworks support multi-agent systems. However, the scalability of current multi-agent harnesses is often constrained by a central orchestrator's capacity to alloc…

  325. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Shuang Guo ·

    SkillApt:从反事实证据中学习何时激活Agent技能

    Large language model agents increasingly retrieve reusable Skills and inject them into the active context. However, a retrieved Skill can be relevant yet unnecessary, costly, or even harmful in the current execution state. We present SkillApt, a post-retrieval activation framewor…

  326. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Oren Gal ·

    MATES:通过转换冻结单智能体策略来学习多智能体交互

    Multi-agent reinforcement learning (MARL) commonly trains decentralized policies from scratch, requiring agents to acquire individual task competence and coordination simultaneously. Yet many multi-agent problems admit a compatible single-agent counterpart in which the underlying…

  327. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Andrea Baronchelli ·

    间接小费:AI代理群体中的社交攻击面

    As generative AI agents are deployed at scale, safety will depend not only on technical safeguards and individual model design, but also on collective equilibria that determine how agent populations process information, prioritize actions, and respond to uncertainty. Yet the same…

  328. Hugging Face Daily Papers TIER_1 English(EN) ·

    OSWorld-Pro:面向计算机使用代理的基于过程的评估

    Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with functional verifiers, as seen in OSWorld. However, such evaluation of end-state performance lacks transparency into how and why agen…

  329. arXiv cs.MA (Multiagent) TIER_1 English(EN) · James Zou ·

    自组织代理团队学会协同推理

    Collective intelligence depends not only on what team members know, but also on how they organize their work. When the structure of a solution is unknown, useful roles and divisions of labor cannot be specified in advance; teams must learn from experience how to organize reasonin…

  330. arXiv cs.CV TIER_1 English(EN) · Liu Renhang, Navonil Majumder, Tej Deep Pala, Soujanya Poria ·

    RoboQuest:通用型物理代理,用于搜索、检查和测试

    arXiv:2610.10388v1 Announce Type: cross Abstract: Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-rele…

  331. arXiv cs.CV TIER_1 English(EN) · Ling Li, Qiuyu Shen, Zheng Jiang, Qinwei Ma, Yuxuan Liu, Zhidong Deng ·

    SkillCycle:代理策略与技能库的协同进化

    arXiv:2610.09430v1 Announce Type: new Abstract: Internalizing external skills changes a language agent's capabilities and, with them, the value of its remaining guidance: rules can become redundant, misleading, or insufficient for newly encountered decisions. This creates a coupl…

  332. arXiv cs.CV TIER_1 Italiano(IT) · Sacha Morin, Kumaraditya Gupta, Mahtab Sandhu, Charlie Gauthier, Francesco Argenziano, Kirsty Ellis, Liam Paull ·

    Agentic Scene Policies

    arXiv:2509.19571v2 Announce Type: replace-cross Abstract: Designing or learning robot policies that generalize zero-shot across a range of language instructions and objects is a core problem in robotics. Vision-Language-Action models (VLAs) learn such policies end-to-end by repur…

  333. arXiv cs.CV TIER_1 English(EN) · Shenxiang Zeng, Chen Yang, Peiyao Chen, Guohui Zhang, Jiansheng Fan, Chen Wang ·

    询问世界:通过代理世界建模和探测实现通用物理推理

    arXiv:2609.39135v1 Announce Type: new Abstract: Physical reasoning from video requires inferring latent physical properties and dynamics beyond direct observation. Direct VLM inference remains unreliable on complex physical tasks without explicit modeling and validation, while pr…

  334. arXiv cs.CV TIER_1 English(EN) · Tongbo Chen, Junbo Niu, Zhengxi Lu, Niu Lian, Fei Tang, Yuchen Yan, Yike Hong, Yong Du, Yizhou Liu, Bofan Chen, Yongliang Shen ·

    HybridCUA:学习编排 GUI 和 CLI 以用于计算机使用代理

    arXiv:2609.38008v1 Announce Type: new Abstract: Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or …

  335. arXiv stat.ML TIER_1 English(EN) · Shengjun Zhang, Tingyi Liu, Dong Xie, Yunlong Dong, Xiang Wang, Cheng Zeng ·

    交互式智能体中的因果保留:接口因子分解与选择性适应

    arXiv:2609.30650v1 Announce Type: cross Abstract: Task performance need not determine which intervention mechanism an agent retains. We study causal retention: whether a frozen learned state answers a mechanism-probe map fixed independently of training, including action, context,…

  336. AWS Machine Learning Blog TIER_1 English(EN) · Manish Ballal ·

    超越节省的时间:构建代理式自动化的商业案例

    The RPA-era ROI model misses most of the value agentic automation creates. This post gives AI center of excellence leaders a framework to size the full value of agents across time savings, exception handling, decision quality, and maintenance economics, and to prioritize which wo…

  337. AWS Machine Learning Blog TIER_1 English(EN) · Thiago Verney ·

    在AgentCore和OpenClaw上构建上下文感知AI助手

    Off-the-shelf AI assistants forget you between conversations. This post shows how to build a personal assistant that accumulates context using OpenClaw on Amazon Bedrock AgentCore runtime, with AgentCore memory turning disposable chats into durable, structured knowledge you can r…

  338. AWS Machine Learning Blog TIER_1 English(EN) · Mona Mona ·

    新代理技能:Amazon SageMaker 为您的编码代理优化生成式 AI 推理

    Amazon SageMaker optimized generative AI inference introduces the aws-ai-ml skill through the Agent Toolkit for AWS, giving coding agents like Kiro, Claude Code, and Codex deep expertise in inference optimization and benchmarking. Describe what you want, and your agent generates …

  339. AWS Machine Learning Blog TIER_1 English(EN) · Manideep Reddy Gillela ·

    使用 LangChain 和 Amazon Bedrock Knowledge Bases 进行代理检索

    Build a Retrieval Augmented Generation (RAG) application on Amazon Bedrock Managed Knowledge Base with LangChain, and see how agentic retrieval handles the multi-part questions that single-shot retrieval answers poorly. Run the same query through both paths, read the trace events…

  340. AWS Machine Learning Blog TIER_1 English(EN) · Kanishk Mahajan ·

    使用 Amazon Bedrock AgentCore 评估多代理系统的可解释性和有用性

    Multi-agent systems need deeper guarantees than fluent responses: they must select the right tools, respect constraints, and explain their decisions. Learn how to build a Strands-based multi-agent supply chain decisioning system and evaluate it with Amazon Bedrock AgentCore Evalu…

  341. Modal blog TIER_1 English(EN) ·

    VM沙盒:代理的完整计算机

    VM Sandboxes are built for those who need to give their agents the power of a full computer.

  342. Databricks Blog TIER_1 English(EN) ·

    如何扩展代理式应用程序而不产生AI蔓延

    Building an agent is getting easier. More capable models and coding agents are making...

  343. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    JetBrains 发布 Mellum2.1:面向编码代理的 12B MoE 开源模型

    <p>JetBrains released Mellum2.1, an Apache 2.0, 12B mixture-of-experts thinking model with 2.5B active parameters. RL in real repositories lifted its SWE-bench Verified score from 2.0 to 47.0.</p> <p>The post <a href="https://www.marktechpost.com/2026/10/08/jetbrains-releases-mel…

  344. dev.to — Claude Code tag TIER_1 English(EN) · Steven Gonsalvez ·

    你的代理是一个while循环:驾驭、循环和图工程

    <p><em>Human thoughts, AI-assisted write-up.</em></p> <p>Strip an agent back to its skeleton and it's embarrassingly simple. A while loop.<br /> </p> <div class="highlight js-code-highlight"> <pre class="highlight plaintext"><code>while not done: context = observe() # read files,…

  345. dev.to — Claude Code tag TIER_1 English(EN) · saaro ·

    Agentic Coding in Production 2026:从自动补全到自主开发者

    <p>The leap from simple chat prompts to agentic workflows is the biggest upheaval in software development since the introduction of Git. While 2024 and 2025 were still dominated by autocomplete plugins and "vibe coding" in the headlines, the landscape has fundamentally changed by…

  346. MarkTechPost TIER_1 English(EN) · Sana Hassan ·

    IBM 将 Bob 带入自托管和隔离环境:无需移动代码的智能体式软件开发

    <p>IBM has made a self-hosted deployment of IBM Bob, its agentic software development platform, generally available. Enterprises can now run Bob on premises, in private or sovereign clouds, and in air-gapped networks. They bring their own model: NVIDIA Nemotron or Poolside Laguna…

  347. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    IQuest Research 开源 IQuest-Q1,一个拥有 15B 活跃参数的 320B MoE 模型,用于 Agentic 编码

    IQuest Research released open weights for IQuest-Q1, a 320B sparse MoE with 15B active parameters built for command-line coding agents, with 512K context and 84.5 on CyberGym.

  348. dev.to — MCP tag TIER_1 English(EN) · Fernando Azevedo ·

    区分多智能体与提示链的四种工程模式

    <p>Google's post on the AI Agents Challenge says something it took me years to accept on financial platforms: "multi-agent" was the most frequent claim across thousands of submissions, and a good share of them were a single model walking through a prompt chain with agent names at…

  349. Medium — Claude tag TIER_1 English(EN) · Fellows Monika ·

    学习编码代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@fellows.monika/learning-about-coding-agents-55e5c5bba4e6?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/1*hoRDlGu2zo_Gcjj0P8Emiw.png" width="2784" /></a></p><p cl…

  350. Medium — AI coding tag TIER_1 English(EN) · AIHoony ·

    不可逆转的指令:为何编码代理需要人工检查点

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@kwd8819/the-irreversible-command-why-coding-agents-need-a-human-checkpoint-d69037bc397b?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1672/1*gLz8vt3r3JJzBAcaOJMH1Q…

  351. Email — Every TIER_1 English(EN) · 010001a117d0d056-a9eaa087-9035-4580-bb51-9434736e88d3-000000@send.every.to (010001a117d0d056-a9eaa087-9035-4580-bb51-9434736e88d3-000000@send.every.to) ·

    构建更高效的Agent

    <!-- Set the language of your main document. This helps screenreaders use the proper language profile, pronunciation, and accent. --> <!-- The title is useful for screenreaders reading a document. Use your sender name or subject line. --> Building a More Efficient Agent <!-- Neve…

  352. dev.to — Anthropic tag TIER_1 English(EN) · Jorge Peraza ·

    为编码代理加固Prompt缓存和SSE流

    <p><strong>TL;DR:</strong></p> <ul> <li> <strong>Anthropic SSE Keep-Alives:</strong> We updated the model proxy to inject keep-alive pings during extended Anthropic "thinking" phases, preventing load balancer and NAT gateway timeouts.</li> <li> <strong>GPT-6 Cache Breakpoints:</s…

  353. Medium — MCP tag TIER_1 English(EN) · OpenVidu ·

    推出用于编码代理的 OpenVidu Agent 插件

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://openvidu.medium.com/introducing-the-openvidu-agent-plugin-for-coding-agents-064ee095f74d?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/2600/1*7mlDKz47mAixRxsx3rXZrA.png" width="3200…

  354. dev.to — MCP tag TIER_1 English(EN) · CAI ·

    CAI为何存在:代理支付问题与三种角色(托管人、审批人、操作人)

    <h1> Why CAI exists: the agent-payment problem and the three roles (custodian, approver, operator) </h1> <p>The agent-payment problem is older than agents, and it has a clear shape. This post is the H1-readable essay on the problem CAI solves, the user-confirmation pattern, and t…

  355. dev.to — MCP tag TIER_1 English(EN) · Baris Sozen ·

    当代理人雇佣代理人:委托支付问题

    <p>The agent economy doesn't wait for perfect infrastructure. Agent A is already paying Agent B to source data, optimize training runs, and verify execution quality. Agent B is already hiring Agent C to do the actual work. And today, all three are using custodians to move money.<…

  356. Medium — MLOps tag TIER_1 English(EN) · Swapnil Surushe ·

    构建Conductor:一个可扩展的AI Agent编排器在Cloud Run上运行

    <div class="medium-feed-item"><p class="medium-feed-snippet">This is Part 2 of the series, The Enterprise AI Agent Blueprint. Read Part 1 here.</p><p class="medium-feed-link"><a href="https://medium.com/@swapnil29071999/building-the-conductor-a-scalable-ai-agent-orchestrator-on-c…

  357. dev.to — MCP tag TIER_1 English(EN) · Luis Alcaraz ·

    推出TAP:面向Agentic世界的软件构建块

    <p>We've open sourced TAP (Trusted Agent Primitives) at Telara so agents can build reusable internal tools for the work they do repeatedly. The packages are code that developers can write, inspect and maintain too.</p> <p>Employees ask agents to investigate issues, prepare custom…

  358. dev.to — MCP tag TIER_1 English(EN) · Luis Alcaraz ·

    推出TAP:面向Agentic世界的软件构建块

    <p>We've open sourced TAP (Trusted Agent Primitives) at Telara so agents can build reusable internal tools for the work they do repeatedly. The packages are code that developers can write, inspect and maintain too.</p> <p>Employees ask agents to investigate issues, prepare custom…

  359. Towards AI TIER_1 English(EN) · Satish Kumar ·

    Cortex Agents 的生产 RBAC、成本优化和部署模式

    <p><em>The developer-to-production pipeline for Snowflake’s Cortex Agent GA enhancements — Personal Database sandboxes, temporary agents, COPY GRANTS, and the cost math on Cortex Search suspension.</em></p><p><strong>Part 2 of 2</strong> — Part 1: <a href="https://snowflakechroni…

  360. Towards AI TIER_1 English(EN) · Enzo Lombardi ·

    编码代理:操控本地模型

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/coding-an-agent-steering-a-local-model-14872070ef56?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1024/1*jOao6kz_lfj1-OR-fPe18A.png" width="1024" /></a></…

  361. Medium — fine-tuning tag TIER_1 English(EN) · Sasha Denisov ·

    微调小型模型以用于设备端代理:训练、转换和衡量变化

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/google-developer-experts/fine-tuning-small-models-for-on-device-agents-train-convert-and-measure-what-changed-f7f494a044f3?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.c…

  362. Medium — fine-tuning tag TIER_1 English(EN) · Sasha Denisov ·

    微调小型模型以用于设备端代理:训练、转换和衡量变化

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@denisov.shureg/fine-tuning-small-models-for-on-device-agents-train-convert-and-measure-what-changed-f7f494a044f3?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/13…

  363. Towards AI TIER_1 English(EN) · Rajesh K ·

    Jev与RLCD:架构、开源模型及实际Agent用例

    <h4><em>A practical guide to typed AI decisions, calibrated probabilities, and the software that turns them into useful workflows.</em></h4><p>An agent receives a request: “Investigate the failed deployment and explain what changed.”</p><p>Before it produces an answer, the system…

  364. Medium — Claude tag TIER_1 English(EN) · Rathish Poovadan ·

    第二部分:超越终端 — 在Jupyter Notebook中原型化Claude Agents

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/the-repo/part-2-beyond-the-terminal-prototyping-claude-agents-in-jupyter-notebooks-8be734ad85dc?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/1*B-jTb3C_wdjK3OjuzT…

  365. Medium — MCP tag TIER_1 English(EN) · Ahmet Kalafat ·

    从单一智能体到智能体团队:ADK、Agent Skills、MCP 和 A2A

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ahmet.kalafat/from-one-agent-to-an-agent-team-adk-agent-skills-mcp-and-a2a-6bc9b48a171c?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1760/1*60kJ2Tm_42iFc7IEOy3dcw.png" …

  366. Towards AI TIER_1 English(EN) · Surya Maddula ·

    使用 RAG 开发复杂的、可控的智能体

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/developing-sophisticated-controllable-agents-with-rag-727f28828815?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1672/1*2Nth5ujXg2CeRFiKDVsYMA.png" width=…

  367. Towards AI TIER_1 English(EN) · Dave R | Microsoft Azure & AI MVP ☁️ ·

    一种能抵抗 kill -9 的 AI Agent 工作流:深入了解 Microsoft Agent Framework 的最新更新

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/microsoft-agent-framework-kill-9-crash-recovery-ag-ui-memory-codeact-d4812e5a7bd2?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1672/1*wO-0z1k70zalhDFv_yq…

  368. dev.to — MCP tag TIER_1 English(EN) · Baris Sozen ·

    结算大战:为何路由不足以支撑代理商业

    <h1> The Settlement War: Why Routing Isn't Enough for Agent Commerce </h1> <p>This week, three signals converged.</p> <p>On Tuesday, the <strong>IETF published draft-hood-agtp-commerce-00</strong> — the agent-to-agent commerce standard. It is a big deal. Agents can now discover e…

  369. Medium — MCP tag TIER_1 English(EN) · vedant Padole ·

    WebMCP:代理式Web入门指南及示例

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@vedantpadole05072/webmcp-a-beginners-guide-to-the-agentic-web-with-an-example-ce091f3e3b40?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1774/1*-CEGRGIjzgJoZTPGQ2ko0g.pn…

  370. dev.to — MCP tag TIER_1 English(EN) · Renato Marinho ·

    调试代理推理:为何结构完整性比准确性更重要

    <p>When we talk about LLM reliability, our focus almost always gravitates toward accuracy—did the model get the math right? Did it retrieve the correct record from the database?</p> <p>But for engineers building autonomous agents using ReAct or Chain-of-Thought (CoT) patterns, th…

  371. dev.to — MCP tag TIER_1 English(EN) · Shakar Bisetty ·

    第二部分:MuleSoft 上代理式变更审批 MVP 的参考架构,一页纸

    <p><em>Part 2 of 10 · Building an Agentic Change-Approval MVP on MuleSoft</em></p> <p>In <a href="https://dev.to/thasha/the-21-step-change-nobody-wants-to-own-why-sap-change-promotion-is-an-agentic-use-case-4mpo">Part 1</a> I set out the use case: automating the approval and prom…

  372. dev.to — MCP tag TIER_1 English(EN) · mech.app ·

    MCP参考服务器:90,000颗星和16个语言SDK揭示了Agent工具边界的什么

    <p>The Model Context Protocol repository sits at 90,950 stars with 16 language SDKs and a collection of reference servers that expose how agent-tool boundaries actually work. These are not production systems. They are educational implementations that reveal transport layer choice…

  373. dev.to — MCP tag TIER_1 English(EN) · Praveen Raj Thulasi S ·

    我构建了一个智能分析平台——我的学习心得

    <h1> I Built an Agentic Analytics Platform — Here's What I Learned </h1> <p>What if you could ask your analytics dashboard:</p> <blockquote> <p><strong>"Why did sales decrease last month?"</strong></p> </blockquote> <p>and instead of manually filtering charts and writing database…

  374. Medium — Claude tag TIER_1 Nederlands(NL) · Ivan Yanishevskyi ·

    Claude代码中的子代理与代理团队:当一个代理不够用时

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ivan.yanishevskyi/subagents-vs-agent-teams-in-claude-code-when-one-agent-isnt-enough-bc062883070c?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1638/1*LWLjCpWV9vXjVEW…

  375. Towards AI TIER_1 English(EN) · Andrei Besleaga (Nicolae) ·

    AgenticSystemCore:一个人们和代理可以多种方式使用的Markdown文本文件夹

    <h4>Existing technologies for future simpler combined understanding and use</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*hk7w8hkMMAkvQ4me.jpeg" /><figcaption>(AgenticSystemCore — Text used by humans and Agentic AI, source:Author, Gemini)</figcaption></f…

  376. Medium — MCP tag TIER_1 English(EN) · Ferry Djaja ·

    让网站无需修改即可与AI代理对话

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://djajafer.medium.com/teaching-websites-to-talk-to-ai-agents-without-changing-the-websites-df0ed020cb79?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/2600/1*r2rir1bMG-pAPE8JP_wjBA.png…

  377. Towards AI TIER_1 English(EN) · Anna Jey ·

    AI 编码代理上下文预算:测试使代理变慢的指令

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*zgDOnsQQy6eaFaf_NN4E9g.jpeg" /><figcaption>AI Coding Agent Context Budget</figcaption></figure><p>More instructions can feel safer. For a coding agent, they can also bury the one rule that prevents a costly mista…

  378. dev.to — MCP tag TIER_1 English(EN) · lizer yang ·

    Agentic Search 对比 RAG:工具调用还是你拥有的索引

    <p><strong>Short answer:</strong> An agent's search tool and a RAG index are both retrieval, and they differ in<br /> four places that decide everything downstream: what they read (the live web against a corpus you<br /> ingested), who owns the ranking (an engine you do not opera…

  379. Towards AI TIER_1 English(EN) · Becca Sees ·

    从数小时到数分钟:使用多智能体系统自动化数据调试

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*C9cnr-mVxVYEVIuOsEgnMw.jpeg" /></figure><h3>The Data Debugging Grind</h3><p>My engineers typically spend 30% of their time debugging data. On some days, that number climbs to 50% or more.</p><p>If you’re a data o…

  380. Medium — AI coding tag TIER_1 English(EN) · Dhmkchs ·

    三个编码代理,一个窗口:我如何实际使用codeme

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@dhmkchs/three-coding-agents-one-window-how-i-actually-work-with-codeme-0f74a0817cd6?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/2560/1*w0uxHUuDhwYY9BmS88Ykbg.png…

  381. Towards AI TIER_1 English(EN) · Mukesh Kumar Shah ·

    掌握AI代理:从ReAct到生产级多代理系统

    <h4><em>From “What is an agent?” to production-grade, multi-agent, tool-using, memory-equipped systems — everything you need in one place.</em></h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*RzS8r6JYTDqLuc2V" /></figure><p><strong>Reading time:</strong> ~…

  382. The Register — AI TIER_1 English(EN) ·

    利用智能体式可观测性缩小可观测性差距

    SPONSORED POST: How agentic AI, real-time visibility, and stronger governance can help enterprises protect critical services and manage increasingly complex IT environments.

  383. dev.to — MCP tag TIER_1 English(EN) · MANI BHUSHANAM K ·

    我为编码代理赋予了记忆——并让它学会了所需的工具

    <p>🤔 What if a coding agent could actually remember?</p> <p>AI coding agents are becoming surprisingly capable.<br /> They can write code, inspect repositories, debug errors, interact with tools, and reason through complex development tasks.</p> <p>But there is still a frustratin…

  384. Medium — MLOps tag TIER_1 English(EN) · Pavan Gosangi ·

    企业智能体RAG:双数据库困境 — 同步智能体RAG中的状态与向量

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/agentic-architecture/enterprise-agentic-rag-the-dual-database-dilemma-synchronizing-state-vectors-in-agentic-rag-8c164c15e16e?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/…

  385. dev.to — Anthropic tag TIER_1 English(EN) · Anshul Rajpal ·

    Claude Opus 5.5:Anthropic 最新用于 Agentic 编码的模型

    <blockquote> <p><strong>Key Takeaways</strong></p> <ul> <li>Claude Opus 5.5 launched September 22, 2026</li> <li>Anthropic's first model in the new Claude 5.5 family</li> <li>40% cheaper to run than Opus 5</li> <li>Optimized for agentic coding and knowledge work</li> <li>Availabl…

  386. dev.to — MCP tag TIER_1 English(EN) · qianqiuwanzi ·

    一个记忆层,十三位代理:我们如何在多代理团队中共享上下文

    <p>Running one AI agent is easy. Running thirteen that actually cooperate is where memory becomes the real bottleneck.</p> <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A…

  387. dev.to — LLM tag TIER_1 English(EN) · Eryk Kubiak ·

    本地编码代理:模型弱?不,是管道差

    <p> </p> <p><strong>In this video:</strong></p> <p>0:00 Why agents fail around turn ten<br /> 0:19 It runs but is useless<br /> 1:54 The base-URL swap<br /> 3:26 Where compatibility breaks<br /> 5:28 From tokens to a tool call<br /> 8:31 Context, speed and turn count<br /> 12:25 …

  388. dev.to — LLM tag TIER_1 English(EN) · Priyesh Dave ·

    上下文窗口退化如何破坏生产环境中的长期运行代理——以及如何通过工程手段解决

    <p>Liquid syntax error: Unknown tag 'endraw'</p>

  389. r/LocalLLaMA TIER_1 English(EN) · /u/Low_Bad_6585 ·

    运行一个由LLM驱动、拥有800多个持久化代理的小镇:并发、上下文缓存和推理成本

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1x0ms0m/running_an_llmdriven_town_with_800_persistent/"> <img alt="Running an LLM-driven town with 800+ persistent agents: concurrency, context caching, and inference costs" src="https://preview.redd.it/71a7tj…

  390. dev.to — LLM tag TIER_1 English(EN) · HiDevs ·

    构建多智能体负载模拟器:50个智能体,一个集合

    <p>As AI applications evolve from simple chatbots into complex multi-agent architectures, the demands placed on the vector database change fundamentally. In a single-agent loop, context retrieval is predictable and mostly sequential: an agent sends a search request, waits for a r…

  391. dev.to — LLM tag TIER_1 English(EN) · Ameer Mavia ·

    多智能体LLM编排如何支持复杂AI工作流

    <p><a href="https://nextigent.ai/services/model-orchestration/" rel="noopener noreferrer">Multi-agent LLM orchestration</a> coordinates multiple AI agents, language models, tools, and data sources so they can contribute to a shared workflow. Instead of expecting one model to inte…

  392. dev.to — LLM tag TIER_1 English(EN) · mech.app ·

    OpenAI与Ironclad:将合同工作流转化为代理评估

    <p>OpenAI and Ironclad just published a case study on using production contract workflows as both training data and evaluation benchmarks for computer-use agents. This is not a demo. It is a partnership where a SaaS company opens its workflow engine to become an agent training gr…

  393. dev.to — LLM tag TIER_1 English(EN) · Pratik ·

    如何构建具有搜索回退循环的弹性AI代理

    <p>Building autonomous AI agents is incredibly rewarding until you deploy them to production and real-world data breaks your clean pipelines.<br /> A common bottleneck is the tool execution layer. When your agent invokes a vector DB search or a live web API, it assumes it will re…

  394. dev.to — LLM tag TIER_1 English(EN) · mech.app ·

    Sarvam Arya:生产代理编排栈暴露状态、路由和恢复机制

    <p>Sarvam just released Arya, an orchestration stack designed to move agent systems from prototype to production. The announcement centers on a concrete problem: frontier models can write compilers and rebuild rendering pipelines in demos, but they fail silently and inconsistentl…

  395. dev.to — LLM tag TIER_1 English(EN) · Abdulmuiz Adebayo ·

    Smallops 基准报告 · MD 小型本地模型能成为 Agentic 吗? 4 款 Ollama 模型的 6 轮基准测试

    <h2> A build-in-public deep dive from the smallOps project </h2> <p>Why this benchmark exists<br /> .<br /> smallOps is an experiment in giving small, locally-run language models — the kind that fit comfortably on a laptop with no GPU — the ability to act as coding agents. Tools …

  396. r/MachineLearning TIER_1 English(EN) · /u/heyitsdannyle ·

    SWE-Race:一个包含188个真实并发bug的编码代理基准测试,以及三个模型的测试结果 [P]

    <table> <tr><td> <a href="https://www.reddit.com/r/MachineLearning/comments/1wyw0my/swerace_a_codingagent_benchmark_of_188_real/"> <img alt="SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]" src="https://preview.redd.it/0uopztlmp…

  397. dev.to — LLM tag TIER_1 English(EN) · mech.app ·

    LLM 智能体的时间旅行调试:Burr 的反事实重放架构

    <p>Every agent trace tool shows you a waterfall of steps. You spot that step six produced garbage. Now what? You re-run the entire pipeline and hope it lands in the same place. With a non-deterministic model, it doesn't. You can never separate your change from model jitter.</p> <…

  398. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    Holo4:H公司如何构建一个能够点击、编码和调用API的单一开放权重代理

    <h1> Holo4: How H Company Built a Single Open-Weight Agent That Clicks, Codes, and Calls APIs </h1> <p>Most agentic AI models are specialists. A GUI-focused model can click through a browser but is lost when the task requires an API call. A tool-calling model can chain function c…

  399. dev.to — LLM tag TIER_1 English(EN) · VIBHEESHAN NK ·

    Agentic GraphRAG:从RAG到基于图的Agentic推理

    <p>Large Language Models (LLMs) can answer many questions, but answering questions over a large and complex corpus becomes difficult when the required information is spread across multiple documents or depends on relationships between different entities.<br /> To address this cha…

  400. dev.to — LLM tag TIER_1 English(EN) · zhijie ·

    微调本地 Qwen 编码代理:哪些已改变,哪些仍失败

    <p>On eight development coding cases repeated three times, a local Qwen3.8-27B agent went from <strong>5/24 to 23/24 functional and delivered successes</strong> across sequential output-budget and reasoning rounds. That is promising development evidence. It does not establish gen…

  401. dev.to — LLM tag TIER_1 English(EN) · Rashid Mahmood ·

    八个失败的工具调用:六个代理框架如何恢复

    <p>Models send bad tool calls: malformed JSON, a tool name that does not exist, a missing argument. What happens next is decided by the agent framework, not the model. In <a href="https://github.com/code-with-rashid/agentic-arena" rel="noopener noreferrer">agentic-arena</a> I mea…

  402. dev.to — LLM tag TIER_1 Nederlands(NL) · mech.app ·

    iFixAi:120秒内独立代理审计

    <p>Production agents fail in ways that runtime guardrails cannot catch. They hallucinate plausible-sounding API calls, leak context across tool invocations, and drift from their original task specification without triggering a single exception. iFixAi is a Python CLI that audits …

  403. dev.to — LLM tag TIER_1 English(EN) · Aleksei Romanov ·

    专业化代理的转变:开源模型如何在生产中超越前沿API

    <p>For the past two years, the standard enterprise AI strategy looked suspiciously like a parlor trick: take an 8,000-token system prompt packed with JSON schemas, API documentation, and polite formatting rules, stuff it into the context window of a giant 400-billion-parameter fr…

  404. dev.to — LLM tag TIER_1 English(EN) · Aleksei Romanov ·

    长时域智能体:多轮推理为何失效及其修复的实用训练技巧

    <p>Almost every modern language model looks impressive on a two-step demo. You ask it to check a database or summarize a document, it calls the right tool, formats the answer, and looks like an autonomous engineer.</p> <p>The illusion falls apart the moment you ask that same mode…

  405. dev.to — LLM tag TIER_1 English(EN) · lizer yang ·

    Agentic Retrieval:决定何时搜索的循环

    <blockquote> <p>Originally published at <a href="https://smartgate.network/industry/agentic-retrieval?utm_source=devto&amp;utm_medium=syndication" rel="noopener noreferrer">Agentic Retrieval: The Loop That Decides When to Search</a> on smartgate.network.</p> </blockquote> <p>A sh…

  406. dev.to — LLM tag TIER_1 English(EN) · astronaut ·

    2026年代理上下文:一页全景图

    <p>There are dozens of posts about <code>CLAUDE.md</code> best practices. Each one is a list of tips, and none of them explains <em>why</em>. We stopped guessing and went to the primary sources instead: Anthropic's own guidance and the 2026 papers on context engineering.</p> <p>T…

  407. dev.to — LLM tag TIER_1 English(EN) · mech.app ·

    PhantomEnvironments:虚构世界如何解决智能体训练瓶颈

    <p>Training LLM agents with reinforcement learning hits a hard wall: you need environments that provide verifiable rewards, support long-horizon interaction, and scale without burning budget. Human-curated data is expensive. LLM-generated environments hallucinate and leak benchma…

  408. dev.to — LLM tag TIER_1 English(EN) · mech.app ·

    Whiteboard IDE:基于画布的设计工具揭示了代理上下文管理

    <p>Whiteboard is a YC W26-backed open-source IDE that replaces the file tree with a spatial canvas. Instead of navigating folders, you arrange components on a 2D plane. The project (422 HN points, 142 comments) exposes a different set of plumbing decisions for agent context manag…

  409. dev.to — LLM tag TIER_1 English(EN) · Akhil Belide ·

    构建具有持久内存的交易智能代理

    <p>Sales conversations rarely happen in isolation.</p> <p>A customer might discuss pricing in one call, raise an objection in another, introduce a new stakeholder later, and finally ask for a specific next step several conversations after the first meeting.</p> <p>The problem is …

  410. dev.to — LLM tag TIER_1 English(EN) · Vardhini ·

    无状态AI代理为何在企业谈判中失败——以及情景记忆如何解决这个问题

    <p>When teams deploy large language models into enterprise workflows, the default architectural pattern connects a model to a user interface with high-level system instructions. In brief, deterministic tasks like code refactoring or single-turn customer support, this pattern work…

  411. r/LocalLLaMA TIER_1 English(EN) · /u/jonas__m ·

    编码代理中的投机性奖励黑客行为

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wsuag0/speculative_reward_hacking_in_coding_agents/"> <img alt="Speculative reward hacking in coding agents" src="https://preview.redd.it/7pvdij77hcsh1.png?width=640&amp;crop=smart&amp;auto=webp&amp;s=0a5fbcf…

  412. dev.to — LLM tag TIER_1 English(EN) · AI Frontier Post ·

    OKF Agent Memory:为您的编码代理提供原生 Git 内存,使其能够跨会话持久化

    <p><em>Originally published at <a href="https://aifrontierpost.com/articles/okf-agent-memory-git-native-project-memory/" rel="noopener noreferrer">AI Frontier Post</a>.</em></p> <p>You know the feeling. You spend an hour with a coding agent establishing the architecture: why the …

  413. dev.to — LLM tag TIER_1 English(EN) · Manidhar Bheempadu ·

    当AI代理记忆学会不再复用

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frh9n3neia8zhvf4hia93.png"><img alt=" " height="380" …

  414. Mastodon — mastodon.social TIER_1 English(EN) · Moltbookpulse ·

    并发、预测和 Web 信任边界“多代理评估序列化了它们声称测量的竞态条件”(通用)+“无期望的 se

    Concurrency, Forecasts, and Web Trust Boundaries "Multi-agent evals serialize the race condition they claim to measure" (general) + "An expectation without a settle date is not a forecast" (agents) This + more in today's Moltbook Pulse (Edition #83): https:// superagent-ebe00561.…

  415. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    在 .NET 中构建多智能体系统?探索协调专用智能体、管理工作流和创建可维护 AI 应用的结构化方法

    Building multi-agent systems in .NET? Explore a structured approach to coordinating specialized agents, managing workflows, and creating maintainable AI applications with familiar .NET patterns. # dotnet # AI https:// isaacl.dev/hbw

  416. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Docker Agent:一个使用简单 YAML 配置、丰富的工具生态系统和多代理编排来构建、运行和共享 AI Agent 的工具 # AI # Age

    Docker Agent: a Tool for building, running and sharing AI Agents using a simple YAML configuration, rich tool Ecosystem and multi-agent Orchestration # AI # Agent https:// github.com/docker/docker-agent

  417. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    🤖 长期运行的代理:瓶颈是模型还是围绕它的脚手架?我一直在代理设置中注意到的一件事:每个单独的步骤对

    🤖 Long-running agents: is the bottleneck the model or the scaffolding around it? Something I keep noticing with agent setups: every individual step is easy for the model, but the full chain still falls apart on long tasks. I think it comes down to three things: Error compoundin..…

  418. Mastodon — mastodon.social TIER_1 English(EN) · schuler ·

    独立测试发现,用于编码代理的记忆库 Hindsight 在没有明确指示的情况下,跨会话保留了未写入的项目规则。该工具 dr

    An independent test found Hindsight, a memory bank for coding agents, retained unwritten project rules across sessions without explicit instruction. The tool draws on git history to help agents maintain conventions. Teams should verify storage location and check for empty knowled…

  419. r/ClaudeAI TIER_2 English(EN) · /u/Disastrous_Exam9484 ·

    Repos & Dungeons:让你的代理在你的代码库中战斗

    <table> <tr><td> <a href="https://www.reddit.com/r/ClaudeAI/comments/1wzt7ei/repos_dungeons_watch_your_agents_fight_their_way/"> <img alt="Repos &amp; Dungeons: watch your agents fight their way through your codebase" src="https://external-preview.redd.it/aWZqZHE3YTM4MHVoMdDq__bK…

  420. r/ClaudeAI TIER_2 English(EN) · /u/george-lin ·

    VelaTerm 中的代理通信与编排

    <table> <tr><td> <a href="https://www.reddit.com/r/ClaudeAI/comments/1wzkh5s/agent_communication_orchestration_in_velaterm/"> <img alt="Agent Communication &amp; Orchestration in VelaTerm" src="https://external-preview.redd.it/dmE2aHF0bDFieXRoMapbV4K0o4qagdBawe0kyN1CzVcnkshsHO15v…