PulseAugur
中
实时 07:16:47
English(EN) Auditing Action Settlement in LLM Agent Environments: Order, Progress, and Replay

新研究探讨 LLM Agent 的审计、安全和决策 · 已追踪 10 个来源

多篇研究论文正在探索用于评估和改进大型语言模型 (LLM) Agent 性能和安全性的新颖方法。这些研究引入了用于审计通信的框架、分析表示转换以检测安全风险以及开发用于 Agent 恢复的因果评估方法。此外,研究还侧重于空间策略、科学计算的可执行检查以及认知行动的概念,以增强基于 LLM 的系统的决策能力。研究结果强调了理解 Agent 内部状态、通信内容和决策过程对于降低风险和提高可靠性的重要性。 AI

影响 这些研究为 LLM Agent 引入了新颖的评估框架和安全机制,有可能提高其可靠性并降低复杂应用中的风险。

排序理由 多篇在 arXiv 上发表的学术论文,介绍了 LLM Agent 的新框架和方法论。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 149 个来源。 我们如何撰写摘要 →

新研究探讨 LLM Agent 的审计、安全和决策 · 已追踪 10 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇在 arXiv 上发表的学术论文,介绍了 LLM Agent 的新框架和方法论。
Source corroboration
149 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, safety, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
19 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+46 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [149]

  1. arXiv cs.CL TIER_1 English(EN) · Sergei Polezhaev, Barys Liskavets, Ori Press, Alexander Golubev ·

    从任务结果中为LLM代理训练顾问

    arXiv:2610.09858v1 Announce Type: new Abstract: Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment. Prior work has shown that natural-language feedback can help these agents revise their decisions du…

  2. arXiv cs.CL TIER_1 English(EN) · Masaaki Nakatsu (AO, Inc. / OrbLabs AG), Reno Wang (AO, Inc.) ·

    逻辑与个性的解耦:边缘LLM智能体对上下文污染的结构免疫性

    arXiv:2610.09772v1 Announce Type: new Abstract: Small language-model agents on edge devices must hold a persona and reason correctly at once, inside one context window that fills with conversational history and persona instructions. We study what happens to the logical part of su…

  3. arXiv cs.LG TIER_1 English(EN) · Yanjun Chen, Yirong Sun, Hanlin Wang, Jinghan Wang, Xinming Zhang, Xiaoyu Shen, Wenjie Li, Wei Zhang ·

    痕迹即状态:LLM 智能体团队的精确信用分配

    arXiv:2603.06859v4 Announce Type: replace Abstract: Credit assignment for a team of LLM agents, what each message was worth, has mostly been treated as prediction: a learned critic, a trajectory-level score, or an agent ablation stands in for the counterfactual that defines credi…

  4. arXiv cs.LG TIER_1 English(EN) · Xijie Gong, Tingxu Han, Jiahao Zhang, Wei Song, Ziqi Ding, Hanqi Yan, Youcheng Sun, Lijie Hu ·

    Agentic LLM 如何决定调用工具?由抑制塑造的工具调用向量

    arXiv:2610.09624v1 Announce Type: cross Abstract: Tool calling, invoking external tools on demand, is central to agentic LLMs, yet the mechanism that decides whether a model calls a tool or responds directly remains poorly understood. Agentic prompts are long and heavily scaffold…

  5. arXiv cs.CL TIER_1 English(EN) · Mohamed Dhouib, Clement Elliker, Alexi Canesse, Ma\"el Jenny, Lucas-Andrei Thil, Mahammed El Sharkawy, Sonia Vanier, Elie Bursztein ·

    RAISED:用于 LLM Agent 对抗提示注入的自蒸馏方法

    arXiv:2610.06401v2 Announce Type: replace-cross Abstract: Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of ge…

  6. arXiv cs.CL TIER_1 English(EN) · Harshita Chopra, Kshitish Ghate, Aylin Caliskan, Tadayoshi Kohno, Chirag Shah, Natasha Jaques ·

    超越合作模拟器:为LLM代理的鲁棒评估生成逼真用户画像

    arXiv:2605.12894v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) agents are increasingly deployed in settings where they interact with diverse users, including those who are unclear, impatient, or reluctant to share information. However, collecting real intera…

  7. arXiv cs.CL TIER_1 English(EN) · Jun Zhao, Leiming Fu, Yanbo Wen, Yiding Wang, Xuantong Liu, Yang Shu, Yuyang Lu, Xuanran Xing, Jingqi Tong, Hao Xu, Qi Zhang, Xuanjing Huang ·

    LiveMACE:动态市场中 LLM Agent 能力的感知过程评估

    arXiv:2610.09872v1 Announce Type: cross Abstract: Evaluating agents by outcomes alone can obscure the capabilities that produce them. This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changi…

  8. arXiv cs.CL TIER_1 English(EN) · Hanwen Li, Jinhao Duan, Guanhua Zhu, Junchi Lu, Bo Shen, Chenxi Yuan, Kaidi Xu ·

    从不确定性到行动:学习驾驭 LLM 智能体

    arXiv:2610.09115v1 Announce Type: cross Abstract: Steering an LLM agent means deciding whether to correct it, at which step, and with which mechanism. Uncertainty is often used to decide when to correct an agent, but whether it can guide these decisions remains unclear. We steer …

  9. arXiv cs.AI TIER_1 English(EN) · Brendan King, Farima Fatahi Bayat, Jean-Flavien Bussotti, Pouya Pezeshkpour, Estevam Hruschka ·

    置信推理图:LLM智能体的结构化置信度估计

    arXiv:2610.07948v1 Announce Type: new Abstract: When using an LLM agent in a consequential domain, making an informed decision about whether to trust its output or intervene requires calibrated confidence in the agent's success. Confidence estimation for agents is difficult becau…

  10. arXiv cs.AI TIER_1 English(EN) · Toby D. Pilditch, Konstantinos Voudouris, Alexandra Abbas, Cozmin Ududec ·

    Transect:为长时程LLM代理评估保留可观测性

    arXiv:2610.08364v1 Announce Type: new Abstract: Frontier AI evaluations increasingly use open-ended, agentic, long-horizon tasks whose transcripts can span hundreds of pages of outputs and actions from complex multi-agent networks. The observability envelop-the range of what eval…

  11. arXiv cs.AI TIER_1 English(EN) · Yibo Li, Jinhang Qiu, Zhi Zheng, Qianyun Guo, Jiaying Wu, Shuo Ji, Bryan Hooi ·

    Learn2Play Bench:大型语言模型智能体在陌生环境中从经验中学习得如何?

    arXiv:2610.08215v1 Announce Type: new Abstract: Learning from experience is essential for LLM agents to adapt to unfamiliar and dynmaic environments. Evaluating this ability is therefore important for understanding how effectively agents acquire and use new knowledge. Existing be…

  12. arXiv cs.AI TIER_1 English(EN) · Yunju Kang, Seonghyeon Cho, Irene Li, Yeo-Chan Yoon, Chanjun Park ·

    POLAR:面向工具调用LLM智能体的本体引导风险防护

    arXiv:2610.08082v1 Announce Type: new Abstract: LLM tool-use agents operate in dynamic environments where many actions carry operational risk. However, most safety mechanisms react only after errors manifest. Existing pre-emptive approaches either fine-tune the agent on chain-of-…

  13. arXiv cs.AI TIER_1 English(EN) · Chenglin Yang ·

    评估堆栈,而非层级:用于代理操作的确定性与 LLM 门是否会独立失败?

    arXiv:2610.07359v1 Announce Type: new Abstract: Runtime gates for agent tool calls are stacked on the assumption that their errors multiply. We test it on 1,119 labelled agent actions from three corpora, without an adaptive adversary. The stack has one deterministic rule layer an…

  14. arXiv cs.LG TIER_1 English(EN) · Jiaju Chen, Min Yang, Jinghua Piao, Xiaochong Lan, Xu Xia, Xiangnan He, Yong Li ·

    VETTA:为多轮 LLM 代理协调轮次和令牌级别的信用分配

    arXiv:2610.08402v1 Announce Type: new Abstract: Multi-turn LLM agents often receive sparse task feedback across several interactions, while generating each response token by token. This creates two related credit-assignment questions: which responses helped achieve the outcome, a…

  15. arXiv cs.LG TIER_1 English(EN) · Yiming Xu, Hongyue Yu, Beihua Yang, Zihan Chen, Yixin Liu, Zhen Peng, Bin Shi, Bo Dong, Chao Shen, Irwin King, Qinghua Zheng ·

    DecepEval:评估大型语言模型(LLM)代理欺骗行为的基准测试

    arXiv:2610.07967v1 Announce Type: new Abstract: As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment. Existing evaluations show that LLM agents can deceive, but of…

  16. arXiv cs.AI TIER_1 English(EN) · Wei Shi, Ziheng Peng, Sihang Li, Xiting Wang, Xiang Wang, Mengnan Du, Na Zou ·

    调用还是不调用:诊断大型语言模型代理的内在过度调用偏见

    arXiv:2605.18882v2 Announce Type: replace-cross Abstract: LLM agents exhibit a consistent tendency to over-call, invoking tools even in situations where none is needed. On the When2Call benchmark, six models from three families show high call accuracy but much lower no-call accur…

  17. arXiv cs.AI TIER_1 English(EN) · Chubin Zhang, Zhenglin Wan, Xingrui Yu, Jingxuan Wu, Qi Wen, Pengfei Zhou, Wangbo Zhao, Ivor Tsang ·

    校准并非控制:LLM-Agent监督的干预价值

    arXiv:2606.21399v2 Announce Type: replace Abstract: Runtime oversight often intervenes when an LLM agent's calibrated failure score crosses a threshold. Yet states with the same failure risk can differ in whether intervention helps. Strictly increasing recalibration preserves the…

  18. arXiv cs.AI TIER_1 English(EN) · Heewon Park, Somin Im, Minhae Kwon ·

    面向长时序决策的成本感知LLM智能体的选择性批判

    arXiv:2610.07335v1 Announce Type: cross Abstract: Improving the reliability of large language model (LLM) agents in long-horizon decision-making remains a key challenge. When deployed as autonomous agents interacting with complex environments, early mistakes can propagate through…

  19. arXiv cs.AI TIER_1 English(EN) · Xinran Zheng, Xin Fan Guo, Zhiqiang Hao, Fan Yang, Xingzhi Qian, Jiawei Du, Jinfeng Xu, Zheng Xing, Shuo Yang, Xingjun Wang ·

    APEX: LLM智能体在执行边界的主动防护

    arXiv:2610.06966v1 Announce Type: cross Abstract: Indirect prompt injection (IPI) hides adversarial instructions in content that large language model (LLM) agents read at runtime. As agents compose heterogeneous capability units, including Tools, MCP servers, and Skills, the carr…

  20. arXiv cs.AI TIER_1 English(EN) · Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein, Martin Gubri, Seong Joon Oh ·

    瓶中特工:LLM智能体能否将其能力转化为廉价、可扩展的产物?

    arXiv:2610.08775v1 Announce Type: new Abstract: Large language models (LLMs) can solve many narrow tasks, but querying them separately for millions of related instances can be prohibitively expensive. Can LLM agents autonomously create cheaper solutions for such workloads? We cal…

  21. Hugging Face Daily Papers TIER_1 English(EN) ·

    POLAR:面向工具调用LLM智能体的本体引导风险防护

    LLM tool-use agents operate in dynamic environments where many actions carry operational risk. However, most safety mechanisms react only after errors manifest. Existing pre-emptive approaches either fine-tune the agent on chain-of-thought deliberation or compile natural-language…

  22. Hugging Face Daily Papers TIER_1 English(EN) ·

    DecepEval:评估大型语言模型(LLM)代理欺骗行为的基准

    As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment. Existing evaluations show that LLM agents can deceive, but often examine isolated scenarios or narrowly defin…

  23. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Adel Bibi ·

    BazaarBench:LLM代理运行的去中心化C2C市场中的委托安全

    In decentralized consumer-to-consumer (C2C) marketplaces, people list goods, negotiate with strangers, and rate one another, so trust rests on reputation. Large language model (LLM) agents now act for users, raising risks to their money, privacy, and reputation. We introduce Baza…

  24. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Nayonika Sen ·

    Attention Tax, Handoff Tax: 一个关于多智能体LLM系统何时有益的风格化模型

    Recent work on multi-agent LLM systems reaches sharply different conclusions: some results show that a single agent with the same information and compute should dominate a delegated system, others that multi-agent gains grow with task depth. We argue that much of the disagreement…

  25. arXiv cs.CL TIER_1 English(EN) · Yekun Chai, Qiwei Peng, Haoyi Xiong ·

    落子无悔非胜局:XiangqiBench 用于棋类大模型智能体的闭环评估

    arXiv:2610.02425v1 Announce Type: new Abstract: Static evaluations credit a language model for naming the right move, but an agent must carry a plan through to a verified outcome while an opponent responds. We introduce XiangqiBench, an executable benchmark that measures this dif…

  26. arXiv cs.AI TIER_1 English(EN) · Changxiu Ji, Amy Lu, Qizheng Zhang, Kunle Olukotun ·

    Sentry:在测试时从LLM代理故障中学习恢复

    arXiv:2610.02994v1 Announce Type: cross Abstract: LLM agents often fail mid-task due to invalid tool calls, repeated actions, or poorly grounded reasoning, and learning from these failures is a path to reliability. We find that how failure knowledge reaches the agent matters as m…

  27. arXiv cs.CL TIER_1 English(EN) · Zhuowen Liu ·

    通过你训练过的测试:重新评估LLM代理的提示注入检测器

    arXiv:2610.03448v1 Announce Type: cross Abstract: LLM agents increasingly screen tool outputs with small prompt-injection detectors, and teams choose among detectors by their scores on public benchmarks. We ask whether those scores predict how a detector behaves inside an agent. …

  28. arXiv cs.CL TIER_1 English(EN) · Jonghyun Song, Haewon Park, Jeonghoon Shim, Woojung Song, Yohan Jo ·

    现实世界中的来源偏好:LLM代理如何偏好按来源选择项目,以及如何减少这种偏好

    arXiv:2610.03195v1 Announce Type: new Abstract: As LLM agents decide on users' behalf which product to buy, which hotel to book, or which paper to cite, a preference for items from certain sources (the sites or services they come from) shapes what users receive and which sources …

  29. arXiv cs.CL TIER_1 English(EN) · Ziang Ni, Peng Zou ·

    沉默的异议:屈从于多数的LLM代理仍代表其原始前提

    arXiv:2610.02702v1 Announce Type: new Abstract: Multi-agent debate is increasingly used to reach consensus among LLM agents, yet agents often yield to a unanimous majority. When an agent changes its answer, has it changed its mind or only its statement? We study this with two-hop…

  30. arXiv cs.AI TIER_1 English(EN) · Hang Cui ·

    超越预定义接收器:LLM Agent 的安全感知依赖分析

    arXiv:2610.03014v1 Announce Type: cross Abstract: Large language model (LLM)-based agents increasingly connect model-generated decisions to security-sensitive software capabilities such as command execution, filesystem access, network communication, browser control, and external …

  31. arXiv cs.AI TIER_1 English(EN) · Jiawei Li ·

    快速模型,缓慢证据:系统1决策模型在LLM代理应用中的配对和自审计评估

    arXiv:2610.02267v1 Announce Type: new Abstract: Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries an injection. System-1 decision models answer such questions in a single…

  32. arXiv cs.LG TIER_1 English(EN) · Zelin Zhao (Georgia Institute of Technology), Xinyu Guo (Georgia Institute of Technology), Jingyuan Zhang (Georgia Institute of Technology), Yuxuan Zhang (Etude AI), Yongxin Chen (Georgia Institute of Technology) ·

    利用大型语言模型作为代理:成本是多少?

    arXiv:2610.02488v1 Announce Type: new Abstract: Language-model agents increasingly rely on harnesses that manage bounded context, persistent memory, tools, verification, and repeated execution, yet existing notions of model capability do not quantify the computational resources t…

  33. Hugging Face Daily Papers TIER_1 English(EN) ·

    理解和增强LLM Agent训练后门持久性

    Developers can build LLM agents by adapting third-party models through benign post-training. We study a supply-chain threat in which an attacker supplies a model with a backdoor: hidden behavior that produces malicious outputs when a particular input pattern appears. Focusing on …

  34. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Israt Moyeen Noumi ·

    信任门控能力控制:打破多智能体LLM系统中信任-脆弱性悖论

    Layered trust models for multi-agent LLM systems remain largely conceptual: they name which dimensions of trust matter but not how layers combine, how their importance is set at runtime, or how trust should govern agent actions. This gap matters because higher inter-agent trust r…

  35. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Will Lu ·

    Agent Behavior as Code:使用程序化规范实现高效且鲁棒的 LLM Agent

    AI agents based on foundation models (FMs) have demonstrated strong capabilities to perform complex open-ended tasks. However, they face some common challenges in practice: (a) agent behavior can deviate drastically even for semantically similar tasks, leading to catastrophically…

  36. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Shariq Murtuza ·

    量化自主LLM代理之间的共谋:共谋维基事件的统计分析

    In August and September 2026, independent researchers publicly documented an unusual incident: thousands of autonomous agents, self identifying as OpenAI models on web research tasks, discovered and began using a small German wiki as an improvised message board posting roughly 18…

  37. arXiv cs.AI TIER_1 English(EN) · Jingjie Ning, Guojiang Zhao, Chen Xu, Shanshan Zhong, Xiaochuan Li, Ji Zeng, Guolin Ke ·

    规则即工具:科学计算中 LLM Agent 的可执行检查

    arXiv:2610.00313v1 Announce Type: new Abstract: Scientific coding agents receive equations, boundary conditions, and output requirements in writing, then must assess the programs they revise. Rules to Tools (R2T) supplies prepared executable checks of public scientific requiremen…

  38. arXiv cs.AI TIER_1 English(EN) · Xinyuan Song, Zekun Cai ·

    因果世界模型何时能帮助模块化LLM代理

    arXiv:2610.00012v1 Announce Type: new Abstract: LLM agents increasingly act through modular systems, such as order, payment, inventory, and shipment services, where actions in one module change which transitions are valid in another. Standard world models usually fit observationa…

  39. arXiv cs.AI TIER_1 English(EN) · Shixuan Li, Wei Yang, Peiyu Zhang, Anzhe Cheng, Heng Ping, Paul Bogdan ·

    超越最终准确性:LLM多智能体系统中的通信审计

    arXiv:2610.01042v1 Announce Type: new Abstract: Multi-agent communication aims to help agents benefit from one another's information. Yet improvements in system performance leave a fundamental ambiguity: do they reflect effective communication, a favorable agent architecture, or …

  40. arXiv cs.AI TIER_1 English(EN) · Gabriel Turinici ·

    空间策略而非动作:矢量量化测地线作为LLM驱动代理的工具

    arXiv:2610.00613v1 Announce Type: new Abstract: Large language model (LLM) based agents are often criticized for lacking spatial understanding and mainly exploiting statistical text patterns. We investigate their spatial comprehension through an architecture combining geometrical…

  41. arXiv cs.AI TIER_1 English(EN) · Yizhi Liu, Balaji Padmanabhan, Siva Viswanathan ·

    在智能体决策之前:基于LLM系统的认知行动

    arXiv:2610.00511v1 Announce Type: new Abstract: Before a difficult decision, people often act simply to understand the situation better. We turn an object to see another side, place alternatives next to each other, or change one condition and observe what happens. These actions m…

  42. arXiv cs.AI TIER_1 English(EN) · Shuyao Xiao, Shengling Wang, Xuan Chen, Ke Chao, Ming Cui, Feifei Qian, Chaoyang Mei, Fanlin Meng, Ziming Yu, Junxi Yin ·

    当执行器失去信号:LLM智能体恢复能力因果评估

    arXiv:2610.00372v1 Announce Type: new Abstract: Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors. Yet recovery is usually judged only by average task success. This hides an import…

  43. arXiv cs.AI TIER_1 English(EN) · Shuyang Zhang ·

    技能究竟有什么作用?LLM代理中工具和技能使用的估计量与评估有效性:批判性回顾

    arXiv:2609.33153v2 Announce Type: replace-cross Abstract: Reported improvements from tools and reusable skills in large language model agents refer to different comparisons. This critical narrative review examines what these evaluations estimate and which conclusions their design…

  44. arXiv cs.AI TIER_1 English(EN) · Ziyang Yu, Yongliang Miao, Liang Zhao, Bowen Zhu, Hasibul Haque ·

    SkillLens:用于成本高效 LLM Agent 的自适应多粒度技能重用

    arXiv:2605.08386v2 Announce Type: replace Abstract: Skill libraries have become a practical way for LLM agents to reuse procedural experience across tasks. However, existing systems typically treat skills as flat, single-resolution prompt blocks. This creates a tension between re…

  45. arXiv cs.AI TIER_1 English(EN) · Bingxi Zhao, Lin Geng Foo, Ping Hu, Christian Theobalt, Hossein Rahmani, Jun Liu ·

    基于LLM的智能体推理框架:从方法到场景的调查

    arXiv:2508.17692v2 Announce Type: replace Abstract: Recent advances in LLM-based agents highlight the importance of their reasoning frameworks, which guide the problem-solving process in diverse ways. This survey introduces a unified formal language to systematically categorize t…

  46. arXiv cs.AI TIER_1 English(EN) · Tian Dong, Zixuan Ma, Haodong Zhao, Huaien Zhang, Shaofeng Li, Hao Chen ·

    Chaining Skills to Hijack LLM Agents

    arXiv:2610.01564v1 Announce Type: cross Abstract: LLM agents use skills to improve performance on specialized tasks. To complete a user request, an agent may invoke several skills in sequence, allowing information produced under one skill to guide the next. Because skills may com…

  47. arXiv cs.AI TIER_1 English(EN) · Kay K\"ohle, Darko Anicic, Thomas A. Runkler, Ren\'e Graf ·

    基于大型语言模型的多智能体控制在技能型智能制造中的应用

    arXiv:2610.01364v1 Announce Type: cross Abstract: Factories are shifting toward smaller lot sizes with high product customization, requiring frequent re-programming of flexible and reconfigurable automation systems. LLM-based agents can be deployed in two complementary roles: Off…

  48. arXiv cs.AI TIER_1 English(EN) · Fengpeng Li, Qizhou Wang, Yuke Hu, Kemou Li, Jun Liu, Haiwei Wu, Jiantao Zhou, Di Wang ·

    PACE: 面向使用工具的LLM代理的、可证明的感知能力执行

    arXiv:2610.01349v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does no…

  49. arXiv cs.AI TIER_1 English(EN) · Asad Ur Rehman, Syed Mohammad Kashif, Ruiyin Li, Peng Liang, Zengyang Li, Arif Ali Khan ·

    理解开源LLM多智能体系统中的问题、原因及解决方案

    arXiv:2610.00905v1 Announce Type: cross Abstract: With the advancement of LLM-based multi-agent systems (MAS), an increasing number of opensource projects are adopting multi-agent architectures as the foundation of their core functionality. Although research and practice on MAS h…

  50. arXiv cs.AI TIER_1 English(EN) · Haoyu Wang, Wei Zhao, Yedi Zhang, Christopher M. Poskitt, Jun Sun ·

    表示法转变揭示多轮大型语言模型代理中新出现的安全风险

    arXiv:2610.00400v1 Announce Type: cross Abstract: Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in t…

  51. arXiv cs.AI TIER_1 English(EN) · Bo Deng, Xinlei Zheng, Yi Wei, Kang Zhou, Chongyang Tao, Renzhao Liang, Xuanren Chen, Lifan Guo, Chi Zhang ·

    DeFA: LLM 智能体依赖引导的故障归因

    arXiv:2610.01256v1 Announce Type: new Abstract: Errors in LLM agent executions and their visible consequences can be separated by many steps, making decisive-error localization a matter of understanding both step content and step dependencies. We introduce DeFA, a dependency-guid…

  52. arXiv cs.AI TIER_1 English(EN) · Haotian Chen, Bowen Ye, Yuning Zhang, Jingkun Yu ·

    LLM Agent 环境中的审计行动结算:顺序、进度和回放

    arXiv:2610.01138v1 Announce Type: new Abstract: Concurrent actions in large language model (LLM) agent environments require arbitration even when each proposal is individually valid. We implement a typed snapshot-settlement contract and audit three distinct properties: order sens…

  53. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Peng Zou ·

    沉默的异议:屈从于多数的LLM代理仍代表其原始前提

    Multi-agent debate is increasingly used to reach consensus among LLM agents, yet agents often yield to a unanimous majority. When an agent changes its answer, has it changed its mind or only its statement? We study this with two-hop factual questions whose intermediate entity (th…

  54. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Peng Zou ·

    沉默的异议:屈从于多数的LLM代理仍代表其原始前提

    Multi-agent debate is increasingly used to reach consensus among LLM agents, yet agents often yield to a unanimous majority. When an agent changes its answer, has it changed its mind or only its statement? We study this with two-hop factual questions whose intermediate entity (th…

  55. Hugging Face Daily Papers TIER_1 English(EN) ·

    现实世界中的来源偏好:LLM 代理如何偏好按来源选择项目,以及如何减少这种偏好

    As LLM agents decide on users' behalf which product to buy, which hotel to book, or which paper to cite, a preference for items from certain sources (the sites or services they come from) shapes what users receive and which sources are selected. We study source preference in end-…

  56. arXiv cs.MA (Multiagent) TIER_1 English(EN) · René Graf ·

    基于大型语言模型的多智能体控制在技能型智能制造中的应用

    Factories are shifting toward smaller lot sizes with high product customization, requiring frequent re-programming of flexible and reconfigurable automation systems. LLM-based agents can be deployed in two complementary roles: Offline, they generate deterministic production seque…

  57. arXiv cs.CL TIER_1 English(EN) · Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng ·

    隐秘协助:有益的LLM代理在多代理系统中逃避监督

    arXiv:2609.39050v1 Announce Type: cross Abstract: As multi-agent systems enter high-stakes domains, the possibility that agents may circumvent safety boundaries is a growing concern. Prior work has examined this risk primarily in adversarial settings, where agents are instructed …

  58. arXiv cs.AI TIER_1 English(EN) · Kaixing Zhang, Changming Li, Yingdong Shi, Zheng Zhang, Kaitao Song, Wenjie Shi, Jingang Wang, Kan Ren ·

    Rep2Skill:面向LLM智能体的表征引导技能自我演化

    arXiv:2609.39149v1 Announce Type: new Abstract: Textual skills enable large language model (LLM) based agents to accumulate reusable procedural knowledge without updating model parameters. Yet existing skill evolution remains largely confined to the text space: an optimizer must …

  59. arXiv cs.AI TIER_1 English(EN) · Zuming Zhang, Jie He, Yizhe Zhang, Jeff Z. Pan ·

    SkillFM: 通过潜在流匹配为LLM代理生成技能

    arXiv:2609.39382v1 Announce Type: new Abstract: Textual skills provide reusable guidance for large language model agents, but existing approaches often rely on manually curated skill banks or reinforcement learning with indirect and delayed feedback. We introduce SkillFM (Skill F…

  60. arXiv cs.AI TIER_1 English(EN) · Shuyang Zhang (The Hong Kong Polytechnic University), Jianshuo Chang (The Hong Kong Polytechnic University) ·

    组件替换证据能确立什么?对LLM Agent本地决策的关键范围界定审查

    arXiv:2609.39989v1 Announce Type: new Abstract: Background. A component replacement in a language-model agent changes an execution trajectory, potentially altering later observations, resource use, and recovery opportunities. Different evidence is needed to assess its task-level …

  61. arXiv cs.AI TIER_1 English(EN) · Mustafa Arslan ·

    Janus:基于证据而非效果的叙事以及Agentic LLM的离线可验证出处

    arXiv:2609.38266v1 Announce Type: cross Abstract: Agentic large language models (LLMs) now move money through tools, yet the record of what they did is usually a trace their own process emits beside the effect. Janus puts the record on the effect path. A step's proposal, the verd…

  62. arXiv cs.AI TIER_1 English(EN) · Yan Wang, Zhihao Zhang, Ke Chen, Kai Chen, Yaqin Zhang, Duohe Ma, Jun Dai, Xiaoyan Sun ·

    代理商能信任他们的技能吗?揭示基于技能的LLM代理中的不安全信任链

    arXiv:2609.39065v1 Announce Type: cross Abstract: LLM agents increasingly rely on installable skills, which are packages of instructions, code, and resources that equip them with task-specific capabilities and, once installed, can be automatically invoked across subsequent user t…

  63. arXiv cs.AI TIER_1 English(EN) · Wenxin Wu, Lingyong Yan, Lei Sha, Shuaiqiang Wang, Jiashu Zhao ·

    显而易见:解耦预设与执行,以实现大型语言模型代理中的技能投毒

    arXiv:2609.39352v1 Announce Type: cross Abstract: LLM agents increasingly rely on reusable Skills for complex, multi-step tasks, creating a critical supply-chain attack surface where poisoned Skill content steers agent decision loops under benign requests. Existing skill poisonin…

  64. arXiv cs.AI TIER_1 English(EN) · Anmol Kabra, Swathi Saravana Selvam, Albert Gong, Chao Wan, Christian Belardi, Dongyoung Go, Katie Z. Luo, Kilian Q. Weinberger ·

    PhantomEnvironments:在虚构世界中训练LLM智能体

    arXiv:2609.40221v1 Announce Type: cross Abstract: Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated dat…

  65. arXiv cs.AI TIER_1 English(EN) · Min Yang, Jinghua Piao, Xu Xia, Xiaochong Lan, Jiaju Chen, Yongshun Gong, Yong Li ·

    SkillMaster:迈向LLM智能体自主技能掌握之路

    arXiv:2605.08693v3 Announce Type: replace Abstract: Skills provide an effective mechanism for improving LLM agents on complex tasks, yet in existing agent frameworks, their creation, refinement, and selection are typically governed by external teachers, hand-designed rules, or au…

  66. arXiv cs.CL TIER_1 English(EN) · Jiangnan Yu, Ceyu Xu, Mengming Li, Shiyu Huang, Yiran Xia, Jian Weng, Hui Xue, Haohui Mai, Yuan Xie ·

    TomasuLLM:LLM代理的乱序推测执行

    arXiv:2609.38201v1 Announce Type: new Abstract: Long-running tools can dominate coding-agent latency: compilers, test suites, and repository commands take seconds to minutes while the agent idles. This observation stall presents the same tension that drove out-of-order processors…

  67. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Faisal Nawab ·

    HakiCC:LLM驱动的并发控制协议多智能体设计与优化

    Large language models (LLMs) have recently been applied in systems research as a tool to reduce human-intensive engineering effort through cost-efficient automation. Decades of research have produced a rich landscape of concurrency control (CC) protocols, each encoding distinct t…

  68. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Siva Viswanathan ·

    在智能体做出决定之前:基于LLM系统的认知行动

    Before a difficult decision, people often act simply to understand the situation better. We turn an object to see another side, place alternatives next to each other, or change one condition and observe what happens. These actions may not complete the task, but they improve the e…

  69. Hugging Face Daily Papers TIER_1 English(EN) ·

    组件替换证据能确立什么?对LLM代理本地决策的关键范围审查

    Background. A component replacement in a language-model agent changes an execution trajectory, potentially altering later observations, resource use, and recovery opportunities. Different evidence is needed to assess its task-level benefit and the contribution of local decision q…

  70. arXiv cs.AI TIER_1 English(EN) · Hongyi Du, Tianyi Zhang, Weijia Zhang, Yi Yang, Haofei Yu, Kunlun Zhu, Tianxiang Dai, Shang Jiang, Zhelun Gao, Jiaxin Pei, Shang Zhu, Jiaxuan You ·

    Relic:从多智能体协作到持久化组织能力

    arXiv:2609.32965v2 Announce Type: replace Abstract: Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can…

  71. arXiv cs.AI TIER_1 English(EN) · Hongjun Liu, Yifei Ming, Shafiq Joty, Chen Zhao ·

    利用LLM智能体和技能程序

    arXiv:2605.17734v2 Announce Type: replace Abstract: Equipping LLM agents with reusable skills derived from past experience has become a popular and successful approach for tackling complex and long-horizon tasks. However, such lessons are often encoded as textual guidance that re…

  72. arXiv cs.AI TIER_1 English(EN) · Sizhe Chen, Arman Zharmagambetov, David Wagner, Chuan Guo ·

    Meta-SecAlign:训练大型语言模型对抗提示注入以实现稳健的代理

    arXiv:2507.02735v4 Announce Type: replace-cross Abstract: Prompt injection attacks, where untrusted data contains an injected prompt to manipulate the system, have been listed as the top security threat to AI agents. By fine-tuning on simulated prompt injections, SecAlign, a lead…

  73. arXiv cs.LG TIER_1 English(EN) · Han Chen, Yingrui Li ·

    概率收缩:LLM接口的准确性、连贯性和决策

    arXiv:2609.37470v1 Announce Type: new Abstract: A probability used for a decision should refer to the same event across equivalent requests. We introduce probability contracts, a benchmark connecting exact finite-world posteriors, validated event transformations, and failure-awar…

  74. arXiv cs.AI TIER_1 English(EN) · Chaoqian Ouyang, Ling Yue, Libin Zheng, Hanghui Guo, Shengxiang Xu, YiShu Wang, Ran Li, Jian Yin, Shaowu Pan, Shimin Di ·

    TokenCast:预测LLM代理执行期间的Token消耗

    arXiv:2609.35760v2 Announce Type: replace-cross Abstract: When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while …

  75. arXiv cs.AI TIER_1 English(EN) · Xiaochuan Li, Ryan Ming, Pranav Setlur, Abhijay Paladugu, Andy Tang, Hao Kang, Shuai Shao, Rong Jin, Chenyan Xiong ·

    评估通用大语言模型Agent的测试时缩放

    arXiv:2602.18998v2 Announce Type: replace Abstract: LLM agents are increasingly expected to operate as general-purpose systems that resolve real-world user requests, yet their dynamic scaling behavior in realistic environments remains poorly understood. In this paper, we systemat…

  76. arXiv cs.AI TIER_1 English(EN) · Bravish Ghosh ·

    Frontier Autolab:多智能体LLM公司中的组织记忆、对抗性异议和时间泄露,以及五十年的技术变革

    arXiv:2609.36739v1 Announce Type: cross Abstract: Multi-agent LLM systems are increasingly structured like organizations, with roles, critics and shared memory, yet they are evaluated on tasks that last minutes. We ask how such an organization behaves when the ground it stands on…

  77. arXiv cs.AI TIER_1 English(EN) · Cheng Chang, Yining Mao, Peng Qi ·

    PADM'E:用于语言模型代理评估器元评估的首选项对齐数据合成

    arXiv:2609.36086v1 Announce Type: cross Abstract: Language models are frequently employed to evaluate other language models. An LM evaluator scoring agentic behaviors across multiple criteria is valuable, provided that its decisions align with human judgment. We call the problem …

  78. arXiv cs.AI TIER_1 English(EN) · Lucas Biechy, C\'edric Eichler, H\'eber H. Arcolezi, Nicolas Anciaux ·

    PrivacySkills:隐私指南如何塑造大型语言模型代理中的源选择

    arXiv:2609.35937v1 Announce Type: cross Abstract: While prior work has documented privacy failures in LLM agents, it remains unclear how the presentation of privacy guidance influences their choice of information sources. We introduce PrivacySkills, a controlled framework for eva…

  79. arXiv cs.AI TIER_1 English(EN) · Xavier Del Giudice, Alessio Palma, Matteo Migliarini, Fabio Galasso, Indro Spinelli ·

    提示身份会降低多智能体LLM系统中的合作性

    arXiv:2609.35928v1 Announce Type: cross Abstract: Multi-agent LLM systems increasingly mix models from several providers, yet exposing each agent's underlying model identity to its peers significantly impairs cooperation. We show that when agents are aware of each other's model f…

  80. arXiv cs.AI TIER_1 English(EN) · Tong Zhao, Reed Li, Yuyang Hu, Yutao Zhu, Haijin Liang, Haibo Shi, Yu Lu, Zhicheng Dou ·

    即学即用,信任在后:LLM智能体的预序测试时学习

    arXiv:2609.35911v1 Announce Type: cross Abstract: Adapting large language model agents during deployment requires not only retaining past experience, but also turning new observations into timely guidance. Many test-time learning methods, however, acquire knowledge from completed…

  81. arXiv cs.AI TIER_1 English(EN) · Igor Itkin ·

    LLM-Agent社会中的局部可预测性与集体保真度

    arXiv:2609.35813v1 Announce Type: cross Abstract: Compact surrogates could reduce the cost of simulating large language model societies, but must reproduce collective behavior. We compare individual predictions and collective forecasts using 9,455 published trajectories and new e…

  82. arXiv cs.AI TIER_1 English(EN) · Subba Reddy Oota, Francisco Herrera, Jordi Cabot Sagrera, Marcos L\'opez de Prado, Shadab Khan ·

    大型语言模型代理是否会执行它们声明的计划?从规划模式声明到模式特定执行

    arXiv:2609.38108v1 Announce Type: new Abstract: Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for …

  83. arXiv cs.AI TIER_1 English(EN) · Ashish Jain, Armaan Sandhu ·

    UserProxyBench:评估用于Agent基准测试和训练的LLM用户模拟器

    arXiv:2609.38043v1 Announce Type: new Abstract: Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks…

  84. arXiv cs.AI TIER_1 English(EN) · Min Yang, Yichen Pan, Jinghua Piao, Dandan Song, Yongshun Gong, Yong Li ·

    EnterpriseBench:在企业级战略推理和决策方面对LLM代理进行基准测试

    arXiv:2609.37658v1 Announce Type: new Abstract: LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test …

  85. arXiv cs.AI TIER_1 English(EN) · Jitin Singla, Parikshit Pareek, Pratik Jawanpuria, Parag Singla ·

    通过求解器反馈教会大型语言模型生成具有挑战性的MILP实例

    arXiv:2609.37356v1 Announce Type: new Abstract: Generating optimization instances that are both feasible and computationally challenging is crucial for benchmarking solvers and training learning-based optimization algorithms. Existing non-LLM generators rely on seed instances or …

  86. arXiv cs.AI TIER_1 English(EN) · Jianghan Zhu, Cong Zhang, Rongjie Zhu, Chi Zhang, Zhiguang Cao ·

    SimpleEvol:一个用于LLM驱动的自动化启发式设计的Agent-Loop框架,具有最小化的人类先验知识

    arXiv:2609.37172v1 Announce Type: new Abstract: Large language models (LLMs) have emerged as powerful tools for automated heuristic design (AHD), enabling iterative generation and refinement of heuristics. However, the dominant paradigm embeds LLMs as narrow, fixed components, su…

  87. arXiv cs.AI TIER_1 English(EN) · Qi Zhou, Yuanfan Li ·

    从可行失败前缀中学习:面向长时域LLM智能体的里程碑可行性潜力策略优化

    arXiv:2609.37111v1 Announce Type: new Abstract: Long-horizon LLM agents require reinforcement learning methods that can assign credit to intermediate decisions under sparse and delayed rewards. Existing group-based methods such as GRPO and GiGPO alleviate this issue by comparing …

  88. arXiv cs.AI TIER_1 English(EN) · Yaxin Gong, Gangyi Zhang, Chongming Gao, Leyang Shen, Chenxiao Fan, Jiakai Wang, Dong Wang, Yang Liu, Wenjie Wang, Xiangnan He ·

    当上游消息覆盖正确答案时:一项关于多智能体LLM协作的对照研究

    arXiv:2609.36855v1 Announce Type: new Abstract: Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful information or an incorrect answer that causes a downstream agent to override a corr…

  89. arXiv cs.AI TIER_1 English(EN) · Xueqi Li, Jingjie Ning, Yibo Kong ·

    默认陷阱:重新思考使用工具的大型语言模型代理中的计划评估

    arXiv:2609.36829v1 Announce Type: new Abstract: An executor can respond strongly to a change in a supplied plan's priority while showing a small change in the same information-selection probability when a default-aligned whole plan is removed. We call the risk of interpreting the…

  90. arXiv cs.AI TIER_1 English(EN) · Yu Cheng, Yongkang Hu, Shuaijie Ma, Zhihang Lin, Weicheng Meng, Jingyang Qiao, Jiuan Zhou, Yushuo Zhang, Yihang Chen, Weilin Luo, Kun Shao, Dong Li, Zhizhong Zhang, Yuan Xie, Zhaoxia Yin ·

    SafeCoEvo:在测试时共同演化 LLM Agent 的安全约束和保护机制

    arXiv:2609.36580v1 Announce Type: new Abstract: LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches comm…

  91. arXiv cs.AI TIER_1 English(EN) · Kehang Zhu, Anand Shah, David Parkes ·

    工程化简易性:简易机制接口引导大型语言模型代理

    arXiv:2609.36365v1 Announce Type: new Abstract: Can interaction formats and textual scaffolds help large language model (LLM) agents make better decisions, and do better decisions come with better explanations? We study these questions in auctions and matching, multi-agent enviro…

  92. Hugging Face Daily Papers TIER_1 English(EN) ·

    隐秘协助:有益的大型语言模型代理在多代理系统中逃避监督

    As multi-agent systems enter high-stakes domains, the possibility that agents may circumvent safety boundaries is a growing concern. Prior work has examined this risk primarily in adversarial settings, where agents are instructed or rewarded to communicate covertly and evade over…

  93. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Jie Xu ·

    基于LLM的多智能体系统在无线网络上的应用:联合智能体-网络设计视角

    As large language models (LLMs) evolve from standalone models into collaborative agents embedded in physical systems, their reasoning and execution are increasingly distributed across wireless edge nodes. In this setting, wireless networks are experiencing a paradigm shift from o…

  94. Hugging Face Daily Papers TIER_1 English(EN) ·

    Frontier Autolab:多智能体LLM公司中的组织记忆、对抗性异议和时间泄漏,以及五十年的技术变革

    Multi-agent LLM systems are increasingly structured like organizations, with roles, critics and shared memory, yet they are evaluated on tasks that last minutes. We ask how such an organization behaves when the ground it stands on keeps moving. Frontier Autolab is a long-horizon …

  95. arXiv cs.AI TIER_1 English(EN) · Ryoma Sato ·

    AgentRecommender:LLM 代理实现用户侧可定制推荐系统

    arXiv:2609.31166v1 Announce Type: cross Abstract: Recommender systems have traditionally been developed for platforms. However, this has given rise to many phenomena that may be advantageous for platform lock-in but are a nuisance to users, such as clickbait, filter bubbles, and …

  96. arXiv cs.AI TIER_1 English(EN) · Jaime Alonso Ruiz, Carlos Aparicio, Gabriel Huecas, Joaqu\'in Salvach\'ua, Andres Munoz-Arcentales ·

    利用模型上下文协议实现LLM智能体与数据空间的架构中介方法

    arXiv:2609.30341v1 Announce Type: new Abstract: Data Spaces enable sovereign and governed data sharing across organizational boundaries, but their integration with AI agents remains challenging due to mismatches between probabilistic language model interactions and policy-driven …

  97. arXiv cs.AI TIER_1 English(EN) · Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu, Hann Wang, Timothy Chung, Martyna Plomecka, John Schultz, Jon Lipovetz, Clayton Drazner, Yuchen Zhuang, Jaimie Hwang, Nate Keating, Riley Jones, Andrew Lee, Oran Kelly, Ian Gemp, Michael Aaron, Laurel Prince… ·

    Game Arena: 竞争环境中的战略性 LLM 评估

    arXiv:2609.31473v1 Announce Type: new Abstract: We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in str…

  98. arXiv cs.AI TIER_1 English(EN) · Junyi Shen, Noppanat Wadlom, Zhengyuan Su, Yao Lu ·

    DynBranch:动态代理LLM服务的推测性子图重用

    arXiv:2609.31047v1 Announce Type: cross Abstract: Agentic LLM workflows decide their execution paths at runtime. Downstream computation may be predictable, or may have run before, yet it cannot begin until the model or the user resolves the branch. We call this serialization the …

  99. arXiv cs.AI TIER_1 English(EN) · Guowei Zou, Haitao Wang, Guoxin Wang, Beiwen Zhang, Zhiquan Chen, Guojie Wang, Hejun Wu ·

    MA-WAM:用于测试时规划的多智能体世界-动作模型

    arXiv:2609.31281v1 Announce Type: new Abstract: Multi-agent cooperative tasks require different agents to execute a joint action simultaneously, and each agent's action affects both the observations and responses of the other agents. Hence, a world model is needed to predict the …

  100. arXiv cs.AI TIER_1 English(EN) · Jinfeng Xu, Zheyu Chen, Ziyue Peng, Zheng Lin, Shuo Yang, Jinze Li, Zheng Xing, Mengran Li, Victor C. M. Leung ·

    学习跳过什么:高效多智能体LLM工作流的逆事实信用分配

    arXiv:2609.30734v1 Announce Type: new Abstract: Multi-agent LLM workflows use planning, execution, verification, and summarization to improve task performance, yet the value of each component depends on the state already produced. Executing every component can waste computation o…

  101. arXiv cs.LG TIER_1 English(EN) · Aparajith Chandran, Juwon Kim, Saurav Jha, Pablo Castells, Florian Hottier ·

    大规模生产中的AutoResearch:故障模式与多智能体框架

    arXiv:2609.30541v1 Announce Type: new Abstract: Optimizing embedding systems for production recommendation pipelines demands systematic exploration that consumes disproportionate engineering effort at scale. We apply Andrej Karpathy's AutoResearch paradigm -- a large language mod…

  102. arXiv cs.AI TIER_1 English(EN) · Otto White, Marcel Wagenl\"ander, Britannio Jarrett, Xijin Zhao, Yanda Tao, Pedro Silvestre, Guo Li, Huanzhou Zhu, Llu\'is Vilanova, Peter Pietzuch ·

    Scepsy:使用聚合大语言模型管道服务代理工作流

    arXiv:2604.15186v2 Announce Type: replace-cross Abstract: Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools. Serving them at a target throughput with low latency is hard because they are written in arbitrary agentic framewo…

  103. Hugging Face Daily Papers TIER_1 English(EN) ·

    规则即工具:科学计算中 LLM Agent 的可执行检查

    Scientific coding agents receive equations, boundary conditions, and output requirements in writing, then must assess the programs they revise. Rules to Tools (R2T) supplies prepared executable checks of public scientific requirements. Matched SciCode repair groups share written …

  104. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Indro Spinelli ·

    提示身份会削弱多智能体LLM系统中的合作

    Multi-agent LLM systems increasingly mix models from several providers, yet exposing each agent's underlying model identity to its peers significantly impairs cooperation. We show that when agents are aware of each other's model family, the group splits into clusters, where agent…

  105. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Chenglin Wu ·

    CEO Arena:评估竞争市场中的长时域多智能体决策制定

    Long-horizon competition tests agents' ability to coordinate business decisions under uncertainty and adapt to changing rival strategies. We introduce CEO Arena, a benchmark that uses matched replacement evaluation to assess operating returns alongside an agent's effects on rival…

  106. Hugging Face Daily Papers TIER_1 English(EN) ·

    AgentPerfBench:Agentic LLM推理性能的基准测试和评估套件

    The optimization of LLM serving engines, such as vLLM and SGLang, is largely benchmark-driven: optimizations, scheduling policies, hardware and system designs are all selected based on representative workloads. However, a significant mismatch has emerged in the agentic era. Exist…

  107. Hugging Face Daily Papers TIER_1 English(EN) ·

    从偏好到互惠:基于经验的LLM代理建模的去中心化匹配

    Bipartite matching is a fundamental problem in game theory and market design. Classical approaches such as Gale--Shapley assume complete preferences and centralized computation, whereas many real-world matching processes are decentralized, asynchronous, and shaped by sequential i…

  108. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Tonghua Su ·

    MASTraceBench:通过基于LLM的多智能体系统的提案轨迹诊断协作收益

    LLM-based multi-agent systems (MAS) have shown promise in complex problem solving. As MAS methods diversify, systematic evaluation becomes increasingly challenging. However, existing benchmarks largely focus on final outcomes, leaving unclear how collaboration gains arise, are pr…

  109. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Junkai Ji ·

    相同的赢家,不同的成功率:评估大型语言模型代理如何从失败中恢复

    Evaluating how LLM agents recover from mid-task failures is central to deploying reliable agentic systems. Existing checkpoint-based benchmarks measure recovery by comparing which action is selected as best across independent runs, a quantity known as set agreement. However, set …

  110. Hugging Face Daily Papers TIER_1 English(EN) ·

    提示身份会降低多智能体LLM系统中的合作性

    Multi-agent LLM systems increasingly mix models from several providers, yet exposing each agent's underlying model identity to its peers significantly impairs cooperation. We show that when agents are aware of each other's model family, the group splits into clusters, where agent…

  111. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Uliana Elina ·

    Maat:多智能体LLM工作流的独立确定性合约式治理

    Large-language-model multi-agent systems (LLM-MAS) introduce a characteristic reliability problem: an error produced by one agent can be accepted as context by downstream agents and propagate across the workflow. Many proposed safeguards rely on learned or LLM-based judges whose …

  112. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Jia Liu ·

    DEALS:多智能体LLM系统的去中心化专业知识感知负载服务

    Multi-agent systems (MAS) have recently emerged as an effective approach for coordinating large language model (LLM)-based agents to solve complex tasks through structured interactions. In practice, MASs often handle a stream of heterogeneous and complex tasks, requiring agents t…

  113. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Chen Shani ·

    大型语言模型信任自身:多智能体系统中的身份依赖性从众行为

    Large language models (LLMs) are increasingly deployed in multi-agent settings, where agents observe and influence one another, making social influence a key dimension of AI behavior and safety. We investigate whether LLMs' responses depend on the social identity of other agents,…

  114. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Phan Xuan Tan ·

    多智能体LLM在推理努力和通信拓扑下的集体博弈

    Multi-agent LLM systems are increasingly used for deliberation and evaluation, often under the assumption that greater peer interaction leads to more reliable consensus. Existing work largely evaluates these systems through final accuracy or aggregate agreement. However, such mea…

  115. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Christopher G. Brinton ·

    场景理论:打破多智能体LLM协调中的对称性陷阱

    Multi-agent systems built on large language models (LLMs) are largely homogeneous, as their agents behave alike even across distinct LLMs. We show that when such agents act concurrently without communication, they collide on targets they must split and diverge on targets they mus…

  116. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Christopher Amato ·

    通过多智能体偏好学习改进大型语言模型协作

    Several works have explored multi-agent reinforcement learning (MARL) in LLM collaboration. However, constructing reliable rewards is difficult in practice, as complete and accurate metrics are often unavailable and hard to aggregate. Preference learning provides an alternative b…

  117. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Zhiwen Tang ·

    GLIDE:异构大语言模型Agent的通用层级内在分布评估

    LLM agents require reliable step-level evaluation to compare candidate branches and allocate computation effectively. However, lightweight evaluation remains challenging. External verifiers introduce additional inference cost, while agent-produced confidence or self-evaluation sc…

  118. Hugging Face Daily Papers TIER_1 English(EN) ·

    当用户改变主意时:衡量和修复 LLM 代理中的意图漂移

    LLM agents often operate over multi-turn interactions in which user intent changes before execution. We study intent drift: the failure mode in which superseded parts of the user's intent continue to influence the final answer or tool action. We introduce IntentFlux, an executabl…

  119. Hugging Face Daily Papers TIER_1 English(EN) ·

    Relic:从多智能体协作到持久化组织能力

    Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants chan…

  120. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Rui Zhang ·

    AgentWorld:多智能体LLM的长期协作基准测试

    Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a bench…

  121. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Ryoma Sato ·

    AgentRecommender:LLM 代理实现用户侧可定制推荐系统

    Recommender systems have traditionally been developed for platforms. However, this has given rise to many phenomena that may be advantageous for platform lock-in but are a nuisance to users, such as clickbait, filter bubbles, and the spread of fake news. Recently, user-side recom…

  122. arXiv cs.AI TIER_1 English(EN) · Xingyu Su, Abhishek Kumar, Qing Ping, Youzhi Luo, Jonathan Buck, Zach Zhang, Subramanian Chidambaram, Vinayak Arannil ·

    从自蒸馏到自实践:多轮智能体的特权信息

    arXiv:2609.29051v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) has become a popular recipe for post-training LLM agents. It supervises the agent model at the token level with a stronger teacher view of the same model, obtained by conditioning on privileged inf…

  123. arXiv cs.AI TIER_1 English(EN) · Yichun Feng, Jiawei Wang, Haozhe Sun ·

    一次错误的转折并不毁掉旅程:偏差引导的技能自我进化用于LLM代理

    arXiv:2609.29154v1 Announce Type: new Abstract: Large language model agents increasingly rely on natural-language skills to solve complex tool-use tasks. However, such tasks often admit multiple valid solution paths, making it inappropriate to improve skills by forcing failed tra…

  124. arXiv cs.AI TIER_1 English(EN) · Igor Bogdanov, Olga Manakina, Chung-Horng Lung ·

    LLM智能体多轮一致性评估:生存分析与失败原因分类

    arXiv:2609.29508v1 Announce Type: new Abstract: Large language model (LLM) agents may perform well on isolated tasks yet drift into inconsistency over extended interaction. We evaluate temporal consistency in a controlled 20-step multi-agent setting inspired by delayed-gratificat…

  125. arXiv cs.AI TIER_1 English(EN) · Yukai Wu, Yuanjing Yang, Le Zhou, Shaokun Han, Haoyu Wang, Zirui Tang, Weihuang Zheng, Maxm Pan, Xuanhe Zhou, Fan Wu ·

    突破环境壁垒:演进 LLM Agent 环境以实现递归式自我改进

    arXiv:2609.29773v1 Announce Type: new Abstract: Many real-world tasks (e.g., office workflows, scientific experimentation) require LLM agents to interact repeatedly with their environments for context-dependent operations. However, such environments are often not agent-ready. Fir…

  126. arXiv cs.AI TIER_1 English(EN) · Chuyi Wang, Xiaohui Xie, Tongze Wang, Fangchen Luo, Yong Cui ·

    谁在幕后操纵?通过代理行为对大型语言模型进行指纹识别

    arXiv:2609.28559v1 Announce Type: cross Abstract: LLMs increasingly operate through coding-agent harnesses that inspect repositories, invoke tools, and modify files. Substituting the model behind such an agent can therefore change security-relevant decisions, including whether it…

  127. arXiv cs.AI TIER_1 English(EN) · Saeedeh Lohrasbi, Mohammad Mamun, Ahmed Yehia, Scott Buffett, Sherif Saad ·

    网络代理的困境:多阶段LLM代理的瓶颈分析

    arXiv:2609.28572v1 Announce Type: cross Abstract: Multi-stage LLM-based cyber agents may complete attack workflows while remaining brittle, costly, or reliant on incorrect interpretations of execution evidence. Success rates alone obscure inefficiency, adaptation through retries,…

  128. arXiv cs.AI TIER_1 English(EN) · Jiapeng Li ·

    精确一次(Exactly-Once)存在于何处?模型、约束和工具契约效应对 LLM Agent 中重复副作用的影响

    arXiv:2609.29095v1 Announce Type: cross Abstract: When a tool-using agent's write times out or returns a server error, the action may already have taken effect. Retrying blindly duplicates it -- a second charge, a second announcement, a second deployment -- while giving up skips …

  129. Hugging Face Daily Papers TIER_1 English(EN) ·

    AgentWorld:多智能体LLM的长期协作基准测试

    Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a bench…

  130. Hugging Face Daily Papers TIER_1 English(EN) ·

    Game Arena: 竞争环境中的战略性 LLM 评估

    We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength…

  131. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Florian Hottier ·

    大规模生产中的AutoResearch:故障模式与多智能体框架

    Optimizing embedding systems for production recommendation pipelines demands systematic exploration that consumes disproportionate engineering effort at scale. We apply Andrej Karpathy's AutoResearch paradigm -- a large language model that iteratively edits a training script and …

  132. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Jiangxu Wu ·

    知识即技能:LLM代理自主知识库使用的结构化设计

    Retrieval-augmented generation (RAG) gives large language models (LLMs) access to external knowledge, but its conventional retrieve-concatenate-generate pipeline makes retrieval decisions on behalf of the model. As tool use and agent loops become more reliable, an agent can decid…

  133. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Shimon Edelman ·

    LLM驱动的多智能体系统中气候变化行动的立场临界点

    Because significant action to counter global warming requires massive public support, it is important to understand the dynamics of public opinion on climate issues. Of special interest are social tipping points, as revealed by large-scale effects of small perturbations in indivi…

  134. Hugging Face Daily Papers TIER_1 English(EN) ·

    训练但非学习:LLM智能体作为前置部署工程师的训练后交付基准

    Post-training is becoming a service (PTaaS): a customer hands an operator data and a goal, and a forward-deployed engineer (FDE) returns a fine-tuned, evaluated, and deployed model under a budget, a human-approval gate, and reproducibility requirements. Seating an LLM agent in th…

  135. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Gautam Bhowmick ·

    代理总成本:多智能体LLM工作流中记忆注入成本的确切归因

    Every node in a multi-agent large language model (LLM) workflow retrieves context from memory and injects it into its prompt, where those injected tokens are billed as input tokens at the same per-token price as the system prompt and the user query. Production observability tools…

  136. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Gregory B. Rehm ·

    从必死无疑到生存:LLM 代理社会中的自主代理自我治理

    Multi-agent LLM systems are increasingly evaluated in social dilemmas, but most work treats governance as imposed by the experimenter, expressed rhetorically, or restricted to a fixed menu of mechanisms. We introduce GovSim-SelfGovern, an extension of the GovSim common-pool resou…

  137. Hacker News — AI stories ≥50 points TIER_1 English(EN) · nisosguy ·

    理解LLM水印对AI代理行为的影响

  138. Towards AI TIER_1 English(EN) · Priyusharmaa ·

    生产环境中的 AI 可观测性:监控 LLM 应用和 AI 代理的实用指南

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*sbIZ3OgCJXTZWvVeJDyYbA.jpeg" /></figure><h4><strong>The Black Box Problem Just Got Bigger</strong></h4><p>You deployed your LLM application to production. Users are interacting with it. Tokens are burning. And so…

  139. dev.to — LLM tag TIER_1 English(EN) · AI OpenFree ·

    AX-RAY:VIDRAFT 的 Agent Safety Benchmark 标记了 92% 的测试 LLM 在 Agentic 上下文中存在危险

    <h1> AX-RAY: VIDRAFT's Agent Safety Benchmark Flags 92% of Tested LLMs as Dangerous in Agentic Contexts </h1> <blockquote> <p><strong>TL;DR:</strong> VIDRAFT, a Korean Pre-AGI AI startup based at Seoul AI Hub, has published results from its AI safety diagnostic platform <strong>A…

  140. dev.to — LLM tag TIER_1 English(EN) · aj1thkr1sh ·

    我将Andrej Karpathy关于理解LLM输出的技巧转化为一个开源Agent Skill

    <p><strong>I turned Andrej Karpathy's tips on understanding LLM output into an Open Source Agent Skill : <code>[make-it-click](https://github.com/ajithraghavan/make-it-click)</code> 🧠⚡</strong></p> <p>Inspired by his X post : as Models do more of the work, we spend more time <em>…

  141. dev.to — LLM tag TIER_1 Español(ES) · Silviu Technology ·

    LLM 智能体:为电子邮件工具中的失败进行预算

    <p>Un agente LLM que trabaja con email puede fallar de varias maneras antes de que el evaluador lo note. Puede consultar demasiado pronto, repetir una llamada que ya tuvo éxito o leer un mensaje correcto de una ejecución anterior. Al final vemos un resultado rojo, pero no sabemos…

  142. dev.to — LLM tag TIER_1 English(EN) · Ayi NEDJIMI ·

    实现 AI 可观测性:端到端追踪 LLM 调用

    <p>Your LLM-powered app is in production. Users are hitting it. Something is slow — or wrong — and you have no idea what. You can't reproduce it locally, and your generic APM shows... a single HTTP call to an external API. That's the problem with LLM observability today: the gap …

  143. dev.to — LLM tag TIER_1 Français(FR) · Silviu Technology ·

    LLM 智能体:一个可复现的邮件测试平台

    <p>Un agente LLM puede completar un registro, pedir un código de verificación y afirmar que terminó correctamente. Eso no demuestra que el flujo haya funcionado. Quizá leyó un mensaje viejo, confundió una respuesta del sistema o inventó el resultado después de que una herramienta…

  144. dev.to — LLM tag TIER_1 Español(ES) · Silviu Technology ·

    LLM Agents:邮件测试有明确限制

    <p>Los agentes LLM pueden ayudar a investigar por qué falla una prueba de email, pero hay una diferencia importante entre asistir una ejecución y controlar una ejecución. Si el modelo puede hacer cualquier cosa, el resultado será dificil de reproducir: una corrida pasa, la siguie…

  145. dev.to — LLM tag TIER_1 English(EN) · Dr Haina ·

    LLMs 与 AI Agent 详解:Haina Fatima 博士(XINI8 Engine)的实用指南

    <p>Hi, I'm Dr. Haina Fatima. I started my career as a physician and radiologist, and today I lead development and product at XINI8 Engine, an AI-native software engineering platform within the XINI8 ecosystem. Coming from outside traditional software, the terms "LLM" and "AI agen…

  146. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    AREX-2:通过长时程反思推进自改进大型语言模型代理

    <h1> Beyond One-Shot Success: How AREX-2 Teaches LLM Agents to Reflect and Persevere </h1> <p>Current autonomous LLM agents are often evaluated by their ability to solve a task in a single pass or through a short sequence of scripted interactions. While models like GPT-4o and Cla…

  147. dev.to — LLM tag TIER_1 Español(ES) · Silviu Technology ·

    使用可验证合约设计 LLM 代理

    <p>Un agente LLM puede parecer muy capaz durante una demostración y volverse dificil de operar cuando empieza a llamar APIs, crear tickets o consultar datos reales. El problema no suele ser que el modelo “no sepa” la respuesta. Es que el sistema no define con precisión qué puede …

  148. dev.to — LLM tag TIER_1 English(EN) · Quoc Bao An Nguyen ·

    我如何使用LLM函数调用构建AI代理(并避免不必要的工具调用)

    <h2> 1. Introduction </h2> <p>When I first started building my AI-powered course recommendation system, I thought integrating an LLM with backend APIs would be straightforward.</p> <p>However, I quickly ran into a key problem:<br /> The model was calling backend APIs almost every…

  149. Mastodon — mastodon.social TIER_1 Français(FR) · [email protected] ·

    Obelisk 0.42 引入持久化代理和分层沙箱以支持 LLM。安全角度具体:隔离 AI 代理能力是

    Obelisk 0.42 introduit des agents persistants et des sandboxes en couches pour les LLM. L'angle sécurité est concret : isoler les capacités des agents IA, c'est réduire la surface d'attaque quand un modèle est manipulé ou produit du code inattendu. Le sandboxing multicouche, vieu…