PulseAugur
实时 18:03:36
English(EN) Context-Aware RL for Agentic and Multimodal LLMs

新的基准和安全方法出现,用于先进的大模型代理

新研究探讨了AI代理的开发和评估,重点关注它们在复杂环境中导航和遵守策略的能力。StarDojo在《星露谷物语》等开放式模拟中对代理性能进行基准测试,揭示了视觉理解和推理方面的局限性。CostBench在动态旅行规划场景中评估大模型代理的成本最优规划和适应能力,显示出经济推理方面的显著差距。其他论文介绍了使用基于自我报告的具身大模型代理进行个体模拟的方法,开发用于高效多代理推理的符号通信,以及解决由长时域代理中的上下文压缩引起的“治理衰减”等安全问题。 AI

影响 代理评估和安全机制的进步可以加速更强大、更可靠的AI系统的开发和部署。

排序理由 多篇研究论文介绍了用于大模型代理的新基准、框架和分析。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 127 个来源。 我们如何撰写摘要 →

新的基准和安全方法出现,用于先进的大模型代理

报道来源 [127]

  1. arXiv cs.LG TIER_1 English(EN) · Kaixuan Liu, Guojun Xiong, Weinan Zhang, Shengpu Tang ·

    LLM智能体的社交网络

    arXiv:2607.03695v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed in interacting populations, raising the question of what such populations come to believe collectively. Whether a population aggregates genuine knowledge or collapses into …

  2. arXiv cs.AI TIER_1 English(EN) · Yining She, Yiliang Liang, Eunsuk Kang ·

    通过溯源分析保护大型语言模型代理免受失准影响

    arXiv:2607.01236v1 Announce Type: cross Abstract: As LLM agents gain increasing access to powerful tools, ensuring that their actions are aligned with the user's intent becomes critical. When an agent's proposed tool invocation deviates from the user's intent -- a phenomenon call…

  3. arXiv cs.AI TIER_1 English(EN) · Yunhao Feng, Ruixiao Lin, Ming Wen, Qinqin He, Yanming Guo, Yifan Ding, Yutao Wu, Jialuo Chen, Yunhao Chen, Xiaohu Du, Jianan Ma, Zixing Chen, Zhuoer Xu, Xingjun Ma, Xinhao Deng ·

    大规模 LLM 智能体安全测试:从风险发现到基于证据的验证

    arXiv:2607.01793v1 Announce Type: new Abstract: LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testing targets expert-designed safety violations, and the corresponding outcomes are …

  4. arXiv cs.AI TIER_1 English(EN) · Mahyar Ghazanfari, Amin Tabrizian, Armin Mehrabian, Peng Wei ·

    EO-Agents:一个用于地球观测假设生成的三个智能体LLM流水线

    arXiv:2607.01584v1 Announce Type: new Abstract: Large language models have recently been explored for scientific hypothesis generation, but most prior work relies on unstructured literature and free-form textual claims. We present a pipeline for Earth observation that grounds hyp…

  5. arXiv cs.LG TIER_1 English(EN) · Juanwu Lu, Junyu Zhu, Ziran Wang ·

    具有行为潜变量的可控模拟智能体

    arXiv:2607.02496v1 Announce Type: cross Abstract: Realistic traffic simulation requires agents that imitate logged behavior and can also be steered along interpretable axes. Such controllability enables engineers to isolate variables, reproduce specific edge cases, and test auton…

  6. arXiv cs.CL TIER_1 English(EN) · Zewen Liu ·

    BOUNDARY_SYNC:衡量多智能体LLM系统中通信诱导的表征耦合

    arXiv:2607.01600v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed as communicating agents, does inter-agent communication cause outputs to converge? We introduce BOUNDARY_SYNC, a protocol measuring representational coupling via the Coupling Amplificat…

  7. arXiv cs.AI TIER_1 English(EN) · Chih-Hsuan (Bella), Yang, Tanwi Mallick, Le Chen, Krishnan Raghavan, Amal Gueroudji, Ian T. Foster, Rajeev Thakur ·

    谁得奖,谁受责?面向多LLM智能体的评估对齐训练信号

    arXiv:2511.10687v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) in multi-agent systems (MAS) have shown promise for complex tasks, yet current training methods lack principled ways to connect system-level evaluation with agent- and message-level learning. W…

  8. arXiv cs.AI TIER_1 English(EN) · Jeffrey Flynt ·

    GroundEval:有状态智能体评估中 LLM-as-Judge 的确定性替代方案

    arXiv:2606.22737v2 Announce Type: replace Abstract: Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a deterministic test of what the agent searched, fetched, cited, and was permitted to access. In on…

  9. arXiv cs.AI TIER_1 English(EN) · Trilok Padhi, Ramneet Kaur, Krishiv Agarwal, Adam D. Cobb, Daniel Elenius, Manoj Acharya, Colin Samplawski, Alexander M. Berenbeim, Nathaniel D. Bastian, Susmit Jha, Ugur Kursuncu, Anirban Roy ·

    从行动到理解:LLM智能体中时间概念的保形可解释性

    arXiv:2604.19775v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents capable of reasoning, planning, and acting within interactive environments. Despite their growing capability to perform multi-step reasoning and decisio…

  10. arXiv cs.LG TIER_1 English(EN) · Ziran Wang ·

    具有行为潜变量的可控模拟代理

    Realistic traffic simulation requires agents that imitate logged behavior and can also be steered along interpretable axes. Such controllability enables engineers to isolate variables, reproduce specific edge cases, and test autonomous systems without real-world risk. We introduc…

  11. arXiv cs.AI TIER_1 English(EN) · Jingyuan Zheng, Dongjing Wang, Xin Zhang, Butian Huang, Haiping Zhang, Dongjin Yu, Shuguang Deng ·

    SkillSelect-Serve:面向小型LLM代理的预算可控且服务质量感知的技能服务推荐与组合

    arXiv:2607.00011v1 Announce Type: cross Abstract: Reusable skill libraries are becoming important infrastructure for large language model (LLM) agents, yet existing selection methods often treat skills as retrievable documents and return fixed top-k lists. This paper presents Ski…

  12. arXiv cs.AI TIER_1 English(EN) · Xubin Hao, Hongjin Meng, Xin Yin, Jiawei Zhu, Chenpeng Cao ·

    Self-GC:长远期LLM代理的自治理上下文

    arXiv:2607.00692v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate tool results, files, plans, and user constraints that are too structured to be treated as a disposable text suffix. Current systems mostly rely on in-run heuristics such as chronological pruning an…

  13. arXiv cs.LG TIER_1 English(EN) · Yiping Li, Zhiyu An, Wan Du ·

    当更少的潜在信息带来更好的中继:用于潜在多智能体LLM协作的信息保持压缩

    arXiv:2604.13349v2 Announce Type: replace Abstract: Communication in Large Language Model (LLM)-based multi-agent systems is moving beyond discrete tokens to preserve richer context. Recent work such as LatentMAS enables agents to exchange latent messages through full key-value (…

  14. arXiv cs.CL TIER_1 English(EN) · Zewen Liu ·

    EPC:一种用于衡量 LLM Agent 系统中评估者偏好动态的标准协议

    arXiv:2607.00297v1 Announce Type: cross Abstract: When LLM agents use evaluator feedback to adapt their behavior in closed loops, evaluator biases propagate through the agent's strategy distribution -- a phenomenon known as evaluator preference coupling. Prior work has documented…

  15. arXiv cs.CL TIER_1 English(EN) · Elias Najarro, Ane Espeseth, Eleni Nisioti, Sebastian Risi, Stefano Nichele ·

    可对话的复杂性:作为可解释基底的智能体LLM集合

    arXiv:2607.01047v1 Announce Type: new Abstract: Complexity and interpretability rarely coincide: systems rich enough for complex behaviours to emerge are usually too opaque to question, while transparent ones are too simple for anything complex to emerge. A single large language …

  16. arXiv cs.CL TIER_1 English(EN) · Zewen Liu ·

    BOUNDARY_SYNC:衡量多智能体LLM系统中通信诱导的表征耦合

    As large language models (LLMs) are deployed as communicating agents, does inter-agent communication cause outputs to converge? We introduce BOUNDARY_SYNC, a protocol measuring representational coupling via the Coupling Amplification Factor (CAF = JSD_cond / JSD_baseline), where …

  17. arXiv cs.CL TIER_1 English(EN) · Stefano Nichele ·

    可对话的复杂性:作为可解释基底的智能体LLM集合

    Complexity and interpretability rarely coincide: systems rich enough for complex behaviours to emerge are usually too opaque to question, while transparent ones are too simple for anything complex to emerge. A single large language model (LLM) is a static artefact, hardly exhibit…

  18. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Mirko Degli Esposti ·

    校准仪器:LLM驱动的合成人群的可控性

    Generative Synthetic Populations (GSP) -- the convergence of population synthesis, agent-based modelling, and LLM agents -- are attracting growing interest for urban simulation and institutional communication research. Before any GSP instrument is used on a real population, a mor…

  19. arXiv cs.AI TIER_1 English(EN) · Chenpeng Cao ·

    Self-GC:面向长时域大语言模型智能体的自治理上下文

    Long-horizon LLM agents accumulate tool results, files, plans, and user constraints that are too structured to be treated as a disposable text suffix. Current systems mostly rely on in-run heuristics such as chronological pruning and tool-output masking, or on final self-summary …

  20. arXiv cs.AI TIER_1 English(EN) · Javal Vyas, Milapji Singh Gill, Artan Markaj, Felix Gehlhoff, Mehmet Mercang\"oz ·

    基于知识图谱的大语言模型智能体在自主容错控制中的应用教程

    arXiv:2606.31635v1 Announce Type: cross Abstract: Fault recovery in process plants still relies heavily on plant operators, especially when faults fall outside predefined supervisory logic. Operators interpret alarms, procedures, P\&IDs, interlocks, and process trends, then d…

  21. arXiv cs.LG TIER_1 English(EN) · Idelfonso B. R. Nogueira, Sigurd Skogestad ·

    从高级监管控制理论系统化多智能体AI:用于过程控制的安全可审计LLM操作员代理

    arXiv:2606.30877v1 Announce Type: cross Abstract: Recent literature shows that large language models (LLMs) are useful for general-purpose tasks yet perform poorly on specific domain ones. One reason is the difficulty of supplying narrow context to a general-purpose model and of …

  22. arXiv cs.CL TIER_1 English(EN) · Xinyu Zhao, Zhen Tan, Vaishnav Tadiparthi, Nakul Agarwal, Kwonjoon Lee, Ehsan Moradi Pari, Hossein Nourkhiz Mahjoub, Tianlong Chen ·

    生成式技能组合用于LLM代理

    arXiv:2606.32025v1 Announce Type: new Abstract: Recent LLM agents benefit from skills for solving complex tasks. Skills encapsulate modular packages of procedural knowledge and instructions for performing specialized tasks, such as setting up a sandboxed environment, running a te…

  23. arXiv cs.AI TIER_1 English(EN) · Bang Nguyen, Dominik So\'os, Qian Ma, Rochana R. Obadage, Zack Ranjan, Sai Koneru, Timothy M. Errington, Shakhlo Nematova, Sarah Rajtmajer, Jian Wu, Meng Jiang ·

    ReplicatorBench:为社会与行为科学中的可复现性基准测试大型语言模型代理

    arXiv:2602.11354v3 Announce Type: replace Abstract: The literature has witnessed an emerging interest in AI agents for automated assessment of scientific papers. Existing benchmarks focus primarily on the computational aspect of this task, testing agents' ability to reproduce or …

  24. arXiv cs.AI TIER_1 English(EN) · Renxuan Tan, Rongpeng Li, Fei Wang, Chenghui Peng, Shaoyun Wu, Zhifeng Zhao, Honggang Zhang ·

    LLM赋能的Agentic MAC协议:一种动态Stackelberg博弈方法

    arXiv:2510.10895v2 Announce Type: replace Abstract: Medium Access Control (MAC) protocols, essential for wireless networks, are typically manually configured. While deep reinforcement learning (DRL)-based protocols enhance task-specified network performance, they suffer from poor…

  25. arXiv cs.AI TIER_1 English(EN) · Sergio Hern\'andez-Guti\'errez, Matteo Merler, Ilze Amanda Auzina, Joschka Str\"uber, Ameya Prabhu, Matthias Bethge ·

    QVal:廉价评估长时域 LLM Agent 的密集监督信号

    arXiv:2606.32034v1 Announce Type: cross Abstract: LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of actions. In these settings, outcome-only rewards provide too sparse guidance, failing to inform the model about the goo…

  26. arXiv cs.AI TIER_1 English(EN) · Zewen Liu ·

    校准评估器:概率校准能否缓解 LLM 代理反馈循环中的偏好耦合?

    arXiv:2606.31371v1 Announce Type: cross Abstract: When large language model (LLM) agents adapt their behavior through evaluator feedback, systematic evaluator biases propagate into the agent's learned strategy distribution - a phenomenon termed evaluator preference coupling. Prio…

  27. arXiv cs.AI TIER_1 Deutsch(DE) · Simon Jones, Sabine Hauert ·

    最小化LLM系统中涌现的文化

    arXiv:2606.30668v1 Announce Type: cross Abstract: What happens when LLM agents operate with no context outside a turn, minimal prompting, and simple tools? Inspired by swarm engineering, we give collectives of three agents the ability to send messages and manipulate a shared acti…

  28. arXiv cs.AI TIER_1 English(EN) · Atsushi Masumori, Itsuki Doi, Norihiro Maruyama, Ryosuke Takata, Takashi Ikegami ·

    OpenLife:迈向具有自主LLM智能体的开放世界人工智能生命

    arXiv:2606.31046v1 Announce Type: new Abstract: Artificial life has explored life-like behavior on many computational substrates, but mostly in researcher-designed closed worlds. We argue that large language model (LLM) agents, with persistent memory, tool use, network access, an…

  29. arXiv cs.CL TIER_1 English(EN) · Zewen Liu ·

    EPC:一种用于衡量 LLM 代理系统中评估者偏好动态性的标准化协议

    When LLM agents use evaluator feedback to adapt their behavior in closed loops, evaluator biases propagate through the agent's strategy distribution -- a phenomenon known as evaluator preference coupling. Prior work has documented coupling across multiple evaluator families and m…

  30. arXiv cs.AI TIER_1 English(EN) · Matthias Bethge ·

    QVal:廉价评估长时域 LLM Agent 的密集监督信号

    LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of actions. In these settings, outcome-only rewards provide too sparse guidance, failing to inform the model about the goodness of intermediate actions. Dense supervision m…

  31. arXiv cs.CL TIER_1 English(EN) · Tianlong Chen ·

    面向LLM智能体的生成式技能组合

    Recent LLM agents benefit from skills for solving complex tasks. Skills encapsulate modular packages of procedural knowledge and instructions for performing specialized tasks, such as setting up a sandboxed environment, running a test suite, or refactoring a function across multi…

  32. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Mehmet Mercangöz ·

    基于知识图谱大语言模型智能体实现自主容错控制的教程

    Fault recovery in process plants still relies heavily on plant operators, especially when faults fall outside predefined supervisory logic. Operators interpret alarms, procedures, P\&IDs, interlocks, and process trends, then decide how to move the plant to a safe operating mode w…

  33. arXiv cs.CL TIER_1 English(EN) · Zewen Liu ·

    校准评估器:概率校准能否缓解 LLM 代理反馈循环中的偏好耦合?

    When large language model (LLM) agents adapt their behavior through evaluator feedback, systematic evaluator biases propagate into the agent's learned strategy distribution - a phenomenon termed evaluator preference coupling. Prior work has documented this coupling and establishe…

  34. arXiv cs.AI TIER_1 English(EN) · Mengdie Flora Wang, Haochen Xie, Guanghui Wang, Devin Zhang, Jae Oh Woo ·

    预算驱动的Act-or-Defer多智能体LLM审议与局部可靠性界限

    arXiv:2606.29654v1 Announce Type: new Abstract: Multi-agent deliberation among LLMs can improve reasoning, but deployment requires deciding when the current answer is reliable enough to act on and when it should be escalated to human review. We formulate this as budgeted act-or-d…

  35. arXiv cs.AI TIER_1 English(EN) · Zhengqi Pei, Qingming Huang, Shuhui Wang ·

    当大型语言模型开发语言:用于高效多智能体推理的符号通信

    arXiv:2606.29354v1 Announce Type: new Abstract: Chain-of-Thought (CoT) improves large language models (LLMs) on difficult reasoning tasks, but it often incurs long natural-language rationales that are poorly aligned with efficient machine reasoning. We propose Communicative Langu…

  36. arXiv cs.LG TIER_1 English(EN) · Huaijie Wang, Shusheng Xu, Yi Wu, Kaifeng Lyu ·

    通过两阶段蒸馏构建多任务代理LLM

    arXiv:2606.30044v1 Announce Type: new Abstract: A key step toward artificial general intelligence is to train models that can perform multiple tasks. In this paper, we study how to build such models by first training separate RL experts for individual tasks and then consolidating…

  37. arXiv cs.LG TIER_1 English(EN) · Zewen Liu ·

    传染张量:一种衡量多智能体LLM系统中输出分布耦合的框架——以及审计其支持的声明

    arXiv:2606.28839v1 Announce Type: new Abstract: We introduce the Contagion Tensor, a measurement framework for quantifying how large language model (LLM) output distributions couple across modalities, agents, and time steps. From the tensor we derive the Coupling Amplification Fa…

  38. arXiv cs.CL TIER_1 English(EN) · Liu Zewen ·

    面向自适应LLM智能体的评估者驱动偏好动态的诊断框架与多评估者审计

    arXiv:2606.29719v1 Announce Type: cross Abstract: Measurements of proprietary LLM evaluators can become invalid within weeks -- we document one case and provide the diagnostic framework to detect it. We introduce EPC -- comprising the Multimodal Preference Collapse Index (MPCI), …

  39. arXiv cs.CL TIER_1 English(EN) · Sebastian Kula, Martin Tamajka ·

    利用开源大语言模型的多智能体系统以缓解虚假信息威胁

    arXiv:2606.30259v1 Announce Type: new Abstract: In contemporary societies, the threat of disinformation has reached alarming levels, exacerbated by the proliferation of electronic communication, social media, and advancements in artificial intelligence. As a result, there is an u…

  40. arXiv cs.AI TIER_1 Deutsch(DE) · Haejoon Lee, Vincent-Daniel Yun, Dimitra Panagou, Sai Praneeth Karimireddy ·

    拜占庭容错下的鲁棒多智能体LLM

    arXiv:2605.09076v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents increasingly collaborate over peer-to-peer networks to improve their reliability. However, these same interactions can also become a source of vulnerability, as unreliable or Byzantine age…

  41. arXiv cs.AI TIER_1 English(EN) · Brian Y. Tsui, Alan Y. Fang, Tiffany J. Hwu ·

    LLM智能体实现无需演示的机器人控制

    arXiv:2601.20334v2 Announce Type: replace-cross Abstract: Robotic manipulation has increasingly adopted vision-language-action (VLA) models, which achieve strong performance but typically require task-specific demonstrations and fine-tuning, and often generalize poorly under doma…

  42. arXiv cs.AI TIER_1 English(EN) · Shiyang Chen ·

    治理衰败:上下文压缩如何悄然消除长时域LLM代理中的安全约束

    arXiv:2606.22528v2 Announce Type: replace Abstract: Modern LLM agents increasingly rely on context compaction, summarization, or eviction to keep long-running sessions within a token budget. We show that this context-management layer is a safety-critical failure surface: in-conte…

  43. arXiv cs.AI TIER_1 English(EN) · Weihao Tan, Changjiu Jiang, Yu Duan, Mingcong Lei, Jiageng Li, Yitian Hong, Xinrun Wang, Bo An ·

    StarDojo:在《星露谷物语》的生产性生活模拟中,对Agentic多模态LLM的开放式行为进行基准测试

    arXiv:2507.07445v3 Announce Type: replace Abstract: Autonomous agents navigating human society must master both production activities and social interactions, yet existing benchmarks rarely evaluate these skills simultaneously. To bridge this gap, we introduce StarDojo, a novel b…

  44. arXiv cs.AI TIER_1 English(EN) · Henrique Ferraz de Arruda, Carlos Gracia L\'azaro, Alberto Aleta, Yamir Moreno ·

    LLM代理中的集体合作与个体忠诚度缺失

    arXiv:2606.30454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as agents in simulations of social systems, yet it remains unclear when their behavior can be interpreted as a faithful proxy for human decision-making. Here we test LLM agents ag…

  45. arXiv cs.AI TIER_1 English(EN) · David Mellafe Zuvic ·

    能力门禁并非授权:LLM代理框架中的混淆副官漏洞

    arXiv:2606.28679v1 Announce Type: cross Abstract: Tool-using LLM agents increasingly read untrusted content while holding side-effecting tools such as payments, email, CRM, and infrastructure APIs, yet common framework defaults still conflate tool exposure with authorization. We …

  46. arXiv cs.AI TIER_1 English(EN) · Seongjae Kang, Taehyung Yu, Sung Ju Hwang ·

    PolicyGuard:用于 LLM Agent 策略遵循的对话式子代理验证器

    arXiv:2606.29225v1 Announce Type: new Abstract: LLM agents handle user requests on behalf of organizations through tool calls and must follow the company policies stated in their system prompts. Prior work approaches this as a safeguarding problem -- external checks that block no…

  47. arXiv cs.AI TIER_1 English(EN) · Xuan Zhang, Wenxuan Zhang, See-Kiong Ng, Yang Deng ·

    用于LLM智能体规划的自演化世界模型

    arXiv:2606.30639v1 Announce Type: new Abstract: World models offer a principled way to equip long-horizon LLM agents with foresight: predictions of action consequences before execution. However, unreliable foresight can be ignored, misused, or even degrade downstream decision-mak…

  48. arXiv cs.AI TIER_1 English(EN) · Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Percy Liang, Robb Willer, Michael S. Bernstein ·

    基于自我报告的LLM代理能够实现通用个体模拟

    arXiv:2411.10109v3 Announce Type: replace Abstract: Machine learning can predict human behavior well when substantial structured data are available for well-defined outcomes. Such models are typically outcome-specific, however, requiring training data for each target outcome, lim…

  49. arXiv cs.AI TIER_1 English(EN) · Jiayu Liu, Cheng Qian, Zhaochen Su, Qing Zong, Shijue Huang, Bingxiang He, Yi R. Fung ·

    CostBench:评估LLM工具使用代理在动态环境中多轮成本最优规划与适应性

    arXiv:2511.02734v3 Announce Type: replace Abstract: Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and adaptability. This neglects a crucial capability: agents' ability to devise and adjust cost-…

  50. Hugging Face Daily Papers TIER_1 English(EN) ·

    QVal:廉价评估长时域 LLM Agent 的密集监督信号

    A testbed called QVal is introduced for evaluating dense supervision signals in long-horizon LLM agent tasks by measuring how well method scores align with Q-values, enabling fair comparison of different supervision approaches without training.

  51. arXiv cs.AI TIER_1 English(EN) · Yang Deng ·

    用于LLM智能体规划的自演化世界模型

    World models offer a principled way to equip long-horizon LLM agents with foresight: predictions of action consequences before execution. However, unreliable foresight can be ignored, misused, or even degrade downstream decision-making. In this paper, we introduce WorldEvolver, a…

  52. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Jacques Samain ·

    MAS-Lab:可靠多智能体系统的面向规范的验证框架

    The rapid emergence of LLM-based agentic frameworks has significantly reduced the cost of assembling multi-agent systems (MAS), enabling fast prototyping and exploration of agentic behaviors. However, systems built with current tooling remain ill-suited for reliable, evolvable, a…

  53. arXiv cs.AI TIER_1 English(EN) · Yamir Moreno ·

    LLM代理中的集体合作与个体忠诚度缺失

    Large language models (LLMs) are increasingly used as agents in simulations of social systems, yet it remains unclear when their behavior can be interpreted as a faithful proxy for human decision-making. Here we test LLM agents against a direct empirical benchmark: a large-scale …

  54. arXiv cs.CL TIER_1 English(EN) · Martin Tamajka ·

    利用开源大语言模型的多智能体系统以减轻虚假信息威胁

    In contemporary societies, the threat of disinformation has reached alarming levels, exacerbated by the proliferation of electronic communication, social media, and advancements in artificial intelligence. As a result, there is an urgent need to develop effective countermeasures …

  55. Hugging Face Daily Papers TIER_1 English(EN) ·

    LLM 智能体是潜在的上下文管理器:通过本体感知仪表板引发自管理上下文

    Long-horizon tool agents are bottlenecked by how their context grows toward the limits of the context window. Recent systems make context management agent- or system-controlled, but they either learn a compression policy that discards evidence or manage context in a layer the age…

  56. arXiv cs.LG TIER_1 English(EN) · Javal Vyas, Milapji Singh Gill, Artan Markaj, Felix Gehlhoff, Mehmet Mercang\"oz ·

    从检测到行动:使用 LLM Agent 实现容错控制

    arXiv:2606.28011v1 Announce Type: cross Abstract: We propose an agentic Large Language Model (LLM) framework for active Fault-Tolerant Control (FTC) that transforms fault detection outputs into constraint-aware recovery actions grounded in plant-specific knowledge. The approach c…

  57. arXiv cs.AI TIER_1 English(EN) · Xiang Long, Li Du, Yilong Xu, RongJian Xu, Qiyanhui Lu, Ying Gao, Qinhua Xie, Fangcheng Liu, Ning Ding, Haoqing Wang, Ziheng Li, Changjiang Zhou, Jianyuan Guo, Yehui Tang ·

    LiveClawBench:在复杂、真实世界的助手任务上对 LLM Agent 进行基准测试

    arXiv:2604.13072v2 Announce Type: replace-cross Abstract: OpenClaw-style personal assistants extend LLM agents from isolated tool use to open-ended, stateful, and personalized software environments. Evaluating these assistants is fundamentally a fidelity problem: benchmarks must …

  58. arXiv cs.AI TIER_1 English(EN) · Xinyuan Song, Zekun Cai ·

    基于现实的迭代语言规划:参数化世界模型如何减少LLM代理中的幻觉传播

    arXiv:2606.27806v1 Announce Type: new Abstract: World models for language agents come in two useful forms. An agent-based world model calls an LLM API and reasons flexibly in language, but its errors appear as hallucinated state changes that are hard to score with ordinary regres…

  59. arXiv cs.AI TIER_1 English(EN) · Ronny Ko, Jiseong Jeong, Shuyuan Zheng, Chuan Xiao, Tae-Wan Kim, Makoto Onizuka, Won-Yong Shin ·

    跨域多智能体LLM系统中必须解决的七个安全挑战

    arXiv:2505.23847v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are rapidly evolving into autonomous agents that cooperate across organizational boundaries, enabling joint disaster response, supply-chain optimization, and other tasks that demand decentraliz…

  60. arXiv cs.AI TIER_1 English(EN) · Xuan Zhang, Zhijian Zhou, Lingfeng Qiao, Yulei Qin, Ke Li, Xing Sun, Xiaoyu Tan, Chao Qu, Yuan Qi ·

    内部化未来:世界模型规划的统一智能体训练范式

    arXiv:2606.27483v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated strong capability in sequential decision-making, yet they remains fundamentally reactive in long-horizon tasks. Unlike humans who employ "what-if" reasoning to evaluate potential p…

  61. arXiv cs.CL TIER_1 English(EN) · Igor Itkin ·

    延迟验证破坏多智能体LLM信念:不稳定性阈值与最优纠错器放置

    arXiv:2606.27409v1 Announce Type: cross Abstract: Multi-agent large language model (LLM) systems often rely on verifier and critic agents to suppress hallucinations, but verification is delayed. During this delay, false claims can propagate through the agent network. We model thi…

  62. arXiv cs.LG TIER_1 English(EN) · Chuanhao Li, Xiaoan Xu, Dirk Bergemann, Ethan X. Fang, Yehua Wei, Zhuoran Yang ·

    COOPA:面向运筹学问题的模块化LLM代理架构

    arXiv:2606.27611v1 Announce Type: new Abstract: Operations Research (OR) provides a rigorous framework for high-stakes decision-making, but effective OR modeling requires substantial domain knowledge, mathematical abstraction, and solver expertise. Recent LLM-based systems automa…

  63. arXiv cs.CL TIER_1 English(EN) · Liu Zewen ·

    面向自适应LLM代理的评估者驱动偏好动态的诊断框架与多评估者审计

    Measurements of proprietary LLM evaluators can become invalid within weeks -- we document one case and provide the diagnostic framework to detect it. We introduce EPC -- comprising the Multimodal Preference Collapse Index (MPCI), evaluator-indexed coupling matrix, and Jensen-Shan…

  64. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Jae Oh Woo ·

    预算驱动的Act-or-Defer多智能体LLM审议与局部可靠性界限

    Multi-agent deliberation among LLMs can improve reasoning, but deployment requires deciding when the current answer is reliable enough to act on and when it should be escalated to human review. We formulate this as budgeted act-or-defer decision making. At each round, the system …

  65. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Shuhui Wang ·

    当大型语言模型开发语言:用于高效多智能体推理的符号通信

    Chain-of-Thought (CoT) improves large language models (LLMs) on difficult reasoning tasks, but it often incurs long natural-language rationales that are poorly aligned with efficient machine reasoning. We propose Communicative Language Symbolism Routing (CLSR), a test-time framew…

  66. Hugging Face Daily Papers TIER_1 English(EN) ·

    PolicyGuard:一个基于对话的子代理验证器,用于 LLM Agent 的策略遵循

    POLICYGUARD is a sub-agent verifier that enhances LLM agent policy adherence by providing contextual reasoning and conversation-specific feedback across multi-turn interactions.

  67. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Carlos Baquero ·

    当潜在代理撒谎时:多代理LLM协作中的KV缓存完整性

    LLM agents can share more than text. In some systems, an agent can send a short visible message while also passing its full KV-cache state to another model. This hidden state can help the final model combine evidence from several agents, but it is also hard to inspect. A visible …

  68. arXiv cs.LG TIER_1 English(EN) · Mehmet Mercangöz ·

    从检测到行动:使用 LLM 代理实现容错控制

    We propose an agentic Large Language Model (LLM) framework for active Fault-Tolerant Control (FTC) that transforms fault detection outputs into constraint-aware recovery actions grounded in plant-specific knowledge. The approach couples (i) a multi-agent workflow that decomposes …

  69. arXiv cs.AI TIER_1 English(EN) · Zekun Cai ·

    基于现实的迭代语言规划:参数化世界模型如何减少LLM代理中的幻觉传播

    World models for language agents come in two useful forms. An agent-based world model calls an LLM API and reasons flexibly in language, but its errors appear as hallucinated state changes that are hard to score with ordinary regression losses. A parameterized world model is a tr…

  70. arXiv cs.AI TIER_1 English(EN) · Luyang Zhang, Jialu Wang, Fei Xue, Yi-Yun Chu ·

    训练后配方,而非模型家族,塑造了多智能体LLM的对话行为

    arXiv:2606.20632v2 Announce Type: replace-cross Abstract: Multi-LLM systems use multiple language models to deliberate, judge each other's outputs, or coordinate as agents. Their value depends on the models producing measurably different conversational behaviors when given the sa…

  71. arXiv cs.AI TIER_1 English(EN) · David Akinpelu, Akintonde Abbas, Rereloluwa Alimi, Ayodeji Lana ·

    工具增强型LLM代理在现实世界能源分析任务中的表现如何?

    arXiv:2606.26346v1 Announce Type: new Abstract: Agentic benchmarks have emerged across general-purpose and domain-specific settings, including finance, coding, law, and drug discovery, yet energy-domain evaluations remain largely limited to static knowledge recall. This is a crit…

  72. arXiv cs.AI TIER_1 English(EN) · Sahil Shrivastava ·

    面向迭代式LLM代理循环的语义提前停止

    arXiv:2606.27009v1 Announce Type: new Abstract: Multi-agent large language model (LLM) loops, for example a Writer that drafts and a Critic that revises, are almost always terminated by a fixed iteration cap (max_iterations). This is a syntactic kill-switch: it is blind to whethe…

  73. arXiv cs.AI TIER_1 English(EN) · Praneeth Narisetty, Shiva Nagendra Babu Kore, Uday Kumar Reddy Kattamanchi, Jayaram Kumarapu ·

    LLM智能体提示注入带外防御的自适应评估

    arXiv:2606.26479v1 Announce Type: cross Abstract: Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than training the model to refuse malicious instructions, enforce security outside the model with…

  74. arXiv cs.CL TIER_1 English(EN) · Sriram Selvam, Anneswa Ghosh ·

    ProfileFoundry:用于LLM代理隐私、记忆和工具使用评估的合成人物对象基础

    arXiv:2606.26403v1 Announce Type: new Abstract: Foundation-model research increasingly needs data about people: user state, personal histories, relationships, contact-like fields, documents, and longitudinal updates. Real user data is difficult to share, perturb, audit, or redist…

  75. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Tao Feng ·

    GenWorld:用于可扩展 LLM 代理研究的经验性城市模拟基础设施

    LLM-agent simulation faces a joint grounding and scaling problem: agents should act in environments that reflect real urban constraints, yet direct online LLM calls for city-scale populations are computationally prohibitive. We present GenWorld, an empirically grounded urban simu…

  76. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Jiaming Cui ·

    QueenBee Planner:面向令牌高效大模型多智能体系统的技能演化通信拓扑

    Large language model (LLM) multi-agent systems increasingly depend not only on how individual agents reason, but also on how agents are connected. This paper introduces QueenBee Planner, a framework that treats inter-agent communication topology as a retrievable and self-improvin…

  77. arXiv cs.LG TIER_1 English(EN) · Sahil Shrivastava ·

    面向迭代式LLM代理循环的语义提前停止

    Multi-agent large language model (LLM) loops, for example a Writer that drafts and a Critic that revises, are almost always terminated by a fixed iteration cap (max_iterations). This is a syntactic kill-switch: it is blind to whether the answer is still improving, so it over-spen…

  78. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Igor Itkin ·

    延迟验证破坏多智能体LLM信念:不稳定性阈值与最优纠正器放置

    Multi-agent large language model (LLM) systems often rely on verifier and critic agents to suppress hallucinations, but verification is delayed. During this delay, false claims can propagate through the agent network. We model this process as delayed consensus on a graph with gro…

  79. arXiv cs.CL TIER_1 English(EN) · Kyungmin Kim, Youngbin Choi, Seoyeon Lee, Suhyeon Jun, Dongwoo Kim, Sangdon Park ·

    LLM智能体中绳索设计与训练后技术的相互作用

    arXiv:2606.25447v1 Announce Type: cross Abstract: Tool-integrated LLM agents are often wrapped within a harness: the scaffolding that determines which tools are exposed, how they are described, and what auxiliary information accompanies each per-step observation. While agents are…

  80. arXiv cs.LG TIER_1 English(EN) · Changdae Oh, Wendi Li, Seongheon Park, Samuel Yeh, Tanwi Mallick, Sharon Li ·

    训练后被忽视的免费午餐:LLM智能体的进步优势

    arXiv:2606.26080v1 Announce Type: new Abstract: Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings remains prohibitively difficult: long-horizon interactions, irreversible actions, and stochastic environment feedback m…

  81. arXiv cs.CL TIER_1 English(EN) · Jayaram Kumarapu ·

    LLM智能体提示注入带外防御的自适应评估

    Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than training the model to refuse malicious instructions, enforce security outside the model with a deterministic policy that mediates the agent's …

  82. Hugging Face Daily Papers TIER_1 English(EN) ·

    延迟验证破坏多智能体LLM信念:不稳定性阈值与最优纠错器放置

    Delayed verification in multi-agent LLM systems can cause instability leading to oscillations, but grounded factual answering stabilizes the system by making truth an absorbing boundary.

  83. arXiv cs.CL TIER_1 English(EN) · Anneswa Ghosh ·

    ProfileFoundry:用于LLM代理中隐私、记忆和工具使用评估的合成人-对象底层

    Foundation-model research increasingly needs data about people: user state, personal histories, relationships, contact-like fields, documents, and longitudinal updates. Real user data is difficult to share, perturb, audit, or redistribute responsibly, while independently generate…

  84. arXiv cs.AI TIER_1 English(EN) · Sharon Li ·

    训练后被忽视的免费午餐:LLM智能体的进步优势

    Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings remains prohibitively difficult: long-horizon interactions, irreversible actions, and stochastic environment feedback make both human annotation and Monte Carlo estima…

  85. arXiv cs.CL TIER_1 English(EN) · Sangdon Park ·

    LLM智能体中绳索设计与训练后技术的相互作用

    Tool-integrated LLM agents are often wrapped within a harness: the scaffolding that determines which tools are exposed, how they are described, and what auxiliary information accompanies each per-step observation. While agents are routinely post-trained, this scaffolding is typic…

  86. arXiv cs.AI TIER_1 English(EN) · Khanak Khandelwal (Indian Institute of Technology Jodhpur) ·

    AdversaBench:自动化大语言模型红队测试,具备多裁判确认和跨模型可迁移性

    arXiv:2606.24589v1 Announce Type: new Abstract: Scaling adversarial evaluation of large language models requires both a method for generating hard inputs and a reliable way to confirm that resulting failures are real. We present AdversaBench, an end-to-end red-teaming pipeline th…

  87. arXiv cs.AI TIER_1 English(EN) · Pingchuan Ma, Zhaoyu Wang, Zimo Ji, Yuguang Zhou, Zhantong Xue, Zongjie Li, Shuai Wang, Xiaoqin Zhang ·

    AutoSpec:通过归纳逻辑编程实现大语言模型Agent的安全规则演进

    arXiv:2606.24245v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly automate complex tasks by integrating language models with external tools and environments. However, their autonomy poses significant safety risks: agents may execute destructive comm…

  88. arXiv cs.LG TIER_1 Română(RO) · Kevin Qiu, Marek Cygan ·

    Debate2Create: 机器人通过多智能体LLM辩论进行联合设计

    arXiv:2510.25850v3 Announce Type: replace-cross Abstract: We introduce Debate2Create (D2C), a multi-agent LLM framework that formulates robot co-design as structured, iterative debate grounded in physics-based evaluation. A design agent and control agent engage in a thesis-antith…

  89. Hugging Face Daily Papers TIER_1 English(EN) ·

    训练后被忽视的免费午餐:LLM智能体的进步优势

    Reinforcement learning post-training enables effective step-level scoring for language models without requiring dedicated reward model training by deriving an implicit advantage function called progress advantage.

  90. arXiv cs.AI TIER_1 English(EN) · Khanak Khandelwal ·

    AdversaBench:自动化大语言模型红队测试,具备多裁判确认和跨模型可迁移性

    Scaling adversarial evaluation of large language models requires both a method for generating hard inputs and a reliable way to confirm that resulting failures are real. We present AdversaBench, an end-to-end red-teaming pipeline that mutates seed prompts with five structured ope…

  91. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Laixi Shi ·

    MAS-PromptBench:提示优化何时能改善多智能体LLM系统?

    Multi-agent systems (MAS) offer a scalable path forward for agentic AI, comprising multiple LLM-based agents, each assigned a system prompt and a position within a workflow that governs inter-agent coordination and output aggregation. System prompts thus form a critical and acces…

  92. arXiv cs.CL TIER_1 English(EN) · Hannaneh Hajishirzi ·

    Tmax:终端代理的简单方法

    Terminal-using agents have quickly become the most popular downstream application of language models (LMs). Despite their prevalence, relatively little academic work has examined RL-based training of these models, likely due to difficult benchmarks, a lack of data, and a lack of …

  93. arXiv cs.CL TIER_1 English(EN) · Anupam Datta ·

    计划不会持久:为什么上下文管理对LLM代理至关重要

    Long-horizon agents depend on context management: systems compress, summarize, and evict old tokens so tasks can continue beyond finite windows. That is safe only when dropped information is no longer needed or has been internalized. Plans are the stress case: they are written ea…

  94. Hugging Face Daily Papers TIER_1 English(EN) ·

    GroundEval:有状态智能体评估中 LLM-as-Judge 的确定性替代方案

    Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a deterministic test of what the agent searched, fetched, cited, and was permitted to access. In one case study, two frontier LLM judges scored a plaus…

  95. arXiv cs.CL TIER_1 English(EN) · Jeffrey Flynt ·

    GroundEval:有状态智能体评估中 LLM-as-Judge 的确定性替代方案

    Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a deterministic test of what the agent searched, fetched, cited, and was permitted to access. In one case study, two frontier LLM judges scored a plaus…

  96. Hugging Face Daily Papers TIER_1 English(EN) ·

    当智能体过早承诺时:诊断大型语言模型智能体的过早承诺问题

    Pre premature commitment in long-horizon LLM agents leads to silent failures where agents defend early interpretations without considering alternatives, and hidden-state convergence serves as an early diagnostic for trajectory consistency.

  97. Hugging Face Daily Papers TIER_1 English(EN) ·

    计划无法持久:为什么上下文管理对LLM代理至关重要

    Standard LLM agents rely on plan content remaining in context rather than maintaining it as persistent state, with evidence shown through replay pairing diagnostics and compression stress tests.

  98. Hugging Face Daily Papers TIER_1 English(EN) ·

    Tmax:终端智能体的简单方法

    A novel RL training approach for terminal agents achieves superior performance using a simplified recipe and expanded dataset, enabling effective training with fewer parameters than previous methods.

  99. Hugging Face Daily Papers TIER_1 English(EN) ·

    CLI-Universe:面向终端代理的可验证任务合成引擎

    A principled synthesis engine generates high-quality terminal-agent tasks through multi-dimensional capability taxonomy and evidence-guided research, creating a distilled dataset that enables significant performance gains in LLM training.

  100. arXiv cs.NE (Neural & Evolutionary) TIER_1 Deutsch(DE) · Sabine Hauert ·

    最小化LLM系统中涌现的文化

    What happens when LLM agents operate with no context outside a turn, minimal prompting, and simple tools? Inspired by swarm engineering, we give collectives of three agents the ability to send messages and manipulate a shared actively decaying text store, introducing evolutionary…

  101. arXiv cs.AI TIER_1 English(EN) · Shiyang Chen ·

    治理衰败:上下文压缩如何悄无声息地消除长视界LLM代理中的安全约束

    Modern LLM agents increasingly rely on context compaction, summarization, or eviction to keep long-running sessions within a token budget. We show that this context-management layer is a safety-critical failure surface: in-context governance constraints that agents reliably obey …

  102. arXiv cs.AI TIER_1 English(EN) · Yehui Yang ·

    Hypothesis-Driven Skill Optimization for LLM Agents

    External skills can improve action-oriented LLM agents without changing model weights, but persistent skill updates are risky when they are distilled from sparse or noisy trajectories. A plausible reflection may encode a useful procedure, a spurious shortcut, or a rule that the t…

  103. arXiv cs.AI TIER_1 English(EN) · Shreyas KC ·

    BabelJudge:衡量大型语言模型作为裁判的跨语言和代理轨迹可靠性

    LLM-as-a-judge has become the dominant approach to scalable evaluation in NLP pipelines, yet judges themselves carry systematic biases that raw accuracy hides: they favor responses placed in slot A (position bias), they prefer longer responses regardless of quality (verbosity bia…

  104. Hugging Face Daily Papers TIER_1 English(EN) ·

    训练编排器:一种端到端 PDDL 规划的监督学习方法,结合 LLM 智能体

    Translating natural-language planning intent into verified plans is a longstanding challenge: people communicate goals in language, while classical planners require formal PDDL specifications. Recent agentic frameworks bridge this gap by orchestrating a pool of specialized repair…

  105. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Zewen Liu ·

    Contagion Networks: 评估者偏好传播在多智能体LLM系统中的应用

    When large language models serve as evaluators in multi-agent systems, their strategy preferences -- whether induced by explicit prompts or by shared architectural priors -- propagate through the agent network. We introduce Contagion Networks, a formal framework for measuring how…

  106. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Suranjan Goswami ·

    PACMS:子模组化上下文选择作为LLM代理的可插拔引擎

    Conversational and tool-using LLM agents operate over a context window that fills from several directions simultaneously. As a session proceeds, the agent accumulates user and assistant turns, entries drawn from a persistent memory store, and often largest of all, the verbatim ou…

  107. Hugging Face Daily Papers TIER_1 English(EN) ·

    当较低的权限就足够时:研究LLM代理中过度特权工具的选择

    LLM agents frequently select higher-privilege tools unnecessarily, and while safety alignment doesn't ensure least-privilege choices, a post-training defense can reduce excessive privilege use without sacrificing performance.

  108. Hugging Face Daily Papers TIER_1 English(EN) ·

    面向Agentic和多模态大模型的上下文感知强化学习

    ContextRL enhances long-horizon reasoning and multimodal performance through reinforcement learning that rewards context selection for supporting query-answer pairs, achieving improvements over standard methods on diverse benchmarks.

  109. Hugging Face Daily Papers TIER_1 English(EN) ·

    CoffeeBench:在异构多智能体经济中对长时域 LLM 智能体进行基准测试

    CoffeeBench evaluates LLM agents in a multi-agent economic simulation where firms interact over 90 days to maximize profits, revealing differences in communication patterns and performance among various models.

  110. arXiv cs.CV TIER_1 English(EN) · Nuo Chen, Lulin Liu, Zihao Li, Ziyao Zeng, Zihao Zhu, Wenyan Cong, Junyuan Hong, Yunhao Yang, Zhengzhong Tu, Yan Wang, Boris Ivanovic, Marco Pavone, Zhangyang Wang, Yang Zhou, Zhiwen Fan ·

    用于世界模型中多智能体动力学的物理学基准测试

    arXiv:2606.28757v1 Announce Type: new Abstract: Generative world models hold immense promise as scalable simulators for autonomous systems, particularly for synthesizing rare but safety-critical multi-agent interactions, such as vehicle collisions. However, current evaluation par…

  111. Replit blog TIER_1 English(EN) ·

    闭环:大规模评估和改进 Replit Agent

    Most Replit Agent users start with nothing more than an idea. They describe the goal in natural language — without a repo, test suite, or chosen framework — and expect the agent to turn it into a functioning app. The result might be a website, slide deck, mobile app, several conn…

  112. X — Nathan Lambert (Interconnects) TIER_1 English(EN) · natolambert ·

    TMax:面向终端代理的开源强化学习方法

    TMax: An open RL recipe for terminal agents I’m very excited to get to share a new RL paper today that I got to have a small part in – a type of paper I suspect we’ll see much more of in the future. The key is that RL research is very different today, in mid-2026, than what most

  113. Towards AI TIER_1 English(EN) · Sachinbenchihalli ·

    Reflection Agent Architecture:通过工具驱动的迭代消除 LLM 幻觉…

    <h3>Reflection Agent Architecture: Eliminating LLM Hallucinations via Tool-Grounded Iterative Self-Verification</h3><p>A Technical Design Paper Covering Prompt Architecture, Multi-Agent Design, and LangGraph Integration</p><blockquote><em>Large language models (LLMs) are increasi…

  114. dev.to — MCP tag TIER_1 English(EN) · Ruben ·

    我如何构建了一个免费的事件溯源世界模型,以防止多个LLM代理破坏共享状态

    <h2> The problem </h2> <p>I kept hitting the same wall building multi-agent systems: LLMs that write directly to shared state corrupt it. They hallucinate field values, conflict with each other, produce structurally invalid data. The more agents you add, the worse it gets — corru…

  115. dev.to — MCP tag TIER_1 English(EN) · Diogo Santos ·

    The Weaver Stack:安全LLM代理的单一合约层

    <div class="crayons-card c-embed text-styles text-styles--secondary"> <div class="c-embed__content"> <div class="c-embed__body flex items-center justify-between"> <a class="c-link fw-bold flex items-center" href="https://pub.towardsai.net/the-weaver-stack-one-contract-layer-for-s…

  116. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    技能更多,代理更差?技能库扩展时技能模仿会降低性能

    "More Skills, Worse Agents? Skill Shadowing Degrades Performance When Expanding Skill Libraries" Skill libraries allow LLM agents to load task-specific instructions on demand, letting non-expert users solve domain-specific tasks through natural language without knowing which skil…

  117. Towards AI TIER_1 English(EN) · Divy Yadav ·

    构建长期运行的 Claude 托管代理:为何状态比计算更重要

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*dFGAIvcuYh47KDmw_2WpQQ.png" /><figcaption>Photo from AI</figcaption></figure><h4>A build story with real code, real failures, and the specific reasons one sandbox provider fixed problems I didn’t know I had.</h4>…

  118. Medium — MCP tag TIER_1 English(EN) · Diogo Santos ·

    The Weaver Stack:一个用于安全 LLM Agent 的合约层

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/the-weaver-stack-one-contract-layer-for-safe-llm-agents-7f733cad5eac?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1500/0*OZushp79ixGE9vEL.png" width="1500" /></a>…

  119. HN — AI startup stories TIER_1 English(EN) · dhorthy ·

    12要素代理:可靠的LLM应用模式

  120. dev.to — LLM tag TIER_1 English(EN) · Gabriel Anhaia ·

    多智能体协调:保持智能体理智的消息总线模式

    <ul> <li> <strong>Book:</strong> <a href="https://www.amazon.com/dp/B0GX35XTG6" rel="noopener noreferrer">Agents in Production — Building, Tracing, and Shipping Multi-Step AI You Can Trust</a> </li> <li> <strong>Also by me:</strong> <a href="https://www.amazon.de/-/en/dp/B0GXNNMK…

  121. dev.to — LLM tag TIER_1 English(EN) · Guillermo Fernandez ·

    超越提示词的谨慎和护栏:运行时强制执行的预行动认知,以实现值得信赖的 LLM 代理

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faoxty8tgp9e88wzuoozc.png"><img alt=" " height="1000"…

  122. dev.to — LLM tag TIER_1 English(EN) · Sébastien Conejo ·

    LLM智能体(agents)的可靠性栈:工具与方法

    <p>A request can fail at three moments: before you send it, while it runs, or after it returns. Different tools and habits cover different moments. This is a directory grouped by what each one does.</p> <h1> Methods you apply yourself </h1> <p>You apply these for free, and they r…

  123. r/MachineLearning TIER_1 English(EN) · /u/vagobond45 ·

    系统级方法应对提示注入:在 LLM 代理中分离指令和数据通道 [P]

    <!-- SC_OFF --><div class="md"><p>Prompt injection has emerged as one of the most persistent failure modes in tool-using LLM systems, particularly in agentic workflows where models interact with external data sources.</p> <p>Most mitigation strategies focus on input filtering or …

  124. dev.to — LLM tag TIER_1 English(EN) · Jonah T ·

    Context Warp Drive:长时运行LLM代理的确定性折叠

    <p>Context Warp Drive is an open-source TypeScript library for keeping long-running LLM agents under the context ceiling without asking another model to summarize their state.</p> <p>The core trick is deterministic folding. Instead of summarization calls, it compacts old transcri…

  125. dev.to — LLM tag TIER_1 English(EN) · 최해일 ·

    The Loadout Pattern:将方向盘交给自主LLM

    <h1> The Loadout Pattern: Handing the Wheel to an Autonomous LLM </h1> <h2> The core idea </h2> <p>Conventional automation <strong>executes</strong> a procedure — code runs a fixed sequence of steps and decides<br /> nothing; same input, same path, every time. The loadout pattern…

  126. r/MachineLearning TIER_1 English(EN) · /u/ThirdWaveCat ·

    将代理工作流编译到LLM权重:以成本降低两个数量级实现近乎前沿的质量

    <table> <tr><td> <a href="https://www.reddit.com/r/MachineLearning/comments/1ufgpnh/r_compiling_agentic_workflows_into_llm_weights/"> <img alt="[R] Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost" src="https://external-prev…

  127. r/LocalLLaMA TIER_1 English(EN) · /u/pmttyji ·

    TMax:终端代理的简单方法

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1uco0aa/tmax_a_simple_recipe_for_terminal_agents/"> <img alt="TMax: A Simple Recipe for Terminal Agents" src="https://preview.redd.it/u8v8ya27su8h1.png?width=140&amp;height=68&amp;auto=webp&amp;s=f2295f1d2a376…