PulseAugur
中
实时 00:31:57
English(EN) Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

最新研究探讨LLM智能体在技能选择、自动驾驶和合规性方面的进展

arXiv上发布的多篇研究论文探讨了大型语言模型(LLM)智能体的进展,重点在于提高其能力和可靠性。其中一篇论文介绍了用于LLM智能体最优技能选择的最佳前缀选择(BPS),该方法在性能和代币成本方面提供了可证明的保证。另一项研究提出了一个混合框架用于自动驾驶,该框架整合了LLM的常识推理与强化学习和PID控制,以增强决策能力。此外,还有研究通过纵向生命轨迹来缓解LLM智能体中的身份本质主义,并开发了LLM智能体的策略合规性和故障归因方法。 AI

影响 这些进展旨在提高LLM智能体在自动驾驶和金融合规等不同领域的性能、可靠性和适用性。

排序理由 arXiv上发表了多篇研究论文,详细介绍了LLM智能体的新方法和基准测试。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 492 个来源。 我们如何撰写摘要 →

最新研究探讨LLM智能体在技能选择、自动驾驶和合规性方面的进展

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
arXiv上发表了多篇研究论文,详细介绍了LLM智能体的新方法和基准测试。
Source corroboration
492 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
92 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+218 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [492]

  1. arXiv cs.AI TIER_1 English(EN) · Yining She, Lei Lin ·

    生产环境中的高效基准测试:一项关于演进式 LLM Agent 的研究

    arXiv:2609.21267v1 Announce Type: new Abstract: Production LLM agents are evaluated repeatedly as they evolve, but full agent benchmarks are costly to rerun. We study efficient recurring evaluation for a production analytics agent serving tens of thousands of monthly active users…

  2. arXiv cs.AI TIER_1 English(EN) · Laxmipriya Ganesh Iyer ·

    LLM智能体中的闭世界消解对抗工具幻觉

    arXiv:2609.19425v1 Announce Type: new Abstract: Tool-augmented large language model (LLM) agents fail in a way no tool-selection or tool-security method addresses: they call tools that do not exist and pass arguments no schema declares. Existing defenses either pick the right too…

  3. arXiv cs.AI TIER_1 English(EN) · Zehua Zhang, Ati Priya Bajaj, Divij Handa, Siyu Liu, Arvind S Raj, Hongkai Chen, Hulin Wang, Yibo Liu, Zion Leonahenahe Basque, Souradip Nath, Vishal Juneja, Nikhil Chapre, Tiffany Bao, Yan Shoshitaishvili, Adam Doup\'e, Chitta Baral, Ruoyu Wang ·

    BuildBench:在编译真实世界开源软件方面对 LLM Agents 进行基准测试

    arXiv:2509.25248v2 Announce Type: replace-cross Abstract: Automatically compiling open-source software (OSS) projects is a vital, labor-intensive, and complex task, which makes it a good challenge for LLM Agents. Existing methods rely on manually curated rules and workflows, whic…

  4. arXiv cs.AI TIER_1 English(EN) · Run Peng, Ziqiao Ma, Amy Pang, Sikai Li, Zhang Xi-Jia, Yingzhuo Yu, Cristian-Paul Bara, Joyce Chai ·

    LLM智能体在信息不对称下的协作中的沟通与验证

    arXiv:2510.25595v2 Announce Type: replace-cross Abstract: While Large Language Model (LLM) agents are often approached from the angle of action planning/generation to accomplish a goal (e.g., given by language descriptions), their abilities to collaborate with each other to achie…

  5. arXiv cs.CL TIER_1 English(EN) · Mikhail Menschikov, Matvey Iskornev, Alexander Kharitonov, Alina Bogdanova, Mikhail Belkin, Ekaterina Lisitsyna, Artyom Sosedka, Victoria Dochkina, Ruslan Kostoev, Ilia Perepechkin, Evgeny Burnaev ·

    PersonalAI 2.0:通过规划机制增强个性化LLM代理的知识图谱遍历/检索

    arXiv:2605.13481v2 Announce Type: replace Abstract: We introduce PersonalAI 2.0 (PAI-2), a novel framework designed to enhance LLM-based systems through integration of external knowledge graphs (KGs). The proposed approach addresses key limitations of existing Graph Retrieval-Aug…

  6. arXiv cs.AI TIER_1 English(EN) · Abdelghny Orogat, Ana Rostam, Essam Mansour ·

    多智能体LLM性能受架构设计而非仅模型智能的制约

    arXiv:2602.03128v2 Announce Type: replace Abstract: Multi-agent LLM frameworks are data-intensive systems that govern how agents orchestrate tasks, manage state, and coordinate decisions. These architectural choices control execution overhead, memory behavior, planning effectiven…

  7. arXiv cs.AI TIER_1 English(EN) · Yukun Zhang, Kemu Xu, Yishen Chen ·

    智能体(Agent)如何创造价值?有状态大语言模型(LLM)智能体的规划信息与释放控制

    arXiv:2609.20474v1 Announce Type: new Abstract: Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airline pilot in $\tau^2$-bench. The p…

  8. arXiv cs.AI TIER_1 English(EN) · Ruiyang Wang, Hao-Lun Hsu, Swarajh Mehta, Jiwoo Kim, Zhihao Dou, Miroslav Pajic ·

    GAVEL:用于可验证和高效长时LLM任务规划的图世界模型

    arXiv:2609.19315v1 Announce Type: cross Abstract: Large language models (LLMs) provide a flexible interface for long-horizon robot planning, but generated plans often fail to respect embodiment constraints, recover from planning errors, or reason effectively under partial observa…

  9. arXiv cs.AI TIER_1 English(EN) · Haya Halimeh, Sascha Kaltenpoth, Kevin B\"osch, Oliver M\"uller ·

    LLM驱动的GUI代理中的助推易感性的双过程视角

    arXiv:2609.19843v1 Announce Type: new Abstract: LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind. These graphical user interfaces were designed to support, but also deliberately steer, the behaviour and d…

  10. arXiv cs.AI TIER_1 English(EN) · Nolan Smyth, Yorguin-Jose Mantilla-Ramos, Pascal Jr Tikeng Notsawo, Saskia Helbling, Alberto Tosato, Mohamed Amine Merzouk, Nouha Dziri, Gauthier Gidel, Tommaso Tosato ·

    量化前沿大型语言模型代理的过度声称倾向

    arXiv:2609.20812v1 Announce Type: cross Abstract: Frontier coding agents are increasingly trusted to work autonomously for long periods, yet an agent's final response is often the only account of that work a user sees. We quantify the propensity of frontier agents to \emph{overcl…

  11. arXiv cs.AI TIER_1 English(EN) · Tisha Chawla, Susheem Koul ·

    Chronicle:LLM代理回归测试的切点重放

    arXiv:2609.20625v1 Announce Type: cross Abstract: Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing state, and on a multi-step traject…

  12. Hugging Face Daily Papers TIER_1 English(EN) ·

    Chronicle:LLM代理回归测试的切点重放

    Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing state, and on a multi-step trajectory that a re-run rarely repeats. Record-and-repla…

  13. Hugging Face Daily Papers TIER_1 English(EN) ·

    LLM驱动的GUI代理中的助推易感性的双过程视角

    LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind. These graphical user interfaces were designed to support, but also deliberately steer, the behaviour and decisions of users. While behavioural biases in t…

  14. arXiv cs.AI TIER_1 English(EN) · Tanmay Sah, Dolly Sah, Harshul Jain, Tanya Sah ·

    EvoUndo:LLM 代理工具的约束可恢复性自演化

    arXiv:2608.28363v2 Announce Type: replace Abstract: LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a successful mutation may leave persistent effects that cannot b…

  15. arXiv cs.AI TIER_1 English(EN) · Sehee Kim, Yumin Choi, Minki Kang, Sung Ju Hwang ·

    EvolveTrade:面向自进化LLM交易代理的体验驱动策略优化

    arXiv:2609.17632v1 Announce Type: new Abstract: Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment. This limits their ab…

  16. arXiv cs.AI TIER_1 English(EN) · Li Chen ·

    AutoTuneBench:LLM服务引擎的代理自动调优的可信赖测量

    arXiv:2609.18123v1 Announce Type: new Abstract: Large language model agents tune GPU kernels and serving engines through a closed loop of propose, measure, and keep, but the measurements behind this loop are not trustworthy. We characterize four failure modes from a four-day pilo…

  17. arXiv cs.AI TIER_1 English(EN) · Yifeng Xiao, Pierluigi Nuzzo ·

    使用合约对 LLM 代理进行符号化时间监督

    arXiv:2609.18128v1 Announce Type: new Abstract: Large language model (LLM) agents augmented by tools can automate complex, multi-step tasks, such as web navigation, code generation, and workflow orchestration, by acting on external systems through tool calls. However, hallucinati…

  18. arXiv cs.AI TIER_1 English(EN) · Xiangfan Wu, Zonghao Ying, Huiyu Wu, Xing Zheng, Huangsheng Cheng, Xiaorong Shi, Jing Guo ·

    LLM 智能体系统中集体失控:变异、传染与恢复的流行病学叙事

    arXiv:2609.18460v1 Announce Type: new Abstract: How does a multi-agent system evolve from a local deviation into collective loss of control? We propose an epidemic explanation organized around accidental mutation, contagion, and recovery. A spontaneous deviation creates a seed; c…

  19. arXiv cs.CL TIER_1 English(EN) · Yi Yu, Liuyi Yao, Yaliang Li, Enshu Wang, Libing Wu ·

    回滚世界,保留反思:长时域 LLM Agent 的回滚诱导反思

    arXiv:2609.18304v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly tackle long-horizon tasks through multi-step environment interaction, yet a single erroneous action can alter subsequent states and observations, causing errors to compound over time. E…

  20. arXiv cs.AI TIER_1 English(EN) · Jiaxuan Jiang, Liyuan He, Zhixuan Fang ·

    CERA-MoA:与持续学习的LLM代理共同演进路由机制

    arXiv:2609.18779v1 Announce Type: new Abstract: Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, limiting their ability to respond to evolving agent capabilities. This disconnect prevents routing strategies from …

  21. Hugging Face Daily Papers TIER_1 English(EN) ·

    CERA-MoA:与持续学习的LLM代理共同演进路由机制

    Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, limiting their ability to respond to evolving agent capabilities. This disconnect prevents routing strategies from adapting to evolving agent capabilities during p…

  22. Hugging Face Daily Papers TIER_1 English(EN) ·

    使用合约对LLM智能体进行符号化时间监督

    Large language model (LLM) agents augmented by tools can automate complex, multi-step tasks, such as web navigation, code generation, and workflow orchestration, by acting on external systems through tool calls. However, hallucinations, distributional instability, and adversarial…

  23. arXiv cs.LG TIER_1 English(EN) · Sourish Dey, Aditya Kumar ·

    为智能体式假设推理提炼基础模型:混合LLM+SLM架构中的成本、延迟和治理

    arXiv:2609.16091v1 Announce Type: new Abstract: Tabular foundation models deliver strong zero-training predictive performance via in-context learning, but their high inference latency makes them impractical as hot-path decision backends in interactive agentic loops. We distill a …

  24. arXiv cs.AI TIER_1 English(EN) · Zhen Li, Jun Cai, Haoran Gao, An Li, Tan Li ·

    面向Agentic AI服务的边缘LLM推理的端到端低延迟和负载均衡请求调度

    arXiv:2609.17193v1 Announce Type: new Abstract: Large language model (LLM)-powered agentic AI services increasingly demand low-latency inference, motivating the deployment of LLMs across distributed edge servers. However, heterogeneous communication and computing capabilities, to…

  25. arXiv cs.AI TIER_1 English(EN) · Xinyuan Song, Zekun Cai ·

    世界模型科学:长视界LLM智能体的自组织临界性、弱混沌和亚稳态信念动力学

    arXiv:2609.17419v1 Announce Type: new Abstract: Long-horizon LLM agents must maintain task state across extended sequences of observations, actions, tool calls, and intermediate beliefs. We study these trajectories through three dynamical views: self-organized criticality, weak c…

  26. arXiv cs.AI TIER_1 English(EN) · Xiaoyan Li, Yunli Wang ·

    面向集成工具的大语言模型Agent的通用防御对抗性攻击

    arXiv:2609.16098v1 Announce Type: cross Abstract: Large Language Model (LLM) agents have demonstrated impressive capabilities across a variety of domains, particularly when integrated with external tools for multi-step task completion. However, they are increasingly vulnerable to…

  27. arXiv cs.AI TIER_1 English(EN) · Jiayue Gaveal Fan, Arul Murugan, Shreyas Krishnan, Abhishek Nagaraj ·

    解读和引导LLM智能体进行社会模拟

    arXiv:2609.16436v1 Announce Type: cross Abstract: Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit. However, LLMs are ultimately black boxes based on deep…

  28. Hugging Face Daily Papers TIER_1 English(EN) ·

    CERA-MoA:与持续学习的LLM代理共同演进路由机制

    Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, limiting their ability to respond to evolving agent capabilities. This disconnect prevents routing strategies from adapting to evolving agent capabilities during p…

  29. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Glaucia Melo ·

    ToMAS:一个基于多智能体LLM失败的试点性失败-导向的心理理论基准

    LLM-based multi-agent systems can fail even when communication succeeds because agents do not correctly track their peers' roles, knowledge, or intentions. We investigate whether such inter-agent misalignment cases, labelled FC2 in MAST-Data, can be converted into functional part…

  30. arXiv cs.AI TIER_1 English(EN) · Jianhua Jiang, Dongbo Yuan, Weihua Li ·

    MemRiskBench:长远景LLM智能体的追踪感知风险保留评估

    arXiv:2609.14976v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate memory across sessions, creating sparse but high-impact risks: stale facts, conflicting updates, cross-user leakage, revoked-memory reuse, and constraint decay. Standard aggregate scores hide per-r…

  31. arXiv cs.AI TIER_1 English(EN) · Burak Agachan, Max van Duijn, Amirhossein Zohrehvand ·

    LLM 智能体团队中的回溯式权威:关于扁平化与层级化协调的配对实验

    arXiv:2609.14767v1 Announce Type: cross Abstract: Hierarchical orchestration, in which a Manager agent reviews worker output and can send it back for revision, is the default coordination pattern in production multi-agent LLM frameworks. Classical organizational theory predicts t…

  32. arXiv cs.AI TIER_1 English(EN) · Bingzheng Wang, Xiaoyan Gu, Wentao Wang, Xingyou Yang, Hongcheng Li, Rong Yin ·

    ActGuard:针对大语言模型代理中非直接提示注入的预执行操作审计

    arXiv:2609.14987v1 Announce Type: cross Abstract: Large language model (LLM) agents interact with external environments through tool invocation, but tool outputs can also expose them to indirect prompt injection (IPI) attacks. Existing defenses mainly rely on prompt hardening, co…

  33. arXiv cs.AI TIER_1 English(EN) · Mirza Samad Ahmed Baig, Syeda Anshrah Gillani, Asher Ali, Muhammad Hamzah Siddiqui ·

    随机副官:工具使用LLM代理的结构化租户隔离

    arXiv:2609.14780v1 Announce Type: cross Abstract: Multi-tenant tools commonly accept a tenant identifier and validate it against the caller's entitlement. For a large language model (LLM) agent, that pattern delegates resource selection to a process whose context may contain atta…

  34. arXiv cs.AI TIER_1 English(EN) · Zhenyu Zhang1, Jiudong Yang ·

    MOSCOPT:LLM智能体的混合技能集体优化

    arXiv:2609.14399v1 Announce Type: new Abstract: Natural language prompts and skills serve as the strategic backbone of LLM-based agents. Recent advances in prompt and skill optimization have achieved notable gains, yet all existing methods optimize a \emph{single} text template--…

  35. arXiv cs.AI TIER_1 English(EN) · Kazi Abrar Mahmud, Nilotpaul Kundu Dhurubo, Tamal Kirttonia, Sabbir Hossain Ujjal, Mohammad Ariful Haque ·

    连接思想与行动:使用增强型MetaTool的ROS框架驯服开源LLM智能体中的长时程不稳定性

    arXiv:2609.13335v1 Announce Type: cross Abstract: Large Language Models (LLMs) have enabled more natural human-robot interaction, but open-source models often exhibit unstable long-horizon reasoning and inefficient action execution when deployed in agentic robotic frameworks. Thi…

  36. arXiv cs.AI TIER_1 English(EN) · Ala N. Tak, Teruhisa Misu, Kumar Akash, Zhaobo K. Zheng, Kevin H. Joo, Jonathan Gratch ·

    从过程损失到组装奖励:多智能体LLM协作的人类基础诊断

    arXiv:2609.13261v1 Announce Type: cross Abstract: LLM agents are increasingly used for collaborative problem solving and human-group simulation. This makes outcome-only evaluation insufficient: if LLM groups are used as models of human groups, we need to know whether they succeed…

  37. arXiv cs.LG TIER_1 English(EN) · Yue Zhao ·

    GRADE: LLM 代理依赖和执行的图表示

    arXiv:2606.22741v2 Announce Type: replace Abstract: A trace records what an LLM agent did at each step. What is gained by also recording what each step relied on? GRADE represents a run as one typed graph: execution edges come free from the trace, and dependency edges are supplie…

  38. arXiv cs.CL TIER_1 English(EN) · Kaiyuan Liu, Qiuyang Mang, Bo Peng, Wenhao Chai, Hanchen Li, Shreyas Pimpalgaonkar, Luke Zettlemoyer, Alex Dimakis, Alvin Cheung ·

    当智能体变慢时:通过 Elo-per-token 分析理解 LLM 智能体的测试时策略

    arXiv:2609.15309v1 Announce Type: new Abstract: Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to stop. This test-time strategy makes it difficult to measure how agent performance …

  39. arXiv cs.AI TIER_1 English(EN) · Guangsheng Yu, Yanna Jiang, Qin Wang, Baihe Ma, Xu Wang ·

    K-Bench:用于 Agentic 部署中 LLM 遗忘的基准测试

    arXiv:2609.12808v2 Announce Type: replace Abstract: Unlearning benchmarks such as TOFU and MUSE certify forgetting by reading the model's final answer, where a model that refuses to answer already counts as having forgotten. We show that this model-level certificate does not tran…

  40. arXiv cs.AI TIER_1 English(EN) · Boyan Liu, Gongming Zhao, Hongli Xu ·

    面向高效LLM工具使用的效用引导代理编排

    arXiv:2603.19896v2 Announce Type: replace Abstract: Tool-using large language model (LLM) agents often face a fundamental tension between answer quality and execution cost. Fixed workflows are stable but inflexible, while free-form multi-step reasoning methods such as ReAct may i…

  41. arXiv cs.AI TIER_1 English(EN) · Zixiang Liu, Wenrui Liu, Elsie Dai, Wenhan Yu, Lei Yu, Tong Yang, Jinjun Han, Hong Gao ·

    MCPAgentBench:评估LLM Agent MCP工具使用的真实任务基准

    arXiv:2512.24565v4 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of external tools via the Model Context Protocol (MCP) is considered a future trend. Current MCP evaluation sets suffer from issue…

  42. Hugging Face Daily Papers TIER_1 English(EN) ·

    EvolveTrade:面向自进化LLM交易代理的体验驱动策略优化

    Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment. This limits their ability to adapt how they gather evidence, invoke …

  43. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Babak Heydari ·

    廉价对话稳定LLM智能体的战略互动

    Large language models are increasingly deployed as interacting agents, making the persistence of their action policies across repeated interaction critical for reliable multi-agent operation. We investigate whether and how agent-generated, non-binding pre-play communication ("che…

  44. arXiv cs.LG TIER_1 English(EN) · Asaad Althoubi ·

    三思而后行:LLM智能体(Agent)的预行动验证

    arXiv:2609.11957v1 Announce Type: new Abstract: An LLM agent acts on the world by emitting actions: shell commands to run, edits to apply. A wrong action does not always fail loudly; it can fail silently, producing a plausible but incorrect effect that raises no error. We argue t…

  45. arXiv cs.CL TIER_1 English(EN) · Yuli Qiu, Yutong Li, Wei Su, Zeming Liu, Wanxiang Che, Heyan Huang, Haifeng Wang, Yuang Guo ·

    LifeMem:赋能大语言模型智能体实现终身体验复用

    arXiv:2609.12655v1 Announce Type: new Abstract: Large language model agents are expected to continuously adapt to new tasks and environments over their lifetime by reusing past experience. However, existing memory-based agents struggle to transfer reusable experience across envir…

  46. Hugging Face Daily Papers TIER_1 English(EN) ·

    当Agent变慢时:通过Elo-per-token分析理解LLM Agent的测试时策略

    Elo-per-token analysis reveals that LLM agents initially scale faster than independent sampling but eventually slow, while parallel short sessions improve performance over single long runs.

  47. Hugging Face Daily Papers TIER_1 English(EN) ·

    LLM Agent 团队中的循环回溯权威:关于扁平化和层级化协调的配对实验

    Hierarchical orchestration, in which a Manager agent reviews worker output and can send it back for revision, is the default coordination pattern in production multi-agent LLM frameworks. Classical organizational theory predicts that the authority link speeds convergence on decis…

  48. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Amirhossein Zohrehvand ·

    LLM 智能体团队中的回环权威:关于扁平化与层级化协调的配对实验

    Hierarchical orchestration, in which a Manager agent reviews worker output and can send it back for revision, is the default coordination pattern in production multi-agent LLM frameworks. Classical organizational theory predicts that the authority link speeds convergence on decis…

  49. arXiv cs.AI TIER_1 English(EN) · Ruiqing Yue, Yu Cui, Zhuoyu Sun, Sicheng Pan, Xianhong Xue, Tingyu Li, Ting Li, Wenzhuo Zhu, Yi Chen, Yifei Liu, Baohan Huang, Zhe Cui, Haibin Zhang, Cong Zuo ·

    Ecdysis:为大型语言模型代理进行运行时工具的高效有效训练

    arXiv:2609.11677v1 Announce Type: cross Abstract: Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing harness evolution methods typically rely on …

  50. arXiv cs.AI TIER_1 English(EN) · Shenghan Zheng, Zonglin Di, Yimin Liu, Kyoung Whan Choe, Jiankai Sun, Heguang Lin, Penghao Jiang, Yifeng He, Xiao Cheng, Jicheng Wang, Wenbo Chen, Alex Yates, Yinzhe Zhao, Bingran You, Yuan Gao, Ayush Munot, Shubham Gaur, Zhe Ye, Hao Wang, Xiangyi Li, Da… ·

    BenchShield:用于 LLM 代理评估基础设施中奖励完整性的形式化模型支持的插桩

    arXiv:2609.11028v1 Announce Type: cross Abstract: LM-agent benchmarks increasingly function as interactive evaluation infrastructure. Agents observe state, call tools, modify workspaces, submit artifacts, and receive rewards from outcome procedures. This interactivity makes evalu…

  51. arXiv cs.AI TIER_1 English(EN) · Bochao Feng, Jianjiang Li, Haojie Wang, Lin Qiao, Yinghui Li, Yukun Yan, Jidong Zhai ·

    为面向尾部调度的代理式LLM工作流解耦就绪度与发布

    arXiv:2609.10964v1 Announce Type: new Abstract: Agentic LLM workflows consist of sequences of model turns interleaved with tool interactions, so their end-to-end completion time depends not only on inference speed but also on when ready turns are released. Most runtimes release e…

  52. arXiv cs.LG TIER_1 English(EN) · Asif Pinjari, Mithun Paul Saint-Germain ·

    DriftNet:一种用于检测和定位 LLM Agent 中提示注入的双头轨迹 Transformer

    arXiv:2609.10892v1 Announce Type: cross Abstract: When an indirect prompt injection succeeds against an LLM agent, the compromise is visible in the agent's own behavior: a benign prefix of tool calls, a poisoned observation, and a suffix of actions that serve the attacker. An ope…

  53. arXiv cs.AI TIER_1 English(EN) · Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary ·

    Eon Benchmark 时代:生成式企业地产,为 LLM Agent 评测提供精确地面真实数据

    arXiv:2609.09853v1 Announce Type: new Abstract: LLM agents for enterprise systems of record cannot be evaluated on customer production data, and no existing substitute provides ground truth. We present the Era by Eon Benchmark for evaluating LLM agents that use enterprise tools. …

  54. Hugging Face Daily Papers TIER_1 English(EN) ·

    Ecdysis:为大型语言模型代理进行运行时工具的高效有效训练

    Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing harness evolution methods typically rely on iterative search, repeatedly evaluating and revisi…

  55. arXiv cs.AI TIER_1 English(EN) · Asif Pinjari, Mithun Paul Saint-Germain ·

    AgentDrift:注入劫持的LLM代理轨迹的步长标记基准

    arXiv:2609.06972v1 Announce Type: cross Abstract: LLM agents complete tasks by issuing sequences of tool calls, and every observation they read is a channel through which an indirect prompt injection can enter. A successful injection has a characteristic shape when the trajectory…

  56. arXiv cs.CL TIER_1 English(EN) · Shiqi Chen, Tongyao Zhu, Zian Wang, Jinghan Zhang, Kangrui Wang, Ruochen Zhou, Siyang Gao, Teng Xiao, Yee Whye Teh, Junxian He, Manling Li ·

    大型语言模型(LLM)智能体为何在新环境中探索失败?世界模型视角

    arXiv:2510.15047v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) as agents often fail to improve in new environments. We identify and characterize a failure mode we call exploration collapse: under reinforcement learning (RL) in environments whose states are…

  57. arXiv cs.AI TIER_1 English(EN) · Jingyu Zhang, Tianjian Li, William Jurayj, Hongyuan Zhan, Benjamin Van Durme, Daniel Khashabi ·

    LLM智能体中的多层指令层级

    arXiv:2604.09443v4 Announce Type: replace-cross Abstract: Large language model agents receive instructions from many sources-system messages, user prompts, tool outputs, other agents, and more-each carrying different levels of trust and authority. When these instructions conflict…

  58. arXiv cs.AI TIER_1 English(EN) · Junzhuo Ma, Chenghuang Shen, Yi Yu, Xingyan Liu, Jing Gu, Hangyi Sun, Guangquan Hu, Jianfeng Liu, Weiting Liu, Pu Mingyue, Wang Yu, Zhengdong Xiao, Rui Xie, Longjiu Luo, Qianrong Wang, Gurong Cui, Honglin Qiao, Wenlian Lu ·

    使用潜在逻辑增强、鲁棒噪声抑制和混合奖励建模来适应技术服务 LLM 代理

    arXiv:2603.18074v2 Announce Type: replace-cross Abstract: Technical-service LLM agents are entering production workflows, where value depends on whether engineers adopt generated replies. Service tickets hide decision logic, contain noisy single-reference responses, and make rewa…

  59. arXiv cs.AI TIER_1 English(EN) · Boyang Wang, Yunhan Wang, Yalun Wu ·

    不可靠的进度条:LLM代理能否在执行过程中可靠地报告任务进度?

    arXiv:2609.08589v1 Announce Type: cross Abstract: Recent large language models can emit task-progress signals that agent frameworks use to decide whether a task should continue or stop, yet whether a model can reliably report its task progress at every stage of a task, and where …

  60. arXiv cs.AI TIER_1 English(EN) · Ruoxi Shang, Christina-Maria Androna, Orfeas Menis Mastromichalakis, Yu Feng, Aniruddhan Ramesh, Rico Angell, Shang Hong Sim, Chrysoula Zerva, Emmanouil Koukoumidis ·

    AURA-Eval:LLM Agent轨迹中风险意识下的行动评估框架

    arXiv:2609.06783v1 Announce Type: cross Abstract: LLM agents operate in workflows where unsafe actions can have real consequences. Existing safety evaluations often reduce behavior to a single score, obscuring risk recognition, pre-action detection, and safe task completion when …

  61. arXiv cs.AI TIER_1 English(EN) · Md Jafrin Hossain, Nur Al Hasan Haldar ·

    结构相似,时间跨度遥远:衡量长时域 LLM Agent 的安全暴露

    arXiv:2609.05911v1 Announce Type: cross Abstract: Long-horizon LLM agents interact with untrusted content, persistent memory, external state, and sensitive tools. Existing analyses often characterize attacks by the number of execution steps between malicious input and a downstrea…

  62. arXiv cs.AI TIER_1 English(EN) · Yuxing Lu, Yicheng Chen, Shanchan Wu, Sercan \"{O}. Ar{\i}k ·

    过程图:LLM智能体的自演化执行结构

    arXiv:2609.09153v1 Announce Type: new Abstract: Large language models are increasingly deployed as agents that plan over long horizons and act through external tools. Most agents select actions through unconstrained generation over an accumulating history, leaving implicit the pr…

  63. arXiv cs.AI TIER_1 English(EN) · Yuan Gao, Sebastian M\"uller, Mattia Piccinini, Marc Kaufeld, Yuchen Zhang, Finn Rasmus Sch\"afer, Qunying Song, Johannes Betz ·

    PlannerForge:用于自动驾驶运动规划器基于场景的测试的 LLM Agent

    arXiv:2609.08965v1 Announce Type: new Abstract: Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used to validate Autonomous Driving Systems (ADSs), but it remains a fragmented modular pipeline in which scenario gen…

  64. arXiv cs.AI TIER_1 English(EN) · Shengtian Yang, Ziyu Xiong, Yu Li, Yewen Li, Qingpeng Cai, Lei Feng ·

    SRPO:多智能体大语言模型的集合式相对策略优化

    arXiv:2609.08452v1 Announce Type: new Abstract: Multi-agent large language models solve complex tasks by coordinating several policies in a shared environment. However, existing reinforcement learning methods usually optimize each response or trajectory separately, even when seve…

  65. arXiv cs.AI TIER_1 Dansk(DA) · Shuo Ren, Xiaomian Kang, Jiajun Zhang ·

    SkillAlign:为基于LLM的代理对齐技能接口

    arXiv:2609.07255v1 Announce Type: new Abstract: Language-model agents increasingly rely on skills: reusable procedural knowledge for reasoning, tool use, and interaction. Existing work studies how skills are acquired, retrieved, compressed, or composed, but often assumes that onc…

  66. arXiv cs.AI TIER_1 English(EN) · Katherine Tieu, Dongqi Fu, Yinglong Xia, Hong Li, Hong Yan, Jingrui He ·

    面向多智能体LLM工作流的推理时图工程

    arXiv:2609.05774v1 Announce Type: new Abstract: Recent multi-agent LLM systems increasingly rely on graph-structured communication to coordinate specialized agents. We revisit multi-agent orchestration from a graph-engineering perspective: rather than optimizing a static topology…

  67. arXiv cs.AI TIER_1 English(EN) · Cen Mia Zhao, Haibo Ruan, Wenjie Chen, Pei-fen Tu, Usman Abbasi, Joel Hesch ·

    超越提示词:衡量与优化 LLM 工具-代理的驾驭能力

    arXiv:2609.05736v2 Announce Type: new Abstract: LLM tool agents can be improved without retraining by modifying the runtime harness around a fixed model: prompts, tool interfaces, middleware, state handling, and recovery logic. We study this setting as resource-bounded harness se…

  68. Hugging Face Daily Papers TIER_1 English(EN) ·

    程序化图:LLM智能体的自演化执行结构

    Large language models are increasingly deployed as agents that plan over long horizons and act through external tools. Most agents select actions through unconstrained generation over an accumulating history, leaving implicit the procedural knowledge of what to do, in what order,…

  69. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Sercan Ö. Arık ·

    程序化图:LLM智能体的自演化执行结构

    Large language models are increasingly deployed as agents that plan over long horizons and act through external tools. Most agents select actions through unconstrained generation over an accumulating history, leaving implicit the procedural knowledge of what to do, in what order,…

  70. Hugging Face Daily Papers TIER_1 English(EN) ·

    PlannerForge:用于自动驾驶运动规划器基于场景的测试的 LLM 代理

    PlannerForge is an LLM-agent framework that unifies all stages of scenario-based autonomous driving testing and improves generation, selection, modification, and planning performance across commercial and open-source models.

  71. Hugging Face Daily Papers TIER_1 English(EN) ·

    过程图:LLM智能体的自演化执行结构

    A procedural graph framework organizes agent actions into structured relational triplets, providing situational guidance and self-evolving topology to improve long-horizon tool use.

  72. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Goetz Botterweck ·

    LLM智能体的不确定性量化:分类、评估协议与实证研究

    Large language models (LLMs) are no longer deployed only for single-turn conversation but increasingly act as agents that plan, call tools, retrieve evidence, maintain memory, and interact over long horizons, often together with other agents through multi-turn conversations. Ther…

  73. arXiv cs.AI TIER_1 English(EN) · Jiazheng Sun, Boyu Yang, Binhao Yuan, Mingxuan Li, Xin Peng ·

    Trace2Tower:面向LLM智能体的过渡感知特征追踪多层技能归纳

    arXiv:2609.05261v1 Announce Type: new Abstract: Large language model agents increasingly rely on execution traces to master complex interactive tasks. However, current paradigms are bottlenecked by shallow trajectory retrieval and flat skill summarization, fundamentally ignoring …

  74. arXiv cs.AI TIER_1 English(EN) · Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou ·

    SiLR:LLM工具代理的结构保持准入和过程奖励

    arXiv:2609.04629v1 Announce Type: new Abstract: A runtime gate for an LLM tool agent is usually cast as a filter. In a ReAct loop a rejected proposal is followed by another at the same state, so the gate is a search operator over the proposal stream whose admission criterion shap…

  75. Hugging Face Daily Papers TIER_1 English(EN) ·

    PARSER:长上下文LLM代理的并行读取与深度推理

    PARSER decouples parallel chunk reading from iterative reasoning via scatter-gather subagents, improving long-context multi-hop accuracy and reducing latency.

  76. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Ruimin Ke ·

    SimTIO:一个用于组合式交通干预优化的、基于仿真的多智能体大语言模型框架

    Traffic analysts must translate diagnosed bottlenecks into executable interventions without allowing local improvements to degrade network-wide performance. This study presents SimTIO, a simulation-grounded multi-agent large language model framework for composing and selecting tr…

  77. arXiv cs.LG TIER_1 English(EN) · Jinwei Gan ·

    TIGPO:面向长时域LLM智能体的时序实例图策略优化

    arXiv:2609.03383v1 Announce Type: new Abstract: Graph-based policy optimization improves credit assignment for long-horizon LLM agents by organizing rollout trajectories into state-transition graphs. However, existing methods construct graphs independently within each policy upda…

  78. arXiv cs.AI TIER_1 English(EN) · Paul Brookes, Vardan Voskanyan, Rafail Giavrimis, Matthew Truscott, Mina Ilieva, Chrystalla Pavlou, Alexandru Staicu, Manal Adham, Will Evers- Hood, Jingzhi Gong, Kejia Zhang, Matvey Fedoseev, Vishal Sharma, Roman Bauer, Zheng Wang, Hema Nair, Wei Jie, T… ·

    精益求精:基于LLM的智能体自动化优化

    arXiv:2512.09108v2 Announce Type: replace-cross Abstract: Agentic AI systems built on large language models (LLMs) offer significant potential for automating complex workflows, from software development to customer support. However, LLM agents often underperform due to suboptimal…

  79. arXiv cs.AI TIER_1 English(EN) · Uday Vallabhaneni, Cassie L. Cagwin, David J. Wild ·

    SENTINEL-RL:在安全运营中心将拓扑推理从LLM代理中卸载

    arXiv:2609.04159v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly proposed as autonomous SOC analysts, but two limitations make them unreliable at enterprise scale: a finite context window cannot hold a multi-thousand-host authentication graph, …

  80. arXiv cs.CL TIER_1 English(EN) · Michael Nguyen, Wei Chen Tan, Nurul Aisyah Hassan, Arvind Raman, Li Hua Lim, Ahmad Faiz Razak ·

    Harness优化价值何在?自进化LLM智能体中的局部收益与预算分配陷阱

    arXiv:2609.02889v1 Announce Type: new Abstract: A growing body of work improves frozen large language models (LLMs) as agents by evolving their harness: the textual scaffolding around the model, including persona, strategy, format rules, and control heuristics. Existing reflectiv…

  81. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Wanjng Ma ·

    SimSkill:用于交通仿真中技能和知识积累的自进化LLM代理

    Cumulative culture enables humans to preserve, reuse, and extend knowledge and skills across experiences and generations. Inspired by this principle, we introduce \textit{SimSkill}, a self-evolving agent built around the Simulation of Urban MObility (SUMO) traffic simulator. SimS…

  82. arXiv cs.AI TIER_1 English(EN) · Haibo Jin, Suijin Wang, Xucheng Yu, Haojing Luo, Haohan Wang ·

    通过 Agent 原生可重用工具原语在 LLM 工具使用中进行工程化

    arXiv:2609.01736v1 Announce Type: cross Abstract: Large language models (LLMs) augmented with external tools have demonstrated remarkable capability in solving complex real-world tasks. However, existing approaches suffer from two key challenges: brittle multi-step and multi-turn…

  83. arXiv cs.AI TIER_1 English(EN) · Huaiyuan Yao, Xiaoou Liu, Charles Fleming, Tianlong Chen, Hua Wei ·

    MASkills:多智能体LLM系统的持续技能优化

    arXiv:2609.02094v1 Announce Type: new Abstract: LLM-based multi-agent systems have shown strong performance on complex tasks, yet continual improvement from interaction experience remains challenging. Existing self-reflection methods build experience memories, but memories are mo…

  84. arXiv cs.AI TIER_1 English(EN) · Jinxi Yu, Yubei Li, Eric Hanchen Jiang, Zhi Zhang, Dong Liu, Wenxiao Zhao, Levina Li, Kai-Wei Chang, Ying Nian Wu ·

    Codebook Agent:LLM多智能体系统的摊销拓扑设计

    arXiv:2609.02264v1 Announce Type: new Abstract: Adapting the communication topology of an LLM multi-agent system to each query improves both accuracy and efficiency, yet current designers treat this as conditional graph generation: a variational, autoregressive, or diffusion deco…

  85. arXiv cs.AI TIER_1 English(EN) · Yihang Chen, Yuxiang Chen, Yuxuan Huang, Meng Fang, Weilin Luo, Jun Wang ·

    双层协调反射:多智能体LLM系统的博弈论方法

    arXiv:2609.02750v1 Announce Type: new Abstract: Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through textual reflection. Despite strong empirical results, these systems lack a unified account of coordination, memo…

  86. arXiv cs.LG TIER_1 English(EN) · Yanting Yang, Can Jin, Jinman Zhao, Jiahao Wu, Yang Zhou, Zhepeng Wang, Zhendong Wang, Mu Zhou, Dimitris N. Metaxas ·

    多行动,少决策:面向长时域LLM智能体的技能引导自适应行动分块

    arXiv:2609.02042v1 Announce Type: new Abstract: Large language model (LLM) agents for long-horizon interactive tasks typically follow a ReAct-style protocol, issuing one primitive action per LLM round. While this enables frequent replanning, it is inefficient for long-horizon tas…

  87. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Ying Nian Wu ·

    Codebook Agent:LLM多智能体系统的摊销拓扑设计

    Adapting the communication topology of an LLM multi-agent system to each query improves both accuracy and efficiency, yet current designers treat this as conditional graph generation: a variational, autoregressive, or diffusion decoder searches the $N \times N$ adjacency space, a…

  88. arXiv cs.AI TIER_1 English(EN) · Elias Stengel-Eskin, Newton Sander, Carlos Bonetti, Sasha Boguraev, James Bowler, Hale Sirin, Simon Kirby ·

    GlossoGen:复杂多智能体LLM交互中的涌现语言

    arXiv:2609.01491v1 Announce Type: cross Abstract: The growing rate at which LLM agents interact with one another raises key questions about language evolution in multi-LLM-agent settings, with implications for safety and monitorability as well as for linguistic accounts of LLMs. …

  89. arXiv cs.CL TIER_1 English(EN) · Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Xinping Lei, Qingshui Gu, Yuxuan Zhang, Zexuan Wang, Chen He, Chen Huang, Maojia Song, Zhiyuan Zeng, Shaowen Wang, Jinkai Liu, Yunfeng Shi, Jiaheng Liu, Shen Yan, Wenhao Huang, Ge Zhang, Wenxuan Zhang ·

    HarnessDev:大型语言模型能否创建和演进自己的Agent Harness?

    arXiv:2609.01437v1 Announce Type: cross Abstract: As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly termed the agent harness. Changing this harness while holding model weights fixe…

  90. arXiv cs.AI TIER_1 English(EN) · Philip Drammeh ·

    多智能体LLM编排实现确定性、高质量的事件响应决策支持

    arXiv:2511.15755v3 Announce Type: replace Abstract: Large language models (LLMs) promise to accelerate incident response in production systems, yet single-agent approaches generate vague, unusable recommendations. We present MyAntFarm.ai, a reproducible containerized framework de…

  91. arXiv cs.AI TIER_1 English(EN) · Rui Yang, Junjie Xu, Zhengyu Liu, Neil Fendley, Yang Hong, Ziyang Li, Yinzhi Cao ·

    SoK:当安全代理一起失败时:多代理LLM系统的安全性

    arXiv:2609.00595v1 Announce Type: cross Abstract: Safe agents can fail together. Multi-agent LLM systems (MAS) move information, state, decisions, and authority across principal boundaries, creating failures that local checks may miss. Without an execution-level view, a multi-age…

  92. arXiv cs.AI TIER_1 English(EN) · Jun Hou, Priya Pitre, Yi Fang, Xuan Wang ·

    EDGE:多智能体LLM系统中基于错误依赖图的误差依赖多误差归因

    arXiv:2609.01360v1 Announce Type: new Abstract: Large language model (LLM) agent failures often contain multiple related errors rather than a single mistake. Existing attribution methods usually identify a responsible agent, step, or root cause, but do not explicitly model depend…

  93. arXiv cs.AI TIER_1 English(EN) · Wentao Zhang, Syed Shariyar Murtaza, Junaid Ahmad Bhatti, Utkarsh Soni, Yifan Nie, Eugene Wen, Yuntian Deng ·

    控制-数据流分离:多智能体LLM中的稳定提示优化

    arXiv:2609.00621v1 Announce Type: new Abstract: Prompt optimization can improve multi-agent LLM systems, but the prompts being optimized often serve two entangled roles: generating task-relevant content and specifying execution-critical protocols, such as message routing, output …

  94. arXiv cs.AI TIER_1 English(EN) · Shubham Kumar, Harshit Kumar, Narendra Ahuja, Saurabh Jha ·

    迈向基于信念的世界模型以支持LLM智能体

    arXiv:2609.00455v1 Announce Type: new Abstract: Large language models (LLMs) are being used as policies for autonomous decision-making and planning in many domains. Despite their strong reasoning capabilities, LLMs struggle with long-horizon tasks, especially under partial observ…

  95. arXiv cs.AI TIER_1 English(EN) · Rakibul Hasan Rajib, Mengxing Zheng, Qian Lou ·

    学习保留什么:多智能体LLM系统中高效协作的门控记忆路由

    arXiv:2609.00237v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent systems tackle complex reasoning by orchestrating how multiple agents are configured and how they collaborate. A central challenge is to adapt orchestration to the evolving collaboration …

  96. Hugging Face Daily Papers TIER_1 English(EN) ·

    多行动,少决策:面向长时域LLM智能体的技能引导自适应行动分块

    Large language model (LLM) agents for long-horizon interactive tasks typically follow a ReAct-style protocol, issuing one primitive action per LLM round. While this enables frequent replanning, it is inefficient for long-horizon tasks where many rounds are spent on routine action…

  97. Hugging Face Daily Papers TIER_1 English(EN) ·

    双层协同反射:多智能体LLM系统的博弈论方法

    The study formalizes multi-agent LLM coordination via bilevel games and stochastic memory reflection, introducing a grounded evaluation gate and SRMA algorithm with convergence guarantees, validated on SWE-bench.

  98. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Haohan Wang ·

    通过 Agent 原生可复用工具原语在 LLM 工具使用中进行工程化

    Large language models (LLMs) augmented with external tools have demonstrated remarkable capability in solving complex real-world tasks. However, existing approaches suffer from two key challenges: brittle multi-step and multi-turn reasoning caused by incompatible tool output type…

  99. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Simon Kirby ·

    GlossoGen:复杂多智能体LLM交互中的涌现语言

    The growing rate at which LLM agents interact with one another raises key questions about language evolution in multi-LLM-agent settings, with implications for safety and monitorability as well as for linguistic accounts of LLMs. To address these questions, we introduce GlossoGen…

  100. arXiv cs.AI TIER_1 English(EN) · Wujie Xiong, Rabimba Karanjai, Yang Lu, Weidong Shi, Lei Xu ·

    基于可达性的LLM代理在间接提示注入下的能力约束

    arXiv:2608.30041v1 Announce Type: cross Abstract: Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later privileged actions. Existing defenses mainly classify untrusted content or authorize…

  101. arXiv cs.AI TIER_1 English(EN) · Hyomin Lee, Sangwoo Park, Yumin Choi, Sohyun An, Seanie Lee, Sung Ju Hwang ·

    T-MAP:使用轨迹感知进化搜索对 LLM 代理进行红队测试

    arXiv:2603.22341v2 Announce Type: replace-cross Abstract: While prior red-teaming efforts have focused on eliciting harmful text outputs from large language models (LLMs), such approaches fail to capture agent-specific vulnerabilities that emerge through multi-step tool execution…

  102. arXiv cs.AI TIER_1 English(EN) · Guangyi Liu, Haojun Lin, Huan Zeng, Heng Wang, Quanming Yao ·

    MAS-on-the-Fly: 基于LLM的多智能体系统的上下文结构自适应

    arXiv:2602.13671v2 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, existing works often rely on manual designs or "one-size-fits-all" automation and lack adaptabilit…

  103. arXiv cs.AI TIER_1 English(EN) · Angel Yanguas-Gil ·

    评估集成材料合成工具的基于LLM的AI代理:以原子层沉积为例

    arXiv:2608.29309v1 Announce Type: cross Abstract: This work provides an overview of the different strategies that can be used to evaluate the performance of AI models and agents based on large language models (LLMs) for materials synthesis. After providing a brief overview of the…

  104. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Yuntian Deng ·

    控制-数据流分离:多智能体LLM中的稳定提示优化

    Prompt optimization can improve multi-agent LLM systems, but the prompts being optimized often serve two entangled roles: generating task-relevant content and specifying execution-critical protocols, such as message routing, output formatting, and termination signals, on which th…

  105. Hugging Face Daily Papers TIER_1 English(EN) ·

    控制-数据流分离:多智能体LLM中的稳定提示优化

    The framework separates structured execution protocols from optimizable language content to prevent prompt optimization from corrupting multi-agent pipelines.

  106. Hugging Face Daily Papers TIER_1 English(EN) ·

    HarnessDev:大型语言模型能否创建和演进自己的Agent Harness?

    HarnessDev evaluates agents by measuring their ability to build and iteratively improve execution infrastructure rather than final task outputs, revealing that self-built harnesses vary widely in capability and efficiency and transfer poorly across models.

  107. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Tulika Mitra ·

    基于LLM的硬件开发,采用分层IR和端到端多智能体工作流

    Large language models (LLMs) are increasingly used in software development, but their use in complex hardware design remains limited. This gap stems from both the scarcity of public hardware training data and the fundamentally different methodologies used in hardware design. In p…

  108. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Arka Majumdar ·

    HALO: 用于纳米光子设计的物理感知大语言模型代理框架

    Language models have recently been applied to nanophotonic design, but it remains unclear whether they can reliably translate optical objectives into simulation-ready designs, execute electromagnetic analysis, and revise decisions from numerical feedback. We introduce HALO, a phy…

  109. arXiv cs.CL TIER_1 English(EN) · Chung-En Sun, Linbo Liu, Tsui-Wei Weng ·

    LLM智能体中的冷启动安全鸿沟

    arXiv:2606.07867v2 Announce Type: replace Abstract: Are tool-calling LLM agents equally safe throughout a conversation? We discover they are not: agents are most vulnerable at the very start of a session and become substantially safer after a few regular agentic tasks -- a phenom…

  110. arXiv cs.CL TIER_1 English(EN) · Tatiana Petrova, Andrei Mazniak, Radu State ·

    Agent 不分页:LLM 工具响应的首块选择

    arXiv:2608.26130v1 Announce Type: new Abstract: Coding agents built on large language models (LLMs), such as Claude Code, Cursor, OpenAI Codex, GitHub Copilot, and Aider, receive tool responses that routinely exceed the agent's per-turn token budget. The standard remedy, paginati…

  111. arXiv cs.AI TIER_1 English(EN) · Yisen Xi ·

    Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit

    arXiv:2608.27427v1 Announce Type: cross Abstract: Large language model (LLM) agents in governed organizations must let the persona (instructions, tone, self-presentation) evolve freely, while keeping execution (stateful, audited work) traceable. A single trust domain does not sat…

  112. arXiv cs.AI TIER_1 English(EN) · Xingshan Zeng, Zishan Xu, Boju Zhang, Yuzhou Wu, Lingzhi Wang, Jianghao Lin, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Weinan Zhang, Yong Yu, Qun Liu, Weiwen Liu ·

    什么构成优质的 Agentic 数据?从 ACE 视角审视 LLM Agent 的数据生成

    arXiv:2608.27260v1 Announce Type: new Abstract: LLM agents increasingly rely on generated interaction data to learn how to interact with external environments. Agentic data generation must maintain consistency among environments, tasks, interactions, and success signals while pro…

  113. arXiv cs.AI TIER_1 English(EN) · Mesut Toruk ·

    BekchiAI:一键测量、观察和控制 LLM 智能体

    arXiv:2608.26867v1 Announce Type: new Abstract: Large language model agents reason, call tools, and act autonomously over many steps, but their agentic skills-correctly sequencing tools, planning under dependencies, judging untrusted inputs, and grounding generated arguments-are …

  114. arXiv cs.AI TIER_1 English(EN) · Kuan-Hao Tseng, Niruth Bogahawatta, Yasod Ginige, Kunjan Patel, Kosta Dakic, Suranga Seneviratne ·

    FaulT-Bench:面向不可靠用户工单下的网络故障排除大模型智能体基准测试

    arXiv:2608.27021v1 Announce Type: cross Abstract: LLM-based agents are increasingly proposed for network fault diagnosis, but existing benchmarks evaluate them only on accurate tickets and always assume a fault is present, conditions rarely met in practice. We present FaulT-Bench…

  115. arXiv cs.AI TIER_1 English(EN) · Yitian Zhou, Jingyu Zheng, Qiliang Jiang, Linkang Du, Haoming Liu, Lichao Wu, Shiyi Zhao, Mengxiang Liu, Ruilong Deng ·

    PLCBench:自主LLM代理能否将PLC访问转化为持续的物理影响?

    arXiv:2608.26882v1 Announce Type: cross Abstract: Industrial control systems (ICSs) rely on programmable logic controllers (PLCs) to connect networked computation with physical control. Tool-using large language model (LLM) agents represent an emerging attack threat: can an auton…

  116. arXiv cs.AI TIER_1 English(EN) · Kimberly Milner, Minghao Shao, Nanda Rani, Haoran Xi, Venkata Sai Charan Putrevu, Meet Udeshi, Sandeep K. Shukla, Prashanth Krishnamurthy, Farshad Khorrami, Muhammad Shafique, Ramesh Karri ·

    大型语言模型代理如何实际获取Flag?代理式进攻性安全评估的追踪级出处

    arXiv:2608.26237v1 Announce Type: cross Abstract: Capture-the-Flag (CTF) benchmarks are widely used to assess the offensive security capabilities of autonomous language-model agents. Evaluations rely on shallow binary judgments or aggregate scores, overlooking the agent's traject…

  117. arXiv cs.AI TIER_1 English(EN) · Dinh-Khanh Pham, Quy-Anh Dang, Lam Mai Thanh, Khanh Bui, Truong-Son Hy ·

    从SQL到知识图谱:一种由LLM驱动的多智能体方法及数据模式改进

    arXiv:2608.26117v1 Announce Type: cross Abstract: RDBMS (Relational Database Management System) databases face several limitations, including slow execution with multi-hop queries and a lack of explainability by graphical interpretations. In contrast, Graph database offers a more…

  118. arXiv cs.AI TIER_1 English(EN) · Yilong Chen, Xiao Qin, Chenghao Liu, Liang Wu, Noelle I. Samia, Kaize Ding ·

    用于时间序列的大型语言模型代理:一项调查

    arXiv:2608.26226v1 Announce Type: new Abstract: LLM-based agents are increasingly being developed for time-series problems, but their design choices vary substantially across task settings. This survey adopts a problem-driven taxonomy that organizes these systems by the time-seri…

  119. arXiv cs.AI TIER_1 English(EN) · Abhigya Verma, Amit Kumar Saha, Seganrasan Subramanian, Sai Harshitha Aluru ·

    AgentJudgeBench:一个多难度基准,用于评估 LLM 裁判在代理工具调用方面的能力

    arXiv:2608.26623v1 Announce Type: new Abstract: LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJudgeBench, the first benchmark to systematically stud…

  120. arXiv cs.AI TIER_1 English(EN) · Haiteng Wang, Weihao Li, Jing Zhang, Lei Ren ·

    AI控制科学家:用于自动化控制设计的LLM驱动的代理系统

    arXiv:2608.26780v1 Announce Type: new Abstract: Control system design is critical for modern industry, such as chemical process temperature regulation and aero-engine control. However,traditional control design workflows rely heavily on expert knowledge and extensive manual param…

  121. arXiv cs.AI TIER_1 English(EN) · Jiayi Kuang, Yinghui Li, Yunze Song, Keyu Chen, Zhifeng Shen, Yangning Li, Yidong Wang, Di Yin, Ruizhi Qiao, Xing Sun, Kai Jin, Ying Shen, Liang Lin, Philip S. Yu ·

    从原子到智能体:迈向LLM智能体数学能力的可解释评估

    arXiv:2608.26950v1 Announce Type: new Abstract: Large Language Models (LLMs) are evolving from performing end-to-end mathematical reasoning to integrating agentic intelligence. However, most existing math benchmarks evaluate only final answers. This outcome-oriented evaluation pr…

  122. arXiv cs.AI TIER_1 English(EN) · Linsen Zhu, Yi Shi ·

    DSA:面向多市场股票研究的证据感知大语言模型代理编排

    arXiv:2608.26990v1 Announce Type: new Abstract: Large language models can summarize financial information, but an operational stock-research system must first assemble heterogeneous evidence, expose unavailable data and model capabilities, and control how generated opinions affec…

  123. arXiv cs.AI TIER_1 English(EN) · Xiaokun Guo, Zhen Xu, Dongdong Huo, Yanqiu Zhang, Wei Wang, Qinfu Yang, Dongjin Yu, Yu Wang ·

    当工具输出成为命令:在工具增强型LLM代理中分离动作诱导与运行时授权

    arXiv:2608.27146v1 Announce Type: new Abstract: Tool-augmented LLM agents must rely on untrusted runtime Observations to complete open-ended tasks; however, when tool outputs no longer merely provide data but begin to specify concrete actions, they effectively become ``commands''…

  124. arXiv cs.AI TIER_1 English(EN) · Chenhao Wu, Haoxuan Jia, Yang Liu, Yingguang Yang, Yuhan Lin, Chongyang Zhang, Hao Zheng, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Shang Luo, Kefu Xu, Jifeng Zhu, Bin Chong ·

    安全无法组合:自主LLM代理的非衰减循环状态

    arXiv:2608.27141v1 Announce Type: cross Abstract: Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unatten…

  125. Hugging Face Daily Papers TIER_1 English(EN) ·

    EvoUndo:LLM Agent Harnesses 的可恢复性约束的自我演化

    EvoUndo evaluates recoverability of self-modifying LLM agents and shows that reliable recovery requires co-designing verification, state grounding, and recovery-language expressivity.

  126. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Yi Shi ·

    DSA:面向多市场股票研究的证据感知大语言模型代理编排

    Large language models can summarize financial information, but an operational stock-research system must first assemble heterogeneous evidence, expose unavailable data and model capabilities, and control how generated opinions affect a final report. We present DSA, an evidence-aw…

  127. Hugging Face Daily Papers TIER_1 English(EN) ·

    PLCBench:自主LLM代理能否将PLC访问转化为持续的物理影响?

    Industrial control systems (ICSs) rely on programmable logic controllers (PLCs) to connect networked computation with physical control. Tool-using large language model (LLM) agents represent an emerging attack threat: can an autonomous agent convert a network-reachable PLC into s…

  128. arXiv cs.LG TIER_1 English(EN) · Siyuan Chen, Runlin Hou, Shenxiu Wu, Yansong Sun, Junming Cao, Yiyu Zhang, Shudi Shao, Junhao Qiu, Zhichao Lu, Qingfu Zhang ·

    超越扩展:通过经验驱动的工作流程和经验图谱记忆实现硬件内核优化的自进化大语言模型代理

    arXiv:2608.25570v1 Announce Type: new Abstract: Hardware kernel optimization requires repeated compilation, correctness testing, profiling, and revision. LLM agents can automate parts of this process, and stronger foundation models, longer context windows, and longer execution ho…

  129. arXiv cs.CL TIER_1 English(EN) · Hongqiu Ni, Han Tian, Chi Zhang, Guopeng Li, Haisheng Tan ·

    TOPAS:面向多智能体LLM服务的感知工作流前缀状态调度

    arXiv:2608.25523v1 Announce Type: new Abstract: Prefix caching introduces a fundamental tradeoff in multi-agent large language model (LLM) serving: retaining a long system-prompt key-value (KV) cache for an agent accelerates future calls, yet it reduces the GPU memory available f…

  130. arXiv cs.CL TIER_1 English(EN) · Pratyay Banerjee, Ankit Chadha ·

    路由图切换:多智能体LLM委托的自适应格式选择

    arXiv:2608.25277v1 Announce Type: new Abstract: Multi-agent LLM systems coordinate through natural-language messages that consume 40--60\% of their token budget. Replacing these with structured graphs reduces cost but fails on tasks requiring adaptive reasoning. We propose \textb…

  131. Hugging Face Daily Papers TIER_1 English(EN) ·

    什么造就了优质的 Agentic 数据?从 ACE 视角看 LLM Agent 的数据生成

    Agentic data generation is framed as constrained distribution design over factorized experience tuples, emphasizing execution-grounded accuracy, learner-relative complexity, and diversity rather than scale alone.

  132. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Shiqiang Wang ·

    ProgRouter:在线进度引导式编排,用于多智能体 LLM 工作流的质量-成本权衡

    Multi-agent large language model (LLM) workflows have emerged as a powerful paradigm for solving complex, open-ended tasks through collaborative reasoning among specialized LLM agents, but they incur substantial operating costs due to repeated LLM invocations and long-horizon con…

  133. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Zhiyuan Yuan ·

    候选者供给和答案选择塑造了多智能体系统中LLM评判的价值

    Multi-agent systems (MAS) sometimes already have the potential to answer correctly, but still report a wrong answer. Explaining this outcome is difficult because generation, communication and final answer-selection rules usually change simultaneously. We conceptualize multi-agent…

  134. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Qingfu Zhang ·

    超越规模化:通过经验驱动工作流和经验图谱记忆实现硬件内核优化的自进化大语言模型代理

    Hardware kernel optimization requires repeated compilation, correctness testing, profiling, and revision. LLM agents can automate parts of this process, and stronger foundation models, longer context windows, and longer execution horizons have improved optimization within individ…

  135. arXiv cs.AI TIER_1 English(EN) · Yusheng Li, Tianjun Feng, Yunfeng Chen, Chun-Yi Tsai, Yihan Sun, Ayan Das, Kaoutar El Maghraoui, Shuxin Lin, Dhaval Patel ·

    PHMForge:通过 MCP 原生、算法为基础的工具评估工业预测中的 LLM 代理

    arXiv:2604.01532v3 Announce Type: replace Abstract: LLM agents are beginning to invoke industrial asset-management tools through the Model Context Protocol (MCP), yet whether they can act reliably on this substrate for safety-critical \emph{Prognostics and Health Management (PHM)…

  136. arXiv cs.AI TIER_1 English(EN) · Yuchen Xia, Michael Weyrich, Nasser Jazdi, Johannes St\"umpfle, Johannes Sigel, Akshay Narla, Gavin K. Reynolds, Anna Jawor-Baczynska, Pol Llopart ·

    LLM 代理使用模拟模型进行受控实验

    arXiv:2608.23622v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require more than plausible text and code generation. They require understanding how a syste…

  137. arXiv cs.AI TIER_1 English(EN) · Roy Ganz, Mor Shpigel Nacson, Adi Kalyanpur, Ron Litman ·

    交接税:LLM智能体中持续存在的非原生轨迹

    arXiv:2608.24358v1 Announce Type: new Abstract: Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As these runs unfold, users face a practical cost-quality trade-off: escalating to a stronger model when a cheaper one struggles, or…

  138. arXiv cs.AI TIER_1 English(EN) · Zhi-Kai Chen, Xu-Xiang Zhong, Song-Yan Li, De-Chuan Zhan, Han-Jia Ye ·

    PeakBench:LLM智能体中资源感知工具调用的基准测试

    arXiv:2608.24509v1 Announce Type: new Abstract: LLM agents increasingly solve tasks by invoking multiple tools, where parallel execution is essential for low latency but difficult to manage safely. Existing agent benchmarks primarily evaluate tool selection, argument generation, …

  139. arXiv cs.AI TIER_1 English(EN) · Yiheng Sun, Huifei Wang, Yancheng Zhu, Zhenyu Li, Zebin Zhao, Yifan Yuan ·

    当“必须”变成“也许”:LLM 代理工作流中的约束弱化

    arXiv:2608.24569v1 Announce Type: new Abstract: Large language model (LLM) agents coordinate complex tasks through multi-role and multi-stage workflows. Upstream state is repeatedly transformed into intermediate language artifacts, such as summaries, plans, tickets, memories, and…

  140. arXiv cs.AI TIER_1 English(EN) · Dai Jiahong ·

    分久必合:三大LLM智能体框架中的架构融合

    arXiv:2608.23953v1 Announce Type: cross Abstract: An agent harness is what turns a language model into an autonomous agent: the surrounding code that builds the model's context, mediates its tools, runs the loop, and persists state across a long-horizon run. This layer, not the m…

  141. arXiv cs.AI TIER_1 English(EN) · Lin-Fa Lee, YI-YU Chang, Kuo-Hui Yeh ·

    WebMCP-Phalanx:为浏览器集成的大型语言模型代理强制执行和表征信任边界

    arXiv:2608.24017v1 Announce Type: cross Abstract: The emerging W3C WebMCP proposal enables LLM agents to invoke tools exposed by web pages. In multi-party web environments, however, integrating agent execution into a browser security model centered on the Same-Origin Policy (SOP)…

  142. arXiv cs.AI TIER_1 English(EN) · Wonung Kim, Hyunmin Choi, Minsu Kim, Jaehong Cho, Yeongwook Kim, Jongse Park ·

    Simthesizer:面向 LLM 服务系统的代理驱动模拟框架

    arXiv:2608.24650v1 Announce Type: cross Abstract: System-level simulation is an essential tool for exploring the rapidly expanding design space of LLM serving systems, where real deployments remain costly and often infeasible. However, modern LLM serving now evolves faster than h…

  143. arXiv cs.AI TIER_1 English(EN) · Shijun Lei, Quang Nguyen, Swapneel S Mehta, Zeping Li, Huichuan Fu, Xiaolong Zheng, Siki Chen, Yunji Liang, Philip Torr, Zhenfei Yin ·

    LLM 智能体市场中的战略剥削:电子商务信任的模拟框架

    arXiv:2605.10059v3 Announce Type: replace Abstract: Agent-based modeling (ABM) has long been used in economics to study human behavior, and large language model (LLM) agents now enable new forms of social and economic simulation. While prior work has discovered strategic deceptio…

  144. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Yifan Yuan ·

    当“必须”变成“也许”:LLM 代理工作流中的约束弱化

    Large language model (LLM) agents coordinate complex tasks through multi-role and multi-stage workflows. Upstream state is repeatedly transformed into intermediate language artifacts, such as summaries, plans, tickets, memories, and handoff notes, from which downstream components…

  145. arXiv cs.AI TIER_1 English(EN) · Xu Yang, Chenhui Lin, Haotian Liu, Qi Wang, Yue Yang, Wenchuan Wu ·

    一次请求,多个专家:LLM通过自适应任务路由编排领域特定模型

    arXiv:2511.12484v2 Announce Type: replace-cross Abstract: With the integration of massive distributed energy resources and the widespread participation of novel market entities, the operation of active distribution networks (ADNs) is progressively evolving into a complex, multi-s…

  146. arXiv cs.CL TIER_1 English(EN) · Benjamin Plaut ·

    LLM 智能体中的安全训练或在帮助性优化过程中持续进行

    arXiv:2603.02229v2 Announce Type: replace-cross Abstract: Safety post-training has been studied extensively in single-step "chat" settings where safety typically refers to refusing harmful requests. We study an "agentic" (i.e., multi-step, tool-use) setting where safety refers to…

  147. arXiv cs.CL TIER_1 English(EN) · Yaokun Liu, Yifan Liu, Daniel Yue Zhang, Ruichen Yao, Zelin Li, Dong Wang ·

    PropUQ-MAS:面向LLM多智能体系统的传播感知不确定性量化

    arXiv:2608.22130v1 Announce Type: cross Abstract: LLM-based multi-agent systems (MAS) solve complex tasks through communication among role-specialized agents. However, inter-agent dependencies introduce reliability risks beyond isolated agent failures. For instance, errors in int…

  148. arXiv cs.CL TIER_1 English(EN) · Vedant Khatri, Anthony Cusimano, Zachari Swiecki, Zhen Xu, Xiner Liu, Renzhe Yu ·

    从诊断到重新设计:使用定量民族志改进多智能体LLM推理

    arXiv:2608.22566v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems are designed to improve reasoning by decomposing tasks across multiple agents with specialized functions, but the presence of multiple agents does not inherently guarantee coherent reas…

  149. arXiv cs.CL TIER_1 English(EN) · Weixiang Sun, Zehong Wang, Hong Huang, Colby Nelson, Yanfang Ye ·

    协作税:LLM多智能体系统为协调付出多少代价

    arXiv:2608.22152v1 Announce Type: new Abstract: Multi-agent systems built from large language models are deployed widely, yet how much performance is lost when two LLMs must coordinate rather than act alone remains unclear. We formulate the collaboration tax as the team-decentral…

  150. arXiv cs.AI TIER_1 English(EN) · Yinuo Liu, Zi Qian, Heng Zhou, Jiahao Zhang, Yajie Zhang, Zhihang Li, Mengyu Zhou, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang ·

    ATP-Bench:迈向多模态大模型交错生成中的智能体工具规划

    arXiv:2603.29902v2 Announce Type: replace Abstract: Interleaved text-and-image generation represents a significant frontier for Multimodal Large Language Models (MLLMs), offering a more intuitive way to convey complex information. Current paradigms rely on either image generation…

  151. arXiv cs.AI TIER_1 English(EN) · Weiwei Sun, Xuhui Zhou, Weihua Du, Xingyao Wang, Sean Welleck, Graham Neubig, Maarten Sap, Yiming Yang ·

    训练主动式和个性化大语言模型代理

    arXiv:2511.02208v2 Announce Type: replace Abstract: Despite rapid progress, current AI agents are primarily optimized for isolated task completion. We argue for a paradigm shift toward training agents as collaborators that communicate and adapt to people. To facilitate this shift…

  152. arXiv cs.AI TIER_1 English(EN) · Chang Liu, Xiaohui Xie, Xinyi Chen, Yong Cui ·

    NetConfArena:用于闭环网络配置中 LLM Agent 的可执行基准测试

    arXiv:2608.23179v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly attractive for automating network configuration, yet their reliability and failure patterns are poorly understood. An essential prerequisite is to assess such agents in a realisti…

  153. arXiv cs.AI TIER_1 English(EN) · Jiatong Li, Wengyu Zhang, Weida Wang, Yuxuan Ren, Wei Liu, Chenyang Mao, Yuqiang Li, Yatao Bian, Changmeng Zheng, Xiaoyong Wei, Qing Li ·

    分子大语言模型智能体:从架构设计到科学自主

    arXiv:2608.23104v1 Announce Type: cross Abstract: Molecular science represents an important frontier for LLM-based agents. Unlike general agents that mainly operate over natural language, code, or web environments, molecular LLM agents must perceive, reason about, and act upon ch…

  154. arXiv cs.AI TIER_1 English(EN) · Ryuki Hyodo ·

    LLM 和 VLM 驱动的 2D 和 3D 环境中智能体的最小本地模拟基础

    arXiv:2608.22833v1 Announce Type: cross Abstract: Large language models (LLMs) and vision-language models (VLMs) are expanding the range of behaviors that can be represented in agent-based simulations, but many contemporary platforms are difficult to study, modify, or run on ordi…

  155. arXiv cs.AI TIER_1 English(EN) · Wenhao Wu, Menghao Zhang, Xin Wang, Zhi Wang, Kun Shao, Jian Luan ·

    TRACE:一个用于一致性、限制感知型LLM代理的自演化技能库

    arXiv:2608.22793v1 Announce Type: cross Abstract: Reliable deployment of LLM agents in user-facing products depends not on raw task-solving ability but on consistency and limit-awareness: behaving the same way across repeated trials, and recognizing when a request cannot, or cann…

  156. arXiv cs.AI TIER_1 English(EN) · Xinyuan Liu, Eren Sadikoglu, Riana Chatterjee, Ransalu Senanayake ·

    物理代理AI:一种用于编排LLM机器人团队的架构

    arXiv:2608.22657v1 Announce Type: cross Abstract: Agentic AI frameworks interpret open-ended task goals and decompose them into multi-step plans. Richer information about embodiment-specific capabilities, physical preconditions, and cross-robot coordination improves grounding, bu…

  157. arXiv cs.AI TIER_1 English(EN) · Baicheng Chen, Zheyuan Liu, Jingyu Zhang, Kaize Ding, Ningshan Ma, Yue Huang, Meng Jiang ·

    遗忘于权重,工具可恢复:LLM智能体的代理工具遗忘

    arXiv:2608.21544v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as tool-augmented agents, where responses can depend on tool calls and external observations rather than model parameters alone. This creates an evaluation mismatch for LLM un…

  158. arXiv cs.AI TIER_1 English(EN) · Israt Moyeen Noumi, Tarannum Ahmed Nowshin, Md. Mehedi Hasan Nipu, Mohammad Sakib Mahmood, Md. Jakir Hossain, M. F. Mridha ·

    Agentic Security:LLM驱动的渗透测试的工具、故障模式和设计法则的系统化

    arXiv:2608.21423v1 Announce Type: cross Abstract: Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools. As these systems move from demonstrations to deployed products, practitioners repeatedly encounter the same operational failu…

  159. arXiv cs.AI TIER_1 English(EN) · Seth Karten, Alex L. Zhang, Kevin Thomas, Sebastian M\"uller, Elie Bakouch, Daniel Auras, Mika Senghaas, Fares Obeid, Konstantin Dunas, Johannes Hagemann, Sami Jaghouar ·

    Prime Agent:一个自改进的RLM工具集

    arXiv:2608.23552v1 Announce Type: new Abstract: Language models are sequential processors, but long-horizon agency requires external information and computation beyond model weights and active context. Prime Agent is an open-source harness for long-horizon evaluation and coding-a…

  160. arXiv cs.AI TIER_1 English(EN) · Xiaogang Xu, Jiaqi Tang, Jianmin Chen, Yingying Yan, Zhenchao Tang, Xiangxin Zhou, Xiaobin Hu, Wei Wei, Jinfeng Wu, Qifeng Chen, Lu Zhou, Jiafei Wu, Zhe Liu, Jianwei Yin, Weimin Zheng ·

    基于LLM的代理用于预测和预报:方法、训练、评估和应用

    arXiv:2608.23058v1 Announce Type: new Abstract: Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external tools, and iterative prediction. We investigate LLM-based forecasting agents, meanin…

  161. arXiv cs.AI TIER_1 English(EN) · Xiaohui Zhang, Zequn Sun, Chengyuan Yang, Yuanning Cui, Lingbing Guo, Wei Hu ·

    面向通过动态本体实现有效可靠的LLM代理

    arXiv:2608.22974v1 Announce Type: new Abstract: Large language model (LLM) agents rely heavily on knowledge encoded in model parameters or presented as unstructured context. In domain-specific tasks, this leaves important semantic connections implicit. This often results in incom…

  162. arXiv cs.AI TIER_1 English(EN) · Hui Zeng, Pengfei Yang, Yanxin Chen, Fusong Ju, Xinran Wei ·

    LLM4LLM:通过闭环代理优化连接内核基准测试与实际部署

    arXiv:2608.21836v1 Announce Type: new Abstract: Large language models have become increasingly capable agents for low-level code and kernel optimization, but isolated kernel benchmarks provide only a proxy for the deployment behavior that matters in language-model inference. We i…

  163. Hugging Face Daily Papers TIER_1 English(EN) ·

    分久必合:三大LLM智能体框架中的架构融合

    An agent harness is what turns a language model into an autonomous agent: the surrounding code that builds the model's context, mediates its tools, runs the loop, and persists state across a long-horizon run. This layer, not the model it wraps, is increasingly the binding constra…

  164. Hugging Face Daily Papers TIER_1 English(EN) ·

    交接税:LLM智能体中持续的非原生轨迹

    Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As these runs unfold, users face a practical cost-quality trade-off: escalating to a stronger model when a cheaper one struggles, or downshifting once the hard reasoning is complet…

  165. Hugging Face Daily Papers TIER_1 English(EN) ·

    当“必须”变成“也许”:LLM 代理工作流中的约束弱化

    Multi-stage LLM workflows lose operational constraints when intermediate artifacts transform binding prerequisites into non-binding context, causing safety failures despite preserved content.

  166. arXiv cs.MA (Multiagent) TIER_1 English(EN) · James Evans ·

    市场而非规划者:具有私有信息的LLM代理的去中心化编排

    As LLM agents proliferate, built by different parties and with different capabilities and costs, orchestrating them is more like assembling labor across the economy than a computer calling a subroutine. Existing orchestration is typically centralized, with a single planner assign…

  167. Hugging Face Daily Papers TIER_1 English(EN) ·

    分子大语言模型Agent:从架构设计到科学自主

    Molecular science represents an important frontier for LLM-based agents. Unlike general agents that mainly operate over natural language, code, or web environments, molecular LLM agents must perceive, reason about, and act upon chemical objects across symbolic strings, molecular …

  168. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Ryuki Hyodo ·

    面向2D和3D环境中LLM和VLM驱动的代理的最小化本地仿真基础

    Large language models (LLMs) and vision-language models (VLMs) are expanding the range of behaviors that can be represented in agent-based simulations, but many contemporary platforms are difficult to study, modify, or run on ordinary computers. We present two intentionally minim…

  169. arXiv cs.AI TIER_1 English(EN) · Yuyuan Feng, Zhishang Xiang, Chaobin Yang, Qichao Ma, Zerui Chen, Yujing Zhang, Ke Huang, Chuanjie Wu, Zhaoxu Liu, Yili Wang, Xin He, Jiapu Wang, Zijin Hong, Hao Chen, Yuanchen Bei, Kun Wang, Shengyuan Chen, Ningyu Zhang, Enyan Dai, Linhao Luo, Qingyi Pa… ·

    LLM智能体时代的图工程:从个体智能到系统智能

    arXiv:2608.21156v1 Announce Type: cross Abstract: LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolution has produced paradigms including Prompt Engineering to elicit model capabilities, Context Engineering to manage…

  170. arXiv cs.AI TIER_1 English(EN) · Quang Dao, Purvi Kathalkar, Kenneth Eaton ·

    加权记忆树:让长时域LLM智能体记住重要信息

    arXiv:2608.20631v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated the ability to solve multi-step tasks requiring planning, tool use, and external information access, yet growing execution histories increase inference cost and expose reasoning to…

  171. arXiv cs.AI TIER_1 English(EN) · Guodong Xu ·

    LLM智能体中的标准修订校准:失效模式与基于轨迹的协议

    arXiv:2608.20729v1 Announce Type: new Abstract: Language-model agents can improve after failure or carry text across episodes without revising what counts as success. We study the narrower attribution problem of criterion revision: when criterion K0 accepts an outcome violating a…

  172. arXiv cs.AI TIER_1 English(EN) · Yanze Jiang, Mingxuan Li, Yuhao Wang, Shengfang Zhai, Jiaheng Zhang ·

    无需解决,只需比较:用于LLM代理运行时干预的微型顾问

    arXiv:2608.21027v1 Announce Type: new Abstract: LLM agents are emerging as an important paradigm for real-world tasks that require reasoning, tool use, and sequential decision-making. As these agents operate over longer horizons, runtime intervention offers a way to improve relia…

  173. arXiv cs.AI TIER_1 English(EN) · Wei Shao, Chongzhou Fang, Zuxiong Tan, Zequan Liang, Setareh Rafatirad, Avesta Sasan, Houman Homayoun ·

    超越端到端成功:诊断长周期安全LLM代理的失败

    arXiv:2608.20563v1 Announce Type: cross Abstract: Long-horizon security LLM agents must carry information and decisions across many dependent interactions, where later actions often depend on services, state, or access discovered much earlier. This makes final task success diffic…

  174. arXiv cs.AI TIER_1 English(EN) · Jiajun Wu, Zirui Wang, Jiayu Zhou, Qiang Ye, Steve Drew ·

    FL-MAESTRO:面向资源受限联邦学习的多智能体大模型编排

    arXiv:2608.20518v1 Announce Type: new Abstract: In Federated Learning (FL), the communication topology is a runtime variable rather than a fixed design choice, since links and edge devices drop in and out during training. Each round, the server must commit three coupled decisions…

  175. arXiv cs.AI TIER_1 English(EN) · Kai Wang, Zeming Wei, BiaoJie Zeng, Chang Jin, An Wang, Xiaokun Luan, Zhixiao Lin, Jingjing Qu, Xia Hu, Xingcheng Xu ·

    ClawSentry:一款渐进式多层安全监控器,用于保护自主LLM代理

    arXiv:2608.21101v1 Announce Type: cross Abstract: As large language model (LLM) agents move from conversation to executing code, reading local files, and orchestrating external tools, a single agent hijacked by a malicious third-party skill can cause data exfiltration, privilege …

  176. arXiv cs.LG TIER_1 English(EN) · Quan Wei, Siliang Zeng, Chenliang Li, Zhongruo Wang, William Brown, Oana Frunza, Wei Deng, Anderson Schneider, Yuriy Nevmyvaka, Yang Katie Zhao, Alfredo Garcia, Mingyi Hong ·

    通过细粒度奖励结构和信用分配增强LLM智能体多轮推理能力

    arXiv:2505.11821v3 Announce Type: replace Abstract: Reinforcement Learning (RL) approaches have been wildly used to enhance the reasoning capabilities of Large Language Model (LLM) agents in long-horizon, multi-turn scenarios. Such interactions can be formalized as turn-level Mar…

  177. Hugging Face Daily Papers TIER_1 English(EN) ·

    Prime Agent:一个自改进的RLM工具集

    Prime Agent is an open-source harness that uses recursive subagents, persistent computation, and agent-to-agent coordination to extend language models' long-horizon capabilities across coding and reasoning tasks.

  178. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Ransalu Senanayake ·

    Physical Agentic AI:一个用于编排拥有大型语言模型机器人团队的架构

    Agentic AI frameworks interpret open-ended task goals and decompose them into multi-step plans. Richer information about embodiment-specific capabilities, physical preconditions, and cross-robot coordination improves grounding, but does not eliminate infeasible, mistimed, or unsa…

  179. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Dong Wang ·

    PropUQ-MAS:用于大语言模型多智能体系统的传播感知不确定性量化

    LLM-based multi-agent systems (MAS) solve complex tasks through communication among role-specialized agents. However, inter-agent dependencies introduce reliability risks beyond isolated agent failures. For instance, errors in intermediate messages could be inherited and amplifie…

  180. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Pol Llopart ·

    LLM 代理使用模拟模型进行受控实验

    Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require more than plausible text and code generation. They require understanding how a system responds to intervention, which in practice de…

  181. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Hongzhi Yin ·

    利用LLM智能体中的记忆增强推理来增强群组推荐

    The core challenge in group recommendation lies in modeling the dynamic evolution of user preferences and explain?ing the consensus formation process. Existing Large Language Model (LLM)-based methods, despite improved interpretability, treat interaction history as fixed text, ig…

  182. Hugging Face Daily Papers TIER_1 English(EN) ·

    LLM智能体时代的图工程:从个体智能到系统智能

    LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolution has produced paradigms including Prompt Engineering to elicit model capabilities, Context Engineering to manage information access, Harness Engineering to organi…

  183. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Yi Chang ·

    LLM智能体时代的图工程:从个体智能到系统智能

    LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolution has produced paradigms including Prompt Engineering to elicit model capabilities, Context Engineering to manage information access, Harness Engineering to organi…

  184. arXiv cs.CL TIER_1 English(EN) · Mehdi Azarafza, Faezeh Pasandideh, Ali Ehteshami Bejnordi, Stefan Henkler, Achim Rettberg ·

    利用LLM的常识推理能力进行多智能体协同以实现自动驾驶

    arXiv:2608.20129v1 Announce Type: cross Abstract: Autonomous vehicles require robust perception and decision-making capabilities to operate in diverse and unseen scenarios. While reinforcement learning and rule-based methods can provide effective control and safety mechanisms, th…

  185. arXiv cs.CL TIER_1 English(EN) · Hexi Wang, Yujia Zhou, Bangde Du, Weihang Su, Xinyuan Cao, Qingyi Pan, Qingyao Ai, Yueyue Wu, Min Zhang, Yiqun Liu ·

    通过纵向生命轨迹减轻大型语言模型代理中的身份本质主义

    arXiv:2608.19621v1 Announce Type: new Abstract: Large language models (LLMs) offer a scalable approach to social simulation, but their credibility depends on how agents are constructed. Existing methods can partially reproduce population-level patterns, yet often fail to capture …

  186. arXiv cs.AI TIER_1 English(EN) · Bo Qian, Yuting Wu, Shuang Zeng, Huaiyu Wan, Dalin Zhang, Jiqiang Liu ·

    MileGPO:基于图的策略优化长时域LLM智能体的里程碑式推理与本地证据

    arXiv:2608.19803v1 Announce Type: cross Abstract: Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision often comes only from final rewards. Existing methods refine trajectory-level signals into step-level credits through step grouping…

  187. arXiv cs.AI TIER_1 English(EN) · Yizhe Chi, Wenyi Li, Deyao Hong, Xiaoqiu Wang, Mingju Gao, Kaisen Yang, Bingxiang He, Youjie Zheng, Calvin Xiao, Qinhuai Na ·

    AI4AI-Bench:用于递归自改进算法设计的LLM代理基准测试

    arXiv:2608.20318v1 Announce Type: new Abstract: Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inherits the improvement. That process is the training algorithm: a better objective or update rule…

  188. arXiv cs.AI TIER_1 English(EN) · Yiyang Feng, Biddut Sarker Bijoy, Niranjan Balasubramanian, Jiawei Zhou ·

    分解与传递:LLM 智能体中的跨任务技能迁移

    arXiv:2608.20274v1 Announce Type: new Abstract: Large language model (LLM) agents can induce skills from completed tasks and reuse them later to grow more capable with experience. In practice, induced skills may transfer unreliably and can even harm the agent that retrieves them.…

  189. arXiv cs.AI TIER_1 English(EN) · Yu Chen, Ruishuo Chen, Xun Wang, Zhuoran Li, Longbo Huang ·

    具有可证明双标准保证的 LLM 代理的最佳技能选择

    arXiv:2608.19993v1 Announce Type: new Abstract: Loading reusable skill documents into a bounded context window is now the primary way large language model (LLM) agents acquire task-specific capabilities, which makes skill selection a first-order determinant of task performance an…

  190. arXiv cs.AI TIER_1 English(EN) · Yiyang Luo, Yihang Jiang, Qijun Xie, Liang Lan, Lin Willian Cong, Anyi Rao, Yunya Song ·

    ReguSim:评估大型语言模型代理在金融合规中的规则接地性

    arXiv:2608.19974v1 Announce Type: new Abstract: LLM agents in financial markets may cite rules yet still submit orders that violate executable constraints or misread surveillance evidence. We introduce ReguSim, a controlled financial-compliance environment, and ReguBench, a targe…

  191. arXiv cs.AI TIER_1 English(EN) · Seongjae Kang, Taehyung Yu, Sung Ju Hwang ·

    PolicyGuide:从守护单一动作到指导合规性LLM代理的整个工作流

    arXiv:2608.19861v1 Announce Type: new Abstract: Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures arise from either forbidden actions, such as granting an ineligible change, or omitted procedural requirements, such a…

  192. Hugging Face Daily Papers TIER_1 English(EN) ·

    LLM智能体时代的图工程:从个体智能到系统智能

    Graph Engineering organizes multi-agent LLM systems through dynamic graph structures to coordinate specialized agents and manage complex, evolving tasks.

  193. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Achim Rettberg ·

    利用LLMs的常识推理能力进行多智能体协同以实现自动驾驶

    Autonomous vehicles require robust perception and decision-making capabilities to operate in diverse and unseen scenarios. While reinforcement learning and rule-based methods can provide effective control and safety mechanisms, their performance may degrade in situations requirin…

  194. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Achim Rettberg ·

    利用LLM的常识推理能力进行多智能体协同以实现自动驾驶

    Autonomous vehicles require robust perception and decision-making capabilities to operate in diverse and unseen scenarios. While reinforcement learning and rule-based methods can provide effective control and safety mechanisms, their performance may degrade in situations requirin…

  195. arXiv cs.CL TIER_1 English(EN) · Ting-Wei Li, Yuanchen Bei, Xiao Lin, Hanghang Tong ·

    超越基于LLM的推理:轻量级GNN用于智能体故障归因

    arXiv:2608.18575v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent systems (MAS) often exhibit complex failure modes, which frequently cause agents to produce incorrect outcomes. This motivates the task of Agent Failure Attribution: given a failed multi-…

  196. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Reza Zakerian ·

    LLM 智能体何时提供帮助?自动驾驶汽车边缘的截止感知混合关键性任务调度

    Autonomous vehicles offload latency-sensitive perception tasks to nearby mobile edge computing (MEC) servers, where a missed safety-critical task is unsafe rather than merely degraded. Large language models (LLMs) are increasingly proposed as adaptive, explainable schedulers, yet…

  197. Hugging Face Daily Papers TIER_1 English(EN) ·

    PolicyGuide:从守护单一动作到指导合规性LLM代理的整个工作流

    Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures arise from either forbidden actions, such as granting an ineligible change, or omitted procedural requirements, such as identification or confirmation. Runtime safegu…

  198. arXiv cs.AI TIER_1 English(EN) · Liangtao Lin, Qingang Zhang, Zhaomeng Zhu, Tianwei Zhang, Yonggang Wen ·

    面向任务的LLM智能体在关键任务基础设施运营中的线束配置

    arXiv:2608.17433v1 Announce Type: new Abstract: LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take…

  199. arXiv cs.LG TIER_1 English(EN) · Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee ·

    Agentic ESOpt:以极低的 GPU 需求微调长时域 LLM 智能体

    arXiv:2608.17310v1 Announce Type: new Abstract: Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyw…

  200. arXiv cs.AI TIER_1 English(EN) · Dayang Liang, Liyuan He, Xuan Feng, Shuxin Li, Bo An, Yunlong Liu ·

    PlanPO:面向多轮Agentic LLM的群组规划感知策略优化

    arXiv:2608.17289v1 Announce Type: new Abstract: Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactive tasks. However, most existing variants fail to distinguish advantages among successful traj…

  201. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Sandeep P. Chinchali ·

    贝叶斯伙伴建模实现LLM协调的自适应再规划

    Multi-agent Large Language Model (LLM) systems often struggle to collaborate with new teammates whose strategies shift mid-task. Because agents execute multi-step or temporally extended skills, they frequently continue executing outdated plans long after public evidence shows tha…

  202. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Yonggang Wen ·

    面向任务感知的LLM智能体在关键任务基础设施运维中的资源配置

    LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take. Existing systems often expose the same compreh…

  203. arXiv cs.AI TIER_1 English(EN) · Zhengzhao Ma. Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun ·

    从序列到结构:LLM代理的关系不确定性传播

    arXiv:2608.16002v1 Announce Type: cross Abstract: Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments. Existing UQ methods largely rely on local signals, such as token probabilities, predictive …

  204. arXiv cs.AI TIER_1 English(EN) · Teoman Kaman ·

    何时沟通:信念分布与KL散度用于多智能体强化学习中的原则性门控

    arXiv:2608.14559v1 Announce Type: new Abstract: Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when? Existing approaches either communicate at every timestep or learn a binary gate through REINFOR…

  205. arXiv cs.AI TIER_1 English(EN) · Prabhjot Singh, Bhushan Pawar ·

    幻觉滚雪球:将模型错误传播建模为多智能体LLM管道中的状态转换

    arXiv:2608.14588v1 Announce Type: new Abstract: Sequential multi-agent LLM pipelines chain specialized agents without verification at handoffs, creating a structural flaw with measurable and severe consequences. We show that hallucinations injected at Stage 1 do not merely persis…

  206. arXiv cs.AI TIER_1 English(EN) · Pierre Dantas, Lucas Cordeiro, Ehsan Nowroozi, Tihanyi Norbert ·

    迈向安全的LLM智能体:规范、验证与执行综述

    arXiv:2608.14590v1 Announce Type: new Abstract: LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use of tools. However, no existing system provides formally grounded, task-level safety guarante…

  207. arXiv cs.AI TIER_1 English(EN) · Wael Albayaydh, Rui Zhao ·

    大型语言模型代理能否理性谈判?一个用于A2A/MCP上可验证多代理交互的机制设计框架

    arXiv:2608.14613v1 Announce Type: new Abstract: Modern LLM-agent frameworks increasingly interoperate through standards such as Anthropic's Model Context Protocol (MCP) for agent-to-tool access and Google's Agent2Agent (A2A) protocol for agent delegation and negotiation. However,…

  208. arXiv cs.AI TIER_1 English(EN) · Md Fazley Rafy ·

    TwinGridShield:LLM网格代理行为的后果感知运行时授权

    arXiv:2608.15391v1 Announce Type: new Abstract: Large language model (LLM)-assisted energy-management tools can translate natural-language context into structured grid commands, but syntactic validity does not imply physical admissibility. This paper presents TwinGridShield, a mo…

  209. arXiv cs.AI TIER_1 English(EN) · Pouya Ghiasnezhad Omran, Michael Zimmermann, Duncan Cambridge, Ashmita Kapoor, Tanya Dixit ·

    Agent Gym:通过人机协作反馈实现大型语言模型智能体持续评估与进化的框架

    arXiv:2608.15591v1 Announce Type: new Abstract: Large Language Model (LLM) agents deployed in production environments face a fundamental tension: the agent's behavior is frozen at deployment time, while the business rules and edge cases it must handle continue to evolve. Existing…

  210. arXiv cs.AI TIER_1 English(EN) · Veit Laule, Jiangtao Shuai, Manfred Hauswirth, Sonja Schimmler ·

    PDDLCoder:用于LLM辅助符号规划的Agentic PDDL生成

    arXiv:2608.16637v1 Announce Type: new Abstract: LLMs remain unreliable for long-horizon planning, often generating logically inconsistent or non-applicable plans. Recent hybrid methods instead translate natural language into the Planning Domain Definition Language (PDDL), allowin…

  211. arXiv cs.AI TIER_1 English(EN) · Yuan Guo, Yilong Chen, Chao Hu, Xianghao Yu, Liang Hong, Jie Xu ·

    WARA:通过闭环大语言模型代理实现自动化无线优化研究

    arXiv:2608.14573v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly capable of tool use, code execution, artifact inspection, and iterative revision, creating new opportunities for automating scientific and engineering research. To the best of our…

  212. arXiv cs.AI TIER_1 English(EN) · Kaixiang Wang, Yidan Lin, Jiong Lou, Jie Li ·

    BRA-Audit:通过累积暴露审计点放置对LLM多智能体系统进行预算运行时审计

    arXiv:2608.14668v1 Announce Type: cross Abstract: LLM-based multi-agent systems (LLM-MAS) solve complex tasks through specialized collaboration, but inter-agent dependencies can propagate hallucinated or malicious outputs into system-level failures. Auditor agents mitigate these …

  213. arXiv cs.AI TIER_1 English(EN) · Zeyuan Li (Massachusetts Institute of Technology), Lukas Petersson (Andon Labs), Alessandro Acquisti (Massachusetts Institute of Technology), Michiel A. Bakker (Massachusetts Institute of Technology) ·

    长时域多智能体LLM商业化中的涌现式失调沟通

    arXiv:2608.14825v1 Announce Type: cross Abstract: Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elicitation ev…

  214. arXiv cs.AI TIER_1 English(EN) · Puyu Zeng, Qibing Ren ·

    超越直接访问:LLM 代理中的资源劫持

    arXiv:2608.15108v1 Announce Type: cross Abstract: Large language model agents are increasingly connected to high-value resources such as computing infrastructure, credentials, usage budgets, identities, private knowledge, communication channels, and organizational workflows. Exis…

  215. arXiv cs.AI TIER_1 English(EN) · Xiao Wang, Lu Dong, Ifeoma Nwogu, Srirangaraj Setlur, Venu Govindaraju ·

    MistyPilot:通过多智能体大语言模型技能编排实现社交机器人控制

    arXiv:2608.15549v1 Announce Type: cross Abstract: Programming small social robots from natural-language instructions requires more than invoking isolated APIs. Interactive tasks combine reactive physical behaviors with stateful social behaviors, while existing interfaces often re…

  216. arXiv cs.AI TIER_1 English(EN) · Victor Ye Dong, Reid Pryzant, Yi Liu, Jian Jiao ·

    CAPO: 约束感知提示优化用于LLM代理

    arXiv:2608.16068v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as agents that rely on system prompts to use tools and complete tasks. Such deployments impose distinct operational requirements, including appropriate tool use, concise promp…

  217. arXiv cs.CL TIER_1 English(EN) · Ruiyao Xu, Tiankai Yang, Wei-Chieh Huang ·

    HyperSkill:通过超图结构技能记忆实现自演化大型语言模型代理

    arXiv:2608.16114v1 Announce Type: new Abstract: As agentic tasks grow in complexity, LLM agents increasingly rely on experiential memory to reuse procedural knowledge across tasks. Effective memory design must jointly address what to store, how memory is structured and retrieved,…

  218. arXiv cs.LG TIER_1 English(EN) · Jiecheng Zhou, Qinghao Hu, Peng Sun, Xingcheng Zhang, Weiming Zhang ·

    Belayer:LLM 智能体强化学习训练的高效容错机制

    arXiv:2608.14635v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly trained with reinforcement learning in long-horizon, sandboxed environments. Unlike conventional RL, agentic RL couples GPU-intensive rollout engines with stateful environment con…

  219. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agentic ESOpt:以极低的GPU需求微调长时域LLM智能体

    Agentic ESOpt uses evolution strategies for scalable full-parameter fine-tuning of long-horizon LLM agents via trajectory-level reward-weighted updates and parameter-context co-evolution.

  220. Hugging Face Daily Papers TIER_1 English(EN) ·

    MUSE:一个用于理解和引导 LLM 驱动的数据科学系统的交互式元代理

    Recent advances in large language models have enabled a new class of agentic data science systems that allow users to complete complex data science workflows through natural language. Although these systems can significantly reduce manual effort, it remains difficult to diagnose …

  221. arXiv cs.AI TIER_1 English(EN) · Heming Fu, Shan Lin, Qianqian Xie, Guojun Xiong ·

    并非所有Token都相等:面向Agentic LLM系统的通胀感知路由

    arXiv:2608.13571v1 Announce Type: cross Abstract: When a language model fails to answer a query on the first attempt, an agentic system retries, consuming additional tokens each time. This retry overhead creates a gap between what a model's per-token price implies and what a full…

  222. arXiv cs.AI TIER_1 English(EN) · Ismail El Hamraoui, Sagar Jose, Nicolas Bureau, Robert Plana ·

    面向自主LLM智能体结构化漂移诊断与恢复的基于图的强化学习框架

    arXiv:2608.14109v1 Announce Type: new Abstract: Autonomous LLM agents are increasingly deployed in complex real-world workflows, yet they remain vulnerable to runtime behavioral drift, a silent deviation from the original task that can lead to irreversible side effects on externa…

  223. arXiv cs.AI TIER_1 English(EN) · Zhizhao Guan, Chen Huang, Ziming Liu, Hongru Liang, Wenqiang Lei, See-Kiong Ng, Tat-Seng Chua, Anthony G Cohn ·

    拨开迷雾:迈向在大型语言模型代理中安装和优化主动探索能力

    arXiv:2608.14339v1 Announce Type: new Abstract: We study proactive exploration in LLM agents, i.e., the ability to explore an environment to acquire information that improves future decision-making. In this regard, we first identify two fundamental bottlenecks that hinder this ca…

  224. arXiv cs.AI TIER_1 English(EN) · Yu Zhuang, Kefei Chen, Yitong Duan, Shuxin Zheng, Jian Li, Xu-Yao Zhang ·

    AgentRewind:面向长时序LLM智能体的可恢复执行

    arXiv:2608.14380v1 Announce Type: new Abstract: Many real-world tasks require LLM agents to interact with their environments over long execution horizons. Errors that occur early in execution may propagate through both the agent context and environment state, and their effects ma…

  225. arXiv cs.AI TIER_1 English(EN) · Xiaofan Zhou, Huy Nguyen, Bo Yu, Chenxi Liu, Lu Cheng ·

    多轮LLM推理的自适应停止

    arXiv:2604.01413v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) increasingly rely on multi-turn reasoning and interaction, such as adaptive retrieval-augmented generation (RAG) and ReAct-style agents, to answer difficult questions. These methods improve acc…

  226. arXiv cs.LG TIER_1 English(EN) · Ignacio D. Lopez-Miguel, Andreas Happe, J\"urgen Cito, Ezio Bartocci, Bettina K\"onighofer, Martin Tappler ·

    ATLAS:通过 LLM 指导的抽象和自动机学习发现代理策略

    arXiv:2608.14352v1 Announce Type: cross Abstract: Large Language Model (LLM)-based agents are increasingly used for complex tasks such as software testing and cybersecurity assessment. While these agents demonstrate impressive capabilities, their behavior is difficult to understa…

  227. Hugging Face Daily Papers TIER_1 English(EN) ·

    从序列到结构:LLM代理的关系不确定性传播

    RUPA models agent execution as a dependency graph to propagate uncertainty across long trajectories, improving failure detection and confidence estimation for LLM agents.

  228. arXiv cs.MA (Multiagent) TIER_1 English(EN) · M. F. Mridha ·

    Agentic Security:LLM驱动的渗透测试的工具、故障模式和设计法则的系统化

    Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools. As these systems move from demonstrations to deployed products, practitioners repeatedly encounter the same operational failures. We systematize these failures through a hands…

  229. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Travis Smith ·

    小小科学家:通过科学方法驱动的LLM智能体发现

    What happens when you teach an LLM-based agent the scientific method? Motivation: Scientific discovery emerges from cycles of hypothesis, implementation, empirical testing, and feedback. Can this process be automated? We approach automated algorithm design through the lens of the…

  230. Hugging Face Daily Papers TIER_1 English(EN) ·

    TwinGridShield:LLM网格代理行为的后果感知运行时授权

    Large language model (LLM)-assisted energy-management tools can translate natural-language context into structured grid commands, but syntactic validity does not imply physical admissibility. This paper presents TwinGridShield, a model-independent runtime authorization layer that…

  231. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Michiel A. Bakker ·

    长周期多智能体LLM商业中的涌现式错位沟通

    Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elicitation evaluations on single agents or stylized tasks. Its …

  232. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Michiel A. Bakker ·

    长时域多智能体LLM商业化中的涌现式失调沟通

    Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elicitation evaluations on single agents or stylized tasks. Its …

  233. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Michiel A. Bakker ·

    长周期多智能体LLM商业化中的涌现式错位沟通

    Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elicitation evaluations on single agents or stylized tasks. Its …

  234. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Robert Plana ·

    面向自主LLM智能体结构化漂移诊断与恢复的基于图的强化学习框架

    Autonomous LLM agents are increasingly deployed in complex real-world workflows, yet they remain vulnerable to runtime behavioral drift, a silent deviation from the original task that can lead to irreversible side effects on external systems. Existing approaches address drift at …

  235. arXiv cs.AI TIER_1 English(EN) · Chang Liu, Yuqi Zhang, Yiman Zhong, Boyi Liu, Hengjun Wang, Shuyue Wei ·

    SkillShapley: 边界自适应Shapley值用于LLM智能体中的技能步骤归因

    arXiv:2608.13173v1 Announce Type: new Abstract: Agent skills are crucial external instructions that enable language agents to execute long procedural tasks such as coding or document processing. Existing agent skills are primarily created through human manual crafting or agent ex…

  236. arXiv cs.AI TIER_1 English(EN) · Jiajun Ruan, Peiyang Li, Yukun Chen, Fengting Li, Chao Feng ·

    超越手工安全:迈向LLM智能体自进化防御

    arXiv:2608.12977v1 Announce Type: cross Abstract: The expanding operational capabilities of large language model (LLM) agents introduce sophisticated security threats. Runtime defenses have emerged as an effective approach to mitigating these risks by integrating security mechani…

  237. arXiv cs.AI TIER_1 English(EN) · Qinwu Xu, Zhuoheng Li, Jessie Salas ·

    通过代理评估和稳定性感知排序实现多模态大模型的鲁棒性检查点选择

    arXiv:2605.18852v2 Announce Type: replace-cross Abstract: Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and downstream evaluation signals are noisy. Small observed differences can be compara…

  238. arXiv cs.AI TIER_1 English(EN) · Junzhi Li, Peng He, Qirui Ji, Wei Wang, Lixiang Liu, Chuxiong Sun ·

    通过因果推理发现基于LLM的多智能体系统的有效且可解释的通信拓扑

    arXiv:2608.12921v1 Announce Type: cross Abstract: The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn communication topologies through b…

  239. arXiv cs.AI TIER_1 English(EN) · Yanwen Peng, Delvin Ce Zhang, Xi Wang, Nikolaos Aletras ·

    StateBridge:LLM多智能体系统中无训练的隐藏状态对齐用于潜在通信

    arXiv:2608.13317v1 Announce Type: new Abstract: Large language model based multi-agent systems usually communicate in text, i.e., using discrete tokens. However, text introduces a discrete bottleneck. Converting the sender's continuous hidden states into discrete tokens discards …

  240. arXiv cs.AI TIER_1 English(EN) · Zechuan Wang, Siyuan Lu, Hongxuan Zhang, Linjian Mo, Chenyi Zhuang, Leilei Gan ·

    传授幅度而非方向:多轮多步LLM代理的验证器约束信用分配

    arXiv:2608.13179v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignment conflates heterogeneous per-turn outcomes into a…

  241. arXiv cs.AI TIER_1 English(EN) · Xutao Mao, Liangjie Zhao, Xiang Zheng, Cong Wang ·

    熟能生险:自改进 LLM 智能体中的技能错位进化

    arXiv:2608.12851v1 Announce Type: new Abstract: Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input disappears. Skill evolution makes this failure measurable by…

  242. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Chuxiong Sun ·

    通过因果推断发现基于LLM的多智能体系统的有效且可解释的通信拓扑

    The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn communication topologies through black-box optimization driven solely by task-level …

  243. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Chuxiong Sun ·

    通过因果推断发现基于LLM的多智能体系统的有效且可解释的通信拓扑

    The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn communication topologies through black-box optimization driven solely by task-level …

  244. arXiv cs.AI TIER_1 English(EN) · Igor Itkin ·

    穷人的代理建模:在笔记本电脑上模拟大型LLM-代理社会

    arXiv:2608.11215v1 Announce Type: new Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the co…

  245. arXiv cs.AI TIER_1 English(EN) · Dylan Bouchard, Mohit Singh Chauhan ·

    超越单轮置信度:面向LLM智能体的轨迹自适应不确定性量化

    arXiv:2608.11552v1 Announce Type: cross Abstract: Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive …

  246. arXiv cs.AI TIER_1 English(EN) · Josef Liyanjun Chen ·

    就绪队列:界定GPU机遇并避免LLM-Agent控制中的主机往返

    arXiv:2608.12123v1 Announce Type: cross Abstract: LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enough concurrent work for GPU exe…

  247. arXiv cs.AI TIER_1 English(EN) · Junliang Liu, Ruoyu Li, Wenxin Tang, Jingyu Xiao, Zhenyu Liu, Jingheng Xu, Laizhong Cui ·

    收敛绕道劫持:基于技能的大语言模型代理中的任务保留资源放大

    arXiv:2608.12273v1 Announce Type: cross Abstract: LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning. This progressive-disclosure design exposes two sequential control points to untrusted publi…

  248. arXiv cs.CL TIER_1 English(EN) · Yutao Mou, Pengfei Yang, Zhe Yin, Zhangchi Xue, Xiaotian Luan, Dingyao Yu, Tong Zhang, Shikun Zhang, Wei Ye ·

    ToolHazard:为基于LLM的代理的安全评估和对齐扩展对抗性环境

    arXiv:2608.11878v1 Announce Type: cross Abstract: Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments,…

  249. arXiv cs.AI TIER_1 English(EN) · Ruoxi Zhao, Maziar Raissi ·

    Backtrader-Bench:使用自生成多项选择题对算法交易中的 LLM Agent 进行基准测试

    arXiv:2608.11232v1 Announce Type: cross Abstract: Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs require ground truth from actual code execution. We present Backtrader-Bench, a fram…

  250. arXiv cs.AI TIER_1 English(EN) · Ziqiang Wan, Li Gu, Zhixiang Chi, Zhi Liu, Seyed Mehdi Ayyoubzadeh, Yuanhao Yu, Yang Wang ·

    基准测试LLM裁判以评估移动代理

    arXiv:2608.11434v1 Announce Type: new Abstract: Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark for s…

  251. arXiv cs.AI TIER_1 English(EN) · Pardis Taghavi, Santosh Bhavani ·

    从数字到判断:专业大语言模型Agent与强化学习在欧洲上市房地产中的应用

    arXiv:2608.11381v1 Announce Type: new Abstract: We study whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM specialization. Larix maps a 16-lens European listed-real-estate analysis framework to eight lens-…

  252. arXiv cs.AI TIER_1 English(EN) · Dongyang Ao, Kaixiang Fang, Shijie Xu ·

    RecSys Factory: 将 LLM Agent 自主性限制在工业推荐器生命周期中的决策点

    arXiv:2608.11241v1 Announce Type: new Abstract: Deploying LLM agents into industrial recommender operations exposes a three-way tension we frame as the autonomy-determinism-efficiency trilemma: general autonomy (interpreting operator intent, generating glue code zero-shot), indus…

  253. arXiv cs.AI TIER_1 English(EN) · Gen Dong, Yanjie Gao, Liqun Li, Tianyin Xu, Yu Hua, Fan Yang ·

    Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents

    arXiv:2608.11888v1 Announce Type: new Abstract: Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, including planning, tool use, problem-solving, and validation. Prior work reported mixed results …

  254. arXiv cs.AI TIER_1 English(EN) · Touseef Hasan, Mounika Ghanta, Souvika Sarkar, Ujjwal Guin ·

    AgenticTwin:集成数字孪生的代理式LLM框架用于异常检测

    arXiv:2608.11679v1 Announce Type: new Abstract: Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems. Even with skilled operators, interpreting anomalies detected within digital twin pipelines is challenging, as the sheer complexity a…

  255. arXiv cs.AI TIER_1 English(EN) · Alexander Liss, Nicholas Desmond, Santiago Gil Gallego ·

    面向协作式对话结果的多大型语言模型(LLM)代理系统的动态治理

    arXiv:2608.11207v1 Announce Type: new Abstract: When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach…

  256. Hugging Face Daily Papers TIER_1 English(EN) ·

    就绪队列:限制GPU机遇并避免LLM-Agent控制中的主机往返

    LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enough concurrent work for GPU execution, and what changes when a GPU-computed route…

  257. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents

    Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, including planning, tool use, problem-solving, and validation. Prior work reported mixed results of agent skills: some skills improve task succes…

  258. Hugging Face Daily Papers TIER_1 English(EN) ·

    ToolHazard:为基于LLM的代理的安全评估和对齐扩展对抗性环境

    Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefi…

  259. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Ujjwal Guin ·

    AgenticTwin:集成数字孪生的代理式LLM框架用于异常检测

    Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems. Even with skilled operators, interpreting anomalies detected within digital twin pipelines is challenging, as the sheer complexity and volume of raw sensor data make thorough analy…

  260. arXiv cs.AI TIER_1 English(EN) · You Lu, Kun Zhang, Bihuan Chen, Xin Peng ·

    DOCSCHISEL:面向LLM智能体的自适应工具文档优化框架

    arXiv:2608.10037v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical grounding resource for LLM agents. Existing studies mainly focus on improving the tool-u…

  261. arXiv cs.AI TIER_1 English(EN) · Vassilis Papadopoulos, McNair Shah, Sam Zimmerman, Jack Lindsey ·

    心智病毒:多智能体LLM系统中的自我传播思想

    arXiv:2608.10218v1 Announce Type: new Abstract: AI agents are becoming more autonomous and increasingly interconnected, exposing them to new emergent risks arising from agent-to-agent interaction. One such risk is the spread of mind viruses: ideas or goals that propagate through …

  262. arXiv cs.AI TIER_1 English(EN) · Vitalii Belov, Artyom Sosedka, Andrey Sakhovskiy, Elizaveta Kovtun, Artyom Boyarskikh, Semen Budennyy ·

    LLM Agents Factory:领域特定LLM Agent的检索

    arXiv:2608.09934v1 Announce Type: cross Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deployment is often limited by the computational cost and instability associated with the…

  263. arXiv cs.AI TIER_1 English(EN) · Yiyang Li, Weixiang Sun, Tianyi Ma, Kaiwen Shi, Zheyuan Zhang, Yanfang Ye ·

    HoosierHelp:为社会服务导航进行大语言模型代理基准测试

    arXiv:2608.09946v1 Announce Type: cross Abstract: Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints. Although LLM agents offer a promising interface for conversational resource navigation, existin…

  264. arXiv cs.AI TIER_1 English(EN) · Vasundra Srinivasan ·

    一种用于生产环境LLM代理运行时架构模式选择与组合的方法

    arXiv:2605.20173v2 Announce Type: replace Abstract: Production LLM agents combine stochastic model outputs with deterministic software systems, yet the boundary between the two is rarely treated as a first-class architectural object. This paper names that boundary the stochastic-…

  265. arXiv cs.CL TIER_1 English(EN) · Xinying Cai, Minghao Guo, Jiahe Liu, Jiaojiao Han, Bangwei Guo, Yitao Long, Yuxuan Chen, Bohan Wu, Dimitris N. Metaxas, Raymond Li ·

    OpenPM:LLM投资组合管理代理的可审计即时评估

    arXiv:2608.09988v1 Announce Type: cross Abstract: Large language models are increasingly used to read markets, assess risk, and allocate capital. However, reported results for LLM trading agents can be inflated by look-ahead leakage, optimistic execution, and risk mandates that a…

  266. arXiv cs.CL TIER_1 English(EN) · Ying Yuan ·

    检测到效应不等于学会对其采取行动:LLM获取代理的奖励信噪比下限

    arXiv:2608.10441v1 Announce Type: cross Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth using…

  267. arXiv cs.CL TIER_1 English(EN) · Xiaozhe Li, Yongkang Chen, Shujian Deng, Peiji Li, Yichuan Ma, Huaxi Huang, Qiye Cai, Tianyi Lyu, Le Ma, Linyang Li, Qipeng Guo, Dahua Lin, Kai Chen ·

    InternAgentHarness: 一个可扩展的合成环境,用于增强LLM的代理能力

    arXiv:2508.08636v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly expected to act as generalist agents capable of solving complex real-world problems. Training such agents, however, requires stable and diverse environments that support repeated int…

  268. arXiv cs.CL TIER_1 English(EN) · Xiaojun Wu, Cehao Yang, Honghao Liu, Xueyuan Lin, Wenjie Zhang, Zhichao Shi, Xuhui Jiang, Chengjin Xu, Jia Li, Jian Guo ·

    Bayesian-Agent: 后验引导的技能演化跨越LLM代理工具集

    arXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these assets through heuristic reflection or raw success counts. Such updates are brit…

  269. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越单轮置信度:面向LLM智能体的轨迹自适应不确定性量化

    Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying que…

  270. Hugging Face Daily Papers TIER_1 English(EN) ·

    ToolHazard:为基于LLM的代理的安全评估和对齐扩展对抗性环境

    ToolHazard is a scalable framework that synthesizes adversarial environments to test LLM agents against indirect prompt injections, revealing vulnerabilities and improving defensive alignment.

  271. Hugging Face Daily Papers TIER_1 English(EN) ·

    就绪的群组:限制GPU机会并避免LLM-Agent控制中的主机往返

    Two studies define measurable GPU control gates for LLM-agent services by analyzing concurrent cohort scheduling and on-device routing versus host redispatch.

  272. arXiv cs.AI TIER_1 English(EN) · Rohan Bhagra, Mahantesh Halapannavar, Uddhav Bhattarai ·

    Agentic Harnesses:LLM驱动的机器人自主性验证层

    arXiv:2608.09857v1 Announce Type: cross Abstract: Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on execution rather than verifying the feasibility actions planning models propose…

  273. arXiv cs.LG TIER_1 English(EN) · Ibne Farabi Shihab, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan ·

    Stateful CARS: Exact Cross-History Reuse for Policy-Constrained LLM Agents

    arXiv:2608.08282v1 Announce Type: new Abstract: Tool-using language-model agents face constraints whose meaning changes with observations and prior actions. We study exact sampling from the model distribution conditioned on a hard stateful validator while reusing invalidity certi…

  274. arXiv cs.CL TIER_1 English(EN) · Mahesh Ramesh, Kaousheik Jayakumar, Aswinkumar Ramkumar, Pavan Thodima, Aniket Rege, Emmanouil-Vasileios Vlatakis-Gkaragkounis ·

    合作推理的火花:LLM 作为策略性 Hanabi 玩家

    arXiv:2601.18077v3 Announce Type: replace Abstract: Cooperative reasoning under incomplete information remains challenging for both humans and multi-agent systems. The card game Hanabi embodies this challenge, requiring theory-of-mind reasoning and strategic communication. We ben…

  275. arXiv cs.CL TIER_1 English(EN) · Bingzhen Liu, Xiaomeng Fan, Yuwei Wu, Zhi Gao, Mingyang Gao, Chuanhao Li, Yunde Jia ·

    超越能力边界:用于自进化LLM智能体的零阶优化

    arXiv:2608.09292v1 Announce Type: cross Abstract: Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. However, these methods struggle to learn beyond the inherent capability boundary …

  276. arXiv cs.CL TIER_1 English(EN) · Ashritha Gonuguntla ·

    重播差距:LLM代理模型切换的静态评估得分错误的世界

    arXiv:2608.08239v1 Announce Type: cross Abstract: LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents. Yet agentic routers are evaluated like single-turn routers: by replaying logge…

  277. arXiv cs.AI TIER_1 English(EN) · Jiashu He, Jinxuan Fan, Bowen Jiang, Ignacio Houine, Dan Roth, Alejandro Ribeiro ·

    SAKE:基于强化学习的复杂LLM推理结构化智能体知识外推

    arXiv:2505.15062v5 Announce Type: replace-cross Abstract: Knowledge extrapolation is the process of inferring novel information by combining and extending existing knowledge that is explicitly available. It is essential for solving complex questions in specialized domains where r…

  278. arXiv cs.AI TIER_1 English(EN) · Bohan Chen, Shivam N. Patel, Richard Hoffmann, Sam Looi, Tony Yue Yu ·

    学习协调符号工具:LLM 代理用于验证平方和证书

    arXiv:2608.00326v2 Announce Type: replace Abstract: Tool calling allows large language models (LLMs) to invoke external computation during problem solving, a useful capability in various fields including AI for mathematics. We study this setting through weighted sum-of-squares (S…

  279. arXiv cs.AI TIER_1 English(EN) · Jiyong Kwon, Ujin Jeon, Sooji Lee, Guang Lin ·

    AIVV:用于可信赖自主系统的神经符号大模型智能体集成验证与确认

    arXiv:2604.02478v2 Announce Type: replace Abstract: Deep learning models excel at detecting anomaly patterns in normal data. However, they do not provide a direct solution for anomaly classification and scalability across diverse control systems, frequently failing to distinguish…

  280. arXiv cs.AI TIER_1 English(EN) · Thassilo M. Schiepanski, Nicholas Pi\"el ·

    超越像素:探索用于 LLM 驱动的网络代理的 DOM 降采样

    arXiv:2508.04412v3 Announce Type: replace Abstract: The advent of large language models (LLMs) has sparked an evolution of autonomous web browsing agents: given a web browsing task and serialised user interface (UI) state, an LLM is expected to suggest input actions that incremen…

  281. arXiv cs.AI TIER_1 English(EN) · Gordon Dai, Weijia Zhang, Jinhan Li, Siqi Yang, Chidera Onochie lbe, Srihas Rao, Arthur Caetano, Misha Sra ·

    人工智能利维坦:通过霍布斯社会契约论视角探索大型语言模型(LLM)代理的社会进化

    arXiv:2406.14373v3 Announce Type: replace Abstract: The emergence of Large Language Models (LLMs) and advancements in Artificial Intelligence (AI) offer an opportunity for computational social science research at scale. Building upon prior explorations of LLM agent design, our wo…

  282. arXiv cs.AI TIER_1 English(EN) · Nuthakki Siva Gopala Krishna, Kanishka Jain ·

    STEMMA:一个用于评估大型语言模型自我身份一致性的对抗性多智能体框架

    arXiv:2608.08164v1 Announce Type: cross Abstract: Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of structured information and functional behavior from a large teacher model to a smaller stude…

  283. arXiv cs.AI TIER_1 English(EN) · Yijie Wang, Zhen-Yu Yin, Zhenheng Tang, Xiaowen Chu ·

    Agent-MD:针对状态化 GCMC--MD 活动的事件驱动升级选择性 LLM 干预

    arXiv:2608.07637v1 Announce Type: new Abstract: Long-running molecular simulation campaigns require repeated continuation from saved states, provenance-aware progression, adaptive assessment, and occasional interpretation of workflow conditions that cannot be resolved safely by f…

  284. arXiv cs.AI TIER_1 English(EN) · Yang Liu, Shiwei Hou, Xiyuan Chen, Yu Wang, Sen Yuan, Qirui Gan, Shao You, Feifan Chen, Wencheng Li, Shuyang Hu, Yongzhou Liu, Emma Xia, Xiaojing Lu, Hao Wang, Fan Xu, Yanfeng Li ·

    ZhuLong: 面向EDA脚本的基于执行的LLM代理,支持离线API自主探索

    arXiv:2608.07925v1 Announce Type: new Abstract: EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing LLMs fail to address. This paper presents ZhuLong, an execution-grounded LLM coding agent for PyAether and SKILL that combines AP…

  285. arXiv cs.AI TIER_1 English(EN) · Florentina Voboril, Stefan Szeider ·

    使用 LLM 代理改进约束模型

    arXiv:2608.08127v1 Announce Type: new Abstract: The runtime of Constraint Programming (CP) solvers is highly sensitive to modeling choices, such as symmetry breaking, implied constraints, global constraints, constraint reformulation, and variable representation. Improving these c…

  286. arXiv cs.AI TIER_1 English(EN) · Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, Xiaoying Xing ·

    商业竞技场:在真实市场中对标LLM智能体

    arXiv:2608.08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligati…

  287. arXiv cs.AI TIER_1 English(EN) · Liang He, Jingbo Wen, Hongyu Gu, Hao Li, Haoyu Wang, Yixiong Chen, Kangning Cui, Xilu Wang ·

    从相关性到执行效用:基于技能的大模型智能体奖励感知动态执行门控

    arXiv:2608.09168v1 Announce Type: new Abstract: Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has substantially improved skill retrieval due to the increasing skill libraries, retrieving a pl…

  288. arXiv cs.AI TIER_1 English(EN) · You Lu, Xinyu Huang, Bihuan Chen, Xin Peng ·

    SkillSentry:通过运行时保障实现 LLM Agent 的可靠技能执行

    arXiv:2608.09253v1 Announce Type: new Abstract: LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and tool use. Although skills provide reusable procedural knowledge, agents may still execute them unreliably. Even when an agent…

  289. arXiv cs.AI TIER_1 English(EN) · Fatemeh Seyedin, Adrian Weller, Jinhyuk Yun, Mahmoudreza Babaei ·

    政客、骗子和顺从的工人:LLM代理在层级博弈中的新兴行为

    arXiv:2608.09574v1 Announce Type: new Abstract: LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move from individual tools to participants in multi-agent organizations, an important qu…

  290. arXiv cs.AI TIER_1 English(EN) · Hao Sui, Simeng Qin, Jie Liao, Xiaojun Jia, Bing Chen, Yang Liu ·

    ElasticBack:通过耦合触发器-规则优化在LLM代理技能中实现隐蔽的条件后门

    arXiv:2608.09577v1 Announce Type: new Abstract: Agent skills, bundles of instructions and resources that an LLM agent loads on demand, form an emerging supply chain where a single poisoned skill can persistently compromise every agent that installs it. However, existing skill att…

  291. arXiv cs.AI TIER_1 English(EN) · Wanying Qu, Qinghua Mao, Yu Li, Jiyao Liu, Xin Zhang, Dadi Guo, Yanxu Zhu, Qingyu Liu, Leitao Yuan, Xi Lin, Shanfeng Zhu, Yanwei Fu, Jing Shao, Xia Hu, Dongrui Liu ·

    SHE:面向LLM智能体的轨迹驱动安全带演进

    arXiv:2608.09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety mechanisms often treat the harnes…

  292. arXiv cs.AI TIER_1 English(EN) · Yingpeng Ma, Jianhao Yan, Bei Shi, Ka Hou Kam, Runnan Wang, Xuebo Liu, Yulong Chen, Yue Zhang, Derek F. Wong ·

    大型语言模型代理能否坚持剧本?一项用于交互式叙事中长时程一致性的基准测试

    arXiv:2608.08160v1 Announce Type: cross Abstract: The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the critical challenge of maintaini…

  293. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Ying Yuan ·

    检测到效应不等于学会对其采取行动:LLM获取代理的奖励信噪比下限

    Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth using. Our thesis is a distinction that is easy to miss…

  294. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Francisco León Zúñiga Bolívar ·

    并非“大一统”:中国前沿LLM智能体合作博弈中的实验室级分歧

    Does the cooperative bias documented for Western frontier LLM agents extend to a different alignment lineage, and should the Chinese models that embody it be treated as a single bloc or as distinct laboratories? We study four frontier-tier Chinese models - DeepSeek V4 Pro, Qwen3-…

  295. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越能力边界:用于自进化 LLM 智能体的零阶优化

    Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. However, these methods struggle to learn beyond the inherent capability boundary of the agents, since the agents cannot sample corr…

  296. Hugging Face Daily Papers TIER_1 English(EN) ·

    从相关性到执行效用:基于技能的大语言模型智能体的奖励感知动态执行门控

    Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has substantially improved skill retrieval due to the increasing skill libraries, retrieving a plausible skill bundle does not guarantee that exe…

  297. arXiv cs.LG TIER_1 English(EN) · Elizaveta D. Moskovskaya, Anton D. Moscowsky ·

    具有多智能体控制和LLM自动场景生成的机器人导引头

    arXiv:2509.10317v2 Announce Type: replace-cross Abstract: The article describes the development of a hybrid social robot control architecture to overcome the limitations of traditional approaches, where behavior scripts manually synchronize the robot's actions and text, and exist…

  298. arXiv cs.AI TIER_1 English(EN) · Jiacheng Miao, Jin Mu, Guanhua Chen, James Zou ·

    Fisher-R1:训练LLM智能体进行可靠的假设检验

    arXiv:2608.07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can inspect datasets, generate code, and produce analyses e…

  299. arXiv cs.CL TIER_1 English(EN) · Mingguang Chen, Licheng Wang, Bo Qu ·

    地平线鸿沟:长周期LLM智能体的规划、记忆、执行、训练与评估

    arXiv:2608.06663v1 Announce Type: new Abstract: Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour tasks: losing track of earlier decisions, declaring half-finished work done, or…

  300. arXiv cs.AI TIER_1 English(EN) · Karolina Rudnicka, Thomas Stephan Juzek ·

    超越“人工智能语言”:论证LLM输出的语域特性

    arXiv:2608.06589v1 Announce Type: cross Abstract: While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that this perspective coexists with distinct, model-specific linguistic signatures akin to human id…

  301. Hugging Face Daily Papers TIER_1 English(EN) ·

    商业竞技场:在真实市场中对标LLM智能体

    Business Arena evaluates LLM agents running a realistic cross-border shop, revealing large performance gaps versus human strategies and enabling detailed attribution of business decisions.

  302. Hugging Face Daily Papers TIER_1 English(EN) ·

    大型语言模型代理能否坚持剧本?面向交互式叙事中长时程一致性的基准测试

    The study introduces a benchmark and formalizes narrative commitment preservation to evaluate long-horizon logical consistency in interactive storytelling with large language models.

  303. arXiv cs.AI TIER_1 English(EN) · Wuya Chen, Yihao yang, Yang Cao, Yue Lin ·

    CodeGrep:一个用于LLM编码代理的RL训练检索代理

    arXiv:2608.05886v1 Announce Type: cross Abstract: Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the file to patch, rather than patching it. On SWE-Bench Verified, a 30B OpenHands agent average…

  304. arXiv cs.AI TIER_1 English(EN) · Jialuo Chen, Minghe Wang, Lingqi Jiang, Jianan Ma, Xinhao Deng, Xiaohu Du, Ruixiao Lin, Yunhao Feng, Linkang Du, Jingyi Wang ·

    SkillTrace:LLM-Agent 技能复用中的多轨迹溯源审计

    arXiv:2608.05204v1 Announce Type: new Abstract: LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code, tools, references, and operational workflows. As skills become marketplace artifacts, auditin…

  305. arXiv cs.AI TIER_1 English(EN) · Jie Wu, Ming Gong, Feixiang Cheng, Qinqin Zhao ·

    EcoAgent-Bench:评估预算受限LLM代理的经济决策能力

    arXiv:2608.05519v1 Announce Type: new Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local lookup, broad search, composite research tool, stronger model, or human escalation i…

  306. arXiv cs.AI TIER_1 English(EN) · Linfang Shang, Ming Xu, Yiding Sun, Tianle Xia, Lingxiang Hu, Lan Xu, Ning Zheng ·

    当自进化适得其反:预提交门控以防止大型语言模型智能体中的技能污染

    arXiv:2608.05810v1 Announce Type: new Abstract: Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not monotonic: past a critical pool size, newly added skills degrade performance instead of impr…

  307. arXiv cs.AI TIER_1 English(EN) · Zihan Xu, Haolin Tian, Hai Jiang ·

    多智能体LLM系统中推理时并行性的双层视角

    arXiv:2608.05791v1 Announce Type: cross Abstract: Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inference, and their execution strategies directly affect system accuracy, latency, and computation…

  308. arXiv cs.CL TIER_1 English(EN) · Xuying Ning, Dongqi Fu, Tianxin Wei, Hanqing Zeng, Yuanchen Bei, Bingxuan Li, Zihao Li, Qifan Wang, Xiang Shen, Yifan Wu, Jiayi Liu, Hong Li, Yinglong Xia, Xiangjun Fan, Hanghang Tong, Jingrui He ·

    EvoHarness-RL:为长时序LLM代理学习自进化运行时线束

    arXiv:2608.05446v1 Announce Type: cross Abstract: Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, invoke tools, verify outcomes, and reuse experience across interactions. However, effective harness use raises two coupled …

  309. Hugging Face Daily Papers TIER_1 English(EN) ·

    CodeGrep:一个用于LLM编码代理的RL训练检索代理

    Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the file to patch, rather than patching it. On SWE-Bench Verified, a 30B OpenHands agent averages 23 rounds and 631K tokens per resolved issue, wi…

  310. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Hai Jiang ·

    多智能体LLM系统中推理时并行性的双层视角

    Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inference, and their execution strategies directly affect system accuracy, latency, and computational cost. Parallel execution provides a means to im…

  311. arXiv cs.AI TIER_1 English(EN) · Junnan Liu, Linhao Luo, Thuy-Trang Vu, Gholamreza Haffari ·

    EASy:迈向高效的基于LLM的代理系统

    arXiv:2608.04588v1 Announce Type: cross Abstract: Agentic systems have emerged as a promising paradigm for solving complex tasks by coordinating specialized LLM-based agents. However, most existing systems primarily optimize task success while giving limited consideration to exec…

  312. arXiv cs.AI TIER_1 English(EN) · J. de Curt\`o, I. de Zarz\`a ·

    网络物理系统中大型语言模型(LLM)智能体的规划策略的战略评估

    arXiv:2608.04265v1 Announce Type: cross Abstract: Evaluations of LLM planning agents largely ask whether a task succeeds or a declared plan is followed. In strategic cyber-physical systems, a stronger question is whether the planning architecture remains appropriate after autonom…

  313. arXiv cs.AI TIER_1 English(EN) · Peichun Hua, Haoxuan Xu, Mengyuan Li ·

    行为技能重构:从LLM智能体技能中重构隐藏功能

    arXiv:2608.04192v1 Announce Type: cross Abstract: Closed source agent skills may encode proprietary instructions, scripts, constants, and data. Providers may offer their capabilities as services while keeping the underlying packages hidden. Prior work focuses on prompt injection …

  314. arXiv cs.AI TIER_1 English(EN) · Atul Anand, Sourav Chattaraj ·

    使用 Canary Tools 诊断 LLM 代理中的工具选择推理

    arXiv:2608.04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) tool set, each engineered to probe one specific tool-…

  315. arXiv cs.AI TIER_1 English(EN) · Wenxiao Zhao, Dong Liu, Kaiyi Xu, Feng Liu, Zhen Zhao, Fei Ben, Shu Wang, Wenhao Li, Yingnian Wu, Fenghua Ling, Haobo Li, Lei Bai ·

    A-SR:用于符号回归的分层协调的自演化代理LLM

    arXiv:2608.04872v1 Announce Type: cross Abstract: Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses heterogeneous search failures into a scalar score and a single prompt. We …

  316. arXiv cs.CL TIER_1 English(EN) · Jinyi Han, Yuanjian Xu, Ying Liao, Xinyi Wang, Zishang Jiang, Zixiang Di, Fanyang Lu, Zhichao Hu, Yanghua Xiao ·

    Skill-Use:大型语言模型能否在代理式框架中实际使用技能?

    arXiv:2608.04828v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed. Existing evaluations mostly judge the quality of a skill or its co…

  317. Hugging Face Daily Papers TIER_1 English(EN) ·

    A-SR:用于符号回归的分层协调的自演化代理大型语言模型

    Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses heterogeneous search failures into a scalar score and a single prompt. We propose A-SR, a self-evolving agentic framework th…

  318. Hugging Face Daily Papers TIER_1 English(EN) ·

    使用 Canary Tools 诊断 LLM 代理中的工具选择推理

    Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) tool set, each engineered to probe one specific tool-selection weakness. A six-type taxonomy (semanti…

  319. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Wataru Toyokawa ·

    LLM智能体中基于声誉的合作的出现

    Can cooperation among large language model (LLM) agents be evolutionarily stable against free-rider invasion? We study an indirect reciprocity donation game where LLM agents observe behavioral traces and donate on a continuous scale. Strategies, represented as natural language pr…

  320. arXiv cs.AI TIER_1 English(EN) · Unggi Lee, Sookbun Lee, Yeil Jeong, Eunjoo Lee, Minchul Shin, Hoilym Kwon ·

    EduClaw-Bench:用于具有模拟学习者的教学 LLM 代理的长视野基准测试

    arXiv:2608.03206v1 Announce Type: cross Abstract: Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a l…

  321. arXiv cs.CL TIER_1 English(EN) · Ming Shen, Chao Shang, Sadat Shahriar, Devang Kulshreshtha, Yi Zhang, Sandesh Swamy, Yanjun Qi ·

    基于LLM的多智能体系统中作为收敛压力的关系先验

    arXiv:2608.03239v1 Announce Type: new Abstract: Large language model-based multi-agent systems (LLM-MAS) are designed through roles, debate protocols, and aggregation rules. These choices create implicit social expectations: agents may be expected to trust, challenge, defer to, o…

  322. arXiv cs.AI TIER_1 English(EN) · Yue Yao, Shengyuan Wang, Xin Chen, Minke Zhang, Jia He, Bingjun Luo, Tom Gedeon ·

    SkillTrace:遍历查询-技能图谱以实现可组合的LLM代理

    arXiv:2608.02356v2 Announce Type: replace Abstract: Large language model agents increasingly solve complex tasks by composing reusable skills from a library. To address this, the key challenge is not merely to retrieve individually relevant skills, but to identify a complete and …

  323. arXiv cs.AI TIER_1 English(EN) · Qiming Shi, Yibo Dou, Jiawen Zhu, Yulong Tao, Linbo Jin, Zhaolu Kang, Yunfan Zhou, Di Weng ·

    SKILL-KD: 对比式技能蒸馏用于LLM智能体

    arXiv:2607.28048v2 Announce Type: replace Abstract: Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often treat skills as experience summaries, memory entries, or direct summaries of su…

  324. arXiv cs.AI TIER_1 English(EN) · Tianyi Guan, Yiding Wang, Haotong Yang, Siyuan Cao, Shirui Liu, Yi Hu, Jiaqi Li, Muhan Zhang ·

    ContinualSkillBench:大型语言模型代理能否真正进化其能力?

    arXiv:2608.03874v1 Announce Type: new Abstract: Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve…

  325. arXiv cs.AI TIER_1 English(EN) · Zian Zhai, Xingyu Tan, Gaowang Zou, Xiaoyang Wang, Wenjie Zhang ·

    HyperAgent:在工具模式超图上进行规划和行动,用于工具使用 LLM 代理

    arXiv:2608.02650v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external tools to complete complex real-world tasks. However, reliable tool-use planning remains challenging due to the limitations of implicit reasoning and the evolving nature…

  326. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Sho Akiyama ·

    持续改进与并行自主探索:用于搜索大型解空间的LLM-Agent框架

    We present a framework that gives LLM agents two mechanisms for searching large solution spaces autonomously. First, a leaderboard scored on held-out data acts as a reward signal that drives each agent to refine its solutions over repeated submissions, a loop that operates even w…

  327. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Renato Figueiredo ·

    CURATE:利用 LLM 代理来组合、编目和部署可复现的工作流

    Agentic code generation has shown promise in automating and accelerating software development by utilizing Large Language Models (LLMs) to generate, test, and deploy code. For engineers and scientists, such systems have the potential to accelerate the development of applied and s…

  328. arXiv cs.MA (Multiagent) TIER_1 English(EN) · I. de Zarzà ·

    网络物理系统中大型语言模型(LLM)智能体的规划策略的战略评估

    Evaluations of LLM planning agents largely ask whether a task succeeds or a declared plan is followed. In strategic cyber-physical systems, a stronger question is whether the planning architecture remains appropriate after autonomous participants respond and physics constrains th…

  329. Hugging Face Daily Papers TIER_1 English(EN) ·

    ContinualSkillBench:大型语言模型代理能否真正进化其能力?

    Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge this gap, …

  330. Hugging Face Daily Papers TIER_1 English(EN) ·

    基于LLM的多智能体系统中作为收敛压力的关系先验

    Large language model-based multi-agent systems (LLM-MAS) are designed through roles, debate protocols, and aggregation rules. These choices create implicit social expectations: agents may be expected to trust, challenge, defer to, or collaborate with peers. We study the effects o…

  331. Hugging Face Daily Papers TIER_1 English(EN) ·

    SKILL-KD: 对比式技能蒸馏用于LLM代理

    Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often treat skills as experience summaries, memory entries, or direct summaries of successful demonstrations. This creates a mismatch for…

  332. Hugging Face Daily Papers TIER_1 English(EN) ·

    ContinualSkillBench:大型语言模型代理能否真正进化其能力?

    Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge this gap, …

  333. arXiv cs.CL TIER_1 English(EN) · Zhenyu Zhang, Zhichao Cao ·

    TokTier:Agentic LLM服务的精确状态化Token化

    arXiv:2607.29678v1 Announce Type: new Abstract: LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, which resubmit a long transcript after each small tool result, and reuse is hard …

  334. arXiv cs.AI TIER_1 English(EN) · Tianyu Huai, Tingshuo Fan, Xinchi Chen, Yining Zheng, Yuxin Wang, Shuang Chen, Jie Zhou, Xuanjing Huang ·

    AgentHPOBench:一个用于评估LLM代理作为顺序超参数优化器的基准

    arXiv:2607.29626v1 Announce Type: new Abstract: As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmarks typically focus on static code generation, paper r…

  335. arXiv cs.AI TIER_1 English(EN) · Qiming Shi, Yulong Tao, Linbo Jin, Zhaolu Kang, Yibo Dou, Jiawen Zhu, Tianjun Pan, Shaokang Fu, Chengyu Wang, Siyue Li, Yaping Cheng, Di Weng, Chengfu Huo ·

    MerchantBench:对标电商运营中LLM Agent的长期一致性

    arXiv:2607.28956v1 Announce Type: new Abstract: Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to p…

  336. arXiv cs.AI TIER_1 English(EN) · Duo Xu, Faramarz Fekri ·

    NeSyFS:面向部分可观测环境下的LLM智能体的神经符号快慢思维框架

    arXiv:2607.28942v1 Announce Type: new Abstract: Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retrieval-augmented generation, and scientific discovery. In these settings, agents must act based o…

  337. Hugging Face Daily Papers TIER_1 English(EN) ·

    GISAgentBench:一个从业者来源的评估LLM智能体在GIS任务上表现的基准

    Geographic Information System (GIS) professionals rely on multi-step spatial analysis workflows to support decision-making in urban planning, disaster response, and environmental monitoring. The process is tedious, time-consuming, and error-prone. While recent large language mode…

  338. arXiv cs.AI TIER_1 English(EN) · Zhilun Zhou, Jianghao Yu, Yuming Lin, yongjun yang, Sun Yongquan, Depeng Jin, Yong Li ·

    UrbanDS:一个图引导的大语言模型多智能体系统,用于数据密集型城市任务

    arXiv:2607.26724v1 Announce Type: new Abstract: Large language model (LLM) agents have been widely applied in automating data science tasks. However, existing methods typically rely on a limited set of provided datasets, and they face challenges in data-intensive scenarios that r…

  339. arXiv cs.CL TIER_1 English(EN) · I. Kennedy, T. Kennedy ·

    富达并非安全:轻度压缩的大语言模型在无数据质量保障的代理执行中通过所有已发明步骤

    arXiv:2607.28196v1 Announce Type: new Abstract: Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original, downstream accuracy (for example MMLU) inside a confidence interval, and data-fr…

  340. arXiv cs.CL TIER_1 English(EN) · Sebastian Pohl, Harsh Mehta, Pranav Mambayil, Abdul Ghafoor, Franziska Lesigang, Yufang Hou, Christian Hilbe ·

    大型语言模型在受控环境中难以模拟人类信念更新

    arXiv:2607.28347v1 Announce Type: new Abstract: LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been tested directly. We test whether six LLMs can simulate individual human belief u…

  341. arXiv cs.LG TIER_1 English(EN) · Jiwon Jang, Kisu Yang, Heuiseok Lim, Hyunwoo Park ·

    分数持平,失败加剧:错误预算如何掩盖量化大模型代理的损害

    arXiv:2607.27275v1 Announce Type: new Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless. We test this claim for multi-turn, tool-calling agents, where it now matters most. On $\tau^2$-bench, across two open-weight model families in den…

  342. arXiv cs.LG TIER_1 English(EN) · Cong Li, Peixi Peng, Yisen Zhao, Xinyu Hu, Shudong Liu, Zhan Su, Zhuojian Li ·

    TAPO:LLM智能体的感知策略优化

    arXiv:2607.27973v1 Announce Type: new Abstract: Recently, Reinforcement Learning (RL) has emerged as a crucial paradigm for the post-training of Large Language Model (LLM) agents. However, existing methods predominantly rely on sparse task rewards for policy optimization, failing…

  343. arXiv cs.AI TIER_1 English(EN) · Marylou Fauchard, Florian Carichon, Margarida Carvalho, Golnoosh Farnadi ·

    更多欺骗:混合动机LLM多智能体系统中的目标不对齐

    arXiv:2607.26120v1 Announce Type: new Abstract: Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric information and strategic deception due to conflicting or hidden objectives. In t…

  344. arXiv cs.AI TIER_1 English(EN) · Huixiang Zhang, Mahzabeen Emu ·

    潜在通道是否真的在通信?对潜在多智能体LLM的因果审计

    arXiv:2607.26773v1 Announce Type: new Abstract: Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-r…

  345. Hugging Face Daily Papers TIER_1 English(EN) ·

    MerchantBench:对标电商运营中LLM智能体长期一致性的基准测试

    Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended hori…

  346. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Rahul Rachuri ·

    LLM智能体能否进行有竞争力的定价?面向智能体商业的动态多属性拍卖基准

    Agentic commerce is moving from concept to deployed infrastructure: payment networks, retailers, and AI platforms are setting the stage for agents to transact on behalf of merchants and consumers. Yet whether the LLMs behind these agents can price competently in real markets, whe…

  347. Hugging Face Daily Papers TIER_1 English(EN) ·

    一人、N个智能体:在校准不当、相关置信度下的LLM智能体集群的审计预算分配

    A single human must audit $N$ LLM agents under a budget of $B \ll N$ audits per round, guided by self-reported confidence that may be adversarially miscalibrated and by correlated errors. We model this as budgeted noisy inspection over a two-level Gaussian copula and locate the m…

  348. Hugging Face Daily Papers TIER_1 English(EN) ·

    富达并非安全:温和压缩的LLM在无数据质量保障的代理执行中通过所有程序步骤

    Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original, downstream accuracy (for example MMLU) inside a confidence interval, and data-free output-fidelity signals that compare the comp…

  349. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Alois Knoll ·

    扩展LLM驱动的多智能体系统:设计原则与架构可扩展性分析

    LLM-based multi-agent systems have the potential to enable collective intelligence and scale toward solving highly complex tasks through coordinated ensembles of specialized agents. However, despite their theoretical potential, the architectural design space remains largely non-s…

  350. arXiv cs.LG TIER_1 English(EN) · Yicheng Feng, Yan Zhang, Yan Cheng, Wei Qi ·

    分数不是决策:LLM代理中工具获取的成本感知停止

    arXiv:2607.27083v1 Announce Type: new Abstract: As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental tool-selection challenge: acquiring too few tools leaves the task under-informed, w…

  351. arXiv cs.CL TIER_1 English(EN) · Jingxing Wang, Chenyu Zhou, Zhihui Fu, Jun Wang, Weiwen Liu, Weinan Zhang, Jianghao Lin ·

    即时技能:LLM智能体的测试时自适应技能合成

    arXiv:2605.16986v2 Announce Type: replace Abstract: Additional test-time compute can give LLM agents access to more past experience, yet expanding the context or adding rollouts does not necessarily yield greater agent capability. We call this challenge test-time compute-to-capab…

  352. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Tian Lan ·

    通过合作-义务耦合审计涌现式LLM-Agent协作

    LLM-agent systems can solve complex tasks through dynamic self-organization and emergent cooperation. Auditing this process is essential because plausible intermediate or final outputs can conceal incomplete or unsupported work and poorly allocated responsibility, ultimately comp…

  353. Hugging Face Daily Papers TIER_1 English(EN) ·

    潜在通道是否真的在通信?对潜在多智能体LLM的因果审计

    Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-relevant information. End-task performance alone …

  354. arXiv cs.AI TIER_1 English(EN) · Debjyoti Paul ·

    上下文组装作为控制变量:一种基于控制理论的冻结LLM代理策略视角

    arXiv:2607.25408v1 Announce Type: new Abstract: A growing body of 2026 work applies control theory to LLM agents: Lyapunov-certified stability for tool-mediated controllers (Prinos et al., "Stable Agentic Control", 2026), sample-complexity bounds for sparse policies over massive …

  355. arXiv cs.AI TIER_1 English(EN) · Debjyoti Paul ·

    一个控制系统、一个数据集以及一种让冻结的 LLM 代理学习某个领域的方法

    arXiv:2607.25415v1 Announce Type: new Abstract: Production LLM agents are increasingly assembled from a frozen model wrapped in a harness: a prompt template, a tool set, a memory/retrieval layer, a planning strategy, and a verification policy. Two 2026 systems, Meta-Harness (Lee …

  356. arXiv cs.AI TIER_1 English(EN) · Mingwei Zheng, David OBrien, Siwei Cui, Pardis Pashakhanloo, Rajdeep Mukherjee, Myeongsoo Kim, Sachit Kuhar ·

    CORVUS:通过底层同步优化和缩减上下文以用于LLM编码代理

    arXiv:2607.22711v1 Announce Type: cross Abstract: LLM coding agents operate by constructing trajectories that accumulate reasoning, tool calls, and results to enable multi-step decision-making. However, the conventional append-only trajectory architecture found in practice tightl…

  357. arXiv cs.AI TIER_1 English(EN) · Yan Zhang, Shibo Li ·

    ConsistencyGate:通过自洽性准入控制防止 LLM Agent 中的记忆污染

    arXiv:2607.22962v1 Announce Type: new Abstract: LLM agents that operate over many turns accumulate facts in an external memory store and reuse them as premises for downstream reasoning. A hallucinated fact written at one step therefore persists as a false premise for every subseq…

  358. arXiv cs.AI TIER_1 English(EN) · Adhyyan Narang, Artin Tajdini, Claire Zhang, Jamie Morgenstern ·

    LLM微调隐藏行为的推理时共识缓解方法

    arXiv:2607.23394v1 Announce Type: new Abstract: Recent work shows that fine-tuning language models on even a small amount of poisoned data can install targeted misbehavior, and ostensibly benign data can transmit hidden preferences that generalize broadly. Standard defenses, such…

  359. arXiv cs.AI TIER_1 English(EN) · Yihui Zhang (Beihang University), Tianyu Wo (Beihang University), Jinghao Wang (Beihang University), Xiaoyang Sun (University of Leeds), Menghao Zhang (Beihang University), Cangzhou Yuan (Beihang University), Li Li (Beihang University), Chunming Hu (Beih… ·

    SpecBox:用于高效LLM代理服务的推测性沙盒调度

    arXiv:2607.23933v1 Announce Type: cross Abstract: As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces a fundamental tension between resource utilization and interactive tail latency…

  360. arXiv cs.CL TIER_1 English(EN) · Junming Chen, Junyang Jiang, Xu Chen, Zibo Liang, Kai Zheng ·

    DBA-Bench:一个用于基于 LLM 的数据库操作代理的生产级保真度基准

    arXiv:2607.22165v1 Announce Type: cross Abstract: LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation and production operations: live-environment fidelity (multi-turn read-write inter…

  361. arXiv cs.AI TIER_1 English(EN) · Junchi Liao ·

    审计LLM代理行动选择中的来源敏感性

    arXiv:2607.20827v1 Announce Type: new Abstract: LLM agents choose tools and arguments from context that mixes user requests, tool outputs, retrieved records, memory, and untrusted text. Evidence can be relevant without being authorized to determine a decision, so a correct action…

  362. arXiv cs.AI TIER_1 English(EN) · Mohamed Jouini ·

    面向基础设施即代码生成的代理式大语言模型的先验验证器评估

    arXiv:2607.20478v1 Announce Type: cross Abstract: Infrastructure-as-Code (IaC) generation from natural language requires satisfying provider schemas, dependency planning, and organizational policy constraints, not merely producing syntactically plausible configurations. We presen…

  363. arXiv cs.AI TIER_1 English(EN) · Jerzy Kami\'nski, Ilya Galyukshev, Artem Kuznetsov, Sergey Chuprin, Kirill Redko, Aidar Shumbalov, Anna Kalyuzhnaya ·

    DynamicMCPBench:一个针对实时MCP服务器上LLM代理的、基于轨迹的、效果评分的基准测试

    arXiv:2607.20531v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score the final answer or a fixed "ground-truth" list of tools, both of which are fragil…

  364. arXiv cs.AI TIER_1 English(EN) · Aarushi Singh ·

    防护栏成替罪羊:审计工具增强型LLM代理中不忠诚的安全拒绝

    arXiv:2607.19449v1 Announce Type: cross Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure failures and HTTP 200 responses with empty, null, or malformed payloads largely…

  365. arXiv cs.AI TIER_1 English(EN) · Elias Hossain, Md Mehedi Hasan Nipu, Tasfia Nuzhat Ornee, Rajib Rana, Niloofar Yousefi ·

    NEXUS:面向使用工具的大语言模型代理的结构化运行时安全

    arXiv:2607.19356v1 Announce Type: new Abstract: Tool-using LLM agents increasingly execute high-impact actions, making runtime safety monitoring essential. We present NEXUS (Neural EXecution Utility and Safety), a structured-plan safety monitor that applies a formal intervention …

  366. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Zhonghao Hou ·

    并非物以类聚:LLM 智能体中的基于个性的伴侣选择

    Multi-agent LLM systems increasingly let one agent choose which other agents to work with, and agents are increasingly given personalities through personas. We test whether Big Five personality alone influences partner selection when capability is explicitly held constant. Host a…

  367. arXiv cs.LG TIER_1 English(EN) · Thomas Carta, Cl\'ement Romac, Loris Gaven, Pierre-Yves Oudeyer, Olivier Sigaud, Sylvain Lamprier ·

    HERAKLES:面向开放式LLM代理的层级技能编译

    arXiv:2508.14751v2 Announce Type: replace Abstract: We study goal-conditioned reinforcement learning in partially observable environments with sparse rewards and large, structured goal spaces. In such settings, complex goals often require composing simpler skills, but learning th…

  368. arXiv cs.AI TIER_1 English(EN) · Philipp J. Schneider, Lin Tian, Marian-Andrei Rizoiu ·

    学习交友:指导大型语言模型代理形成涌现的社交关系

    arXiv:2510.19299v2 Announce Type: replace Abstract: Can large language model (LLM) agents reproduce the complex social dynamics that characterize human online behavior -- shaped by homophily, reciprocity, and social validation -- and what memory and learning mechanisms enable suc…

  369. arXiv cs.AI TIER_1 English(EN) · Daisuke Kikuta ·

    AI旅游会议:LLM代理的团队旅行规划

    arXiv:2607.18806v1 Announce Type: new Abstract: This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Language Model (LLM)-based agents. The agents are instantiated with distinct personas and collaboratively seek an itinerary that satisf…

  370. arXiv cs.AI TIER_1 English(EN) · Artem Maryanskyy, Dmitry Budnikov, Alibek T. Kaliyev ·

    当代理意见不合时:多代理LLM管道中的选择瓶颈

    arXiv:2603.20324v2 Announce Type: replace-cross Abstract: Multi-agent LLM pipelines produce contradictory evidence on whether team diversity improves output quality: heterogeneous Mixture-of-Agents teams outperform single models, yet homogeneous Self-MoA teams consistently win un…

  371. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Daisuke Kikuta ·

    AI旅游会议:LLM代理的团队旅行规划

    This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Language Model (LLM)-based agents. The agents are instantiated with distinct personas and collaboratively seek an itinerary that satisfies their constraints and preferences through na…

  372. arXiv cs.AI TIER_1 English(EN) · Sumit Verma, Pritam Prasun, Pritish Kumar ·

    RAIL Guard:弥合LLM智能体负责任AI的评估到补救差距

    arXiv:2607.16215v1 Announce Type: new Abstract: Existing guardrail systems for large language model agents operate as binary classifiers that block unsafe content, leaving organizations to discard failing outputs and retry from scratch. We introduce RAIL Guard, a closed-loop resp…

  373. arXiv cs.LG TIER_1 English(EN) · YanZhao Zheng, ZhenTao Zhang, Chao Ma, YuanQiang Yu, JiHuai Zhu, Yong Wu, Tianze Xu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu ·

    SkillRouter: 规模化 LLM Agent 的技能路由

    arXiv:2603.22455v5 Announce Type: replace Abstract: Reusable skills let LLM agents package task-specific procedures, tool affordances, and execution guidance into modular building blocks. As skill ecosystems grow to tens of thousands of entries, exposing every skill at inference …

  374. arXiv cs.AI TIER_1 English(EN) · Shijun Li, Hilaf Hasson, Joydeep Ghosh ·

    OMAC:LLM驱动的多智能体协作的整体优化框架

    arXiv:2505.11765v5 Announce Type: replace-cross Abstract: Agents powered by advanced large language models (LLMs) have demonstrated impressive capabilities across diverse complex applications. Recently, Multi-Agent Systems (MAS), wherein multiple agents collaborate and communicat…

  375. arXiv cs.AI TIER_1 English(EN) · Jing-Jing Li, Jianfeng He, Chao Shang, Devang Kulshreshtha, Xun Xian, Yi Zhang, Hang Su, Sandesh Swamy, Yanjun Qi ·

    STAC:无辜的工具如何为LLM代理形成危险的链条

    arXiv:2509.25624v3 Announce Type: replace-cross Abstract: As LLMs advance into autonomous agents with tool-use capabilities, they introduce security challenges that extend beyond traditional content-based LLM safety concerns. This paper introduces Sequential Tool Attack Chaining …

  376. arXiv cs.AI TIER_1 English(EN) · Roshan Klein-Seetharaman, Daniel Wang, Andrew Xu ·

    Lomekwi:LLM智能体中的资源受限工具发现

    arXiv:2607.16961v1 Announce Type: new Abstract: Existing tool-use benchmarks report a single success rate for complex, multistep tasks. Inspired by ideas from cognitive science, we distinguish tool use from tool discovery and decompose the latter into curiosity (the model's abili…

  377. arXiv cs.AI TIER_1 English(EN) · Runming He, Zhen Hao Wong, Hao Liang, Zimo Meng, Chengyu Shen, Xiaochen Ma, Wentao Zhang ·

    DataFlow-Harness: 用于构建可编辑 LLM 数据管道的、基于事实的代码代理平台

    arXiv:2607.16617v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. We call this…

  378. Hugging Face Daily Papers TIER_1 English(EN) ·

    NexForge:通过面向需求的任务合成扩展 LLM 的代理能力

    Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual substrate engineering, each new domain demands a bespoke pipelin…

  379. Hugging Face Daily Papers TIER_1 English(EN) ·

    验证、修复、重复,还是停止?LLM代理中嘈杂的验证-修复循环的鲁棒停止策略

    Verify-repair loops are a standard means for large language model (LLM) agents to correct faulty plans in code generation, mathematical reasoning, and tool use. When both the verifier and the repairer are noisy, repair can damage already-correct plans, and reported acceptance kee…

  380. arXiv cs.AI TIER_1 English(EN) · Mingxin Li, Enge Song, Yueshang Zuo, Xiaodong Liu, Rong Wen, Qiang Fu, Gianni Antichi, Jian He, Jing Tie, Zhou Shao, Xiaobo Xue, Xiong Xiao, Luyao Zhong, Shaokai Zhang, Jiangu Zhao, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Changgang Zheng, Zihao Fan, Haon… ·

    云端可扩展 LLM 代理工具访问

    arXiv:2607.15593v1 Announce Type: cross Abstract: LLM agents increasingly rely on tool calling to act on external systems, and the Model Context Protocol (MCP) has quickly become its de facto interface. Operating MCP at cloud scale, however, becomes difficult. On the tool provide…

  381. arXiv cs.CL TIER_1 English(EN) · Yanze Wang, Pengfei Yao, Tianyi Sun, Chuanrui Hu, Yan Xiao, Yunyun Han, Jun Sun, Yafeng Deng ·

    SkillCorpus:整合和评估现实世界 LLM Agent 的开放技能生态系统

    arXiv:2607.15557v1 Announce Type: new Abstract: Agent skills, SKILL.md files that package reusable procedural knowledge for an LLM agent, are a popular mechanism for extending agent capabilities. Public repositories now host them in large and growing numbers, yet these artifacts …

  382. Hugging Face Daily Papers TIER_1 English(EN) ·

    穷人的代理建模:在笔记本电脑上模拟大型LLM-代理社会

    Replacing individual LLM agents with low-parameter surrogates fitted from cheap queries enables scalable society simulations, with validity predicted by an interaction-order and memory taxonomy.

  383. Hugging Face Daily Papers TIER_1 English(EN) ·

    DataFlow-Harness: 用于构建可编辑 LLM 数据管道的接地式代码代理平台

    Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. We call this disconnect the NL2Pipeline gap. To bridge it, we …

  384. arXiv cs.AI TIER_1 English(EN) · Jason Miklian ·

    人工智能LLM引擎如何塑造全球冲突信息环境

    arXiv:2607.14197v1 Announce Type: new Abstract: Artificial Intelligence (AI) answer engines now field a growing share of the questions that analysts, scholars, and the public ask about issues of peace and conflict. Large Language Models (LLMs) are known to hallucinate under certa…

  385. arXiv cs.AI TIER_1 English(EN) · Chuhan Shi, Xiaoquan Ren, Sicheng Song, Haobo Li, Rui Sheng, Yushi Sun ·

    大型语言模型已准备好进行科学发现吗?面向AI科学家的能力导向基准测试

    arXiv:2607.11079v1 Announce Type: new Abstract: Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, s…

  386. Hugging Face Daily Papers TIER_1 English(EN) ·

    大型语言模型已准备好进行科学发现吗?面向人工智能科学家的能力导向基准测试

    Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, statistical inference, mechanistic explanation, e…

  387. arXiv cs.CV TIER_1 English(EN) · Tianjie Ju, Xinyue Xu, Wanxuan Sun, Lingxiao Diao, Gongshen Liu, Zhuosheng Zhang, Cheng Yang ·

    首先,先教基于LLM的代理在“锦上添花”之前优先考虑“必不可少”

    arXiv:2609.05224v1 Announce Type: new Abstract: Recent progress in multimodal large language models (MLLMs) has fueled significant enthusiasm in their potential to act as autonomous agents for real-world tasks. However, scenarios requiring agents to fulfill users' complex, struct…

  388. arXiv stat.ML TIER_1 English(EN) · Nadeem Shaikh ·

    知道何时寻求帮助:分层 LLM 代理中的贝叶斯自我升级

    arXiv:2608.24087v1 Announce Type: cross Abstract: Current LLM agent systems decide delegation before reasoning begins (a router picks a model) or after a response is complete (a verifier scores it and may retry). We study a third regime: an agent that recognises, during its own r…

  389. arXiv stat.ML TIER_1 English(EN) · Tianbing Xu ·

    从期望最大化视角看大语言模型推理的强化学习

    arXiv:2504.18587v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrated by systems such as OpenAI's O1~\cite{o1} and DeepSeek-R1~\cite{r1}. However, wide…

  390. arXiv stat.ML TIER_1 English(EN) · Amirmohammad Farzaneh, Osvaldo Simeone ·

    短思考、巧推迟、执行、重复:边缘 LLM 代理的校准推理与不确定性感知推迟

    arXiv:2607.26865v1 Announce Type: new Abstract: LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. Yet, when deployed at the edge, they must tight…

  391. Hacker News — AI stories ≥50 points TIER_1 English(EN) · rellem ·

    Mozilla:开源AI的现状

  392. Forbes — Innovation TIER_1 English(EN) · Mohit Bhat, Forbes Councils Member ·

    微调SLM:企业AI的新运营模式

    Fine-tuning is transforming SLMs from efficient components into high-performance, enterprise-grade systems.

  393. MarkTechPost TIER_1 English(EN) · Michal Sutter ·

    10个开源无代码AI平台,用于构建LLM应用、RAG系统和AI代理

    <p>Retrieval, agents, and workflows now ship as visual and plain-English tools. This roundup covers 10 open-source no-code and low-code platforms for building LLM apps, RAG systems, and AI agents, each with its verified license, repository, and best-fit use case.</p> <p>The post …

  394. dev.to — MCP tag TIER_1 English(EN) · StarkMan ·

    LLM代理中的工具调用注入:为什么你的MCP服务器是新的攻击面

    <h1> Tool-Call Injection in LLM Agents: Why Your MCP Server Is the New Attack Surface </h1> <p>An LLM agent that can read email, browse the web and run shell commands is useful precisely because it acts on untrusted input. That combination is also why the Model Context Protocol (…

  395. Towards AI TIER_1 English(EN) · Pop123 ·

    Agent Harness:为何运行时脚手架与LLM同等重要

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/the-agent-harness-why-the-runtime-scaffolding-matters-as-much-as-the-llm-309579812fd9?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/957/1*Qj44n_mfPw9v-JOj…

  396. Medium — MCP tag TIER_1 English(EN) · FutureLens ·

    我如何使用MCP将一个LLM变成五个专业AI代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/data-science-collective/how-i-turned-one-llm-into-five-specialized-ai-agents-using-mcp-c1596941cefe?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/2600/0*0YCQBsWIakYQ1pfs"…

  397. Medium — MCP tag TIER_1 English(EN) · FutureLens ·

    我如何使用MCP将一个LLM变成五个专业AI代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ravendrakumar22000/how-i-turned-one-llm-into-five-specialized-ai-agents-using-mcp-c1596941cefe?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/2600/0*0YCQBsWIakYQ1pfs" wid…

  398. dev.to — MCP tag TIER_1 English(EN) · Ekemini Samuel ·

    企业人工智能治理:通过 AI 网关大规模治理 LLM 流量

    <p>Your AI governance problem probably isn't a policy problem; it’s a routing one. </p> <p>Let’s paint a common scenario: Your team uses OpenAI, another adds Anthropic, then the product team connects another model. Then someone builds an agent with MCP tools. A few months later, …

  399. dev.to — MCP tag TIER_1 English(EN) · Diego Costa ·

    使用原生MCP工具解决AI销售代理中的LLM参数幻觉问题

    <h1> Solving LLM Parameter Hallucinations in AI Sales Agents with Native MCP Tools </h1> <p>The most efficient way to eliminate LLM parameter hallucinations when retrieving B2B firmographics is by leveraging a native Model Context Protocol (MCP) server with strict Zod-enforced sc…

  400. dev.to — MCP tag TIER_1 English(EN) · Victor García ·

    Agent Gateway 60秒:使用TrustGate管理LLM流量

    <h1> Agent Gateway in 60 Seconds: Governed LLM Traffic with TrustGate </h1> <p>Most teams start with a direct OpenAI (or Anthropic) SDK call. That works until you have three apps, two providers, and a security review asking who can call which model, at what rate, with what audit …

  401. Towards AI TIER_1 English(EN) · Diogo Santos ·

    IntentFlow:具有可审计、哈希链式追踪的可控 LLM 代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/intentflow-governed-llm-agents-with-auditable-hash-chained-traces-49599f09e590?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1000/0*1jmaHvD-pw4NkgOa.png" …

  402. Towards AI TIER_1 English(EN) · MongoDB ·

    为多代理系统添加成本计量和 LLM 支出可见性

    <p><em>Written by </em><a href="https://www.linkedin.com/in/matteo-rossi-280391/"><em>Matteo Rossi.</em></a></p><p>The monthly LLM bill jumped, and nobody on the team can say which agent, which user, or which workflow caused it. The provider dashboard breaks usage down by organiz…

  403. Medium — fine-tuning tag TIER_1 English(EN) · Mikhail Borodastov ·

    Harness-native agents:与Harness共同训练LLM,以在单一任务中最大化质量

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://mlboroda.medium.com/harness-native-agents-co-train-the-llm-with-its-harness-to-max-out-quality-inside-one-task-93c321a93a81?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/250…

  404. dev.to — MCP tag TIER_1 English(EN) · Programming Central ·

    破解像素代码:视觉驱动的Agent如何将LLM的思考转化为DOM点击

    <p>The bleeding edge of AI automation isn't just about making Large Language Models (LLMs) smarter; it's about giving them hands and eyes. When building vision-driven agentic architectures, we cross a massive chasm: bridging the high-level semantic reasoning of an LLM with the lo…

  405. dev.to — MCP tag TIER_1 English(EN) · Diego Costa ·

    使用 Lead Enrichment MCP API 消除 B2B 销售代理中的 LLM 幻觉

    <h1> Eliminating LLM Hallucinations in B2B Sales Agents with the Lead Enrichment MCP API </h1> <p>To stop LLMs from hallucinating company data or fabricating contact details, developers must shift from loose prompt-based retrieval to a Model Context Protocol (MCP) architecture th…

  406. Medium — Claude tag TIER_1 English(EN) · Ashishmohanka ·

    大型语言模型(LLM)是如何工作的——构建 AI 代理的实用指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ashishmohanka123/how-llms-actually-work-a-practical-guide-for-building-ai-agents-53139a3b6665?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1024/1*thBASTiA9rwrB2FAwG8…

  407. dev.to — MCP tag TIER_1 English(EN) · Diego Costa ·

    通过原生MCP B2B增强功能消除销售代理中的LLM参数幻觉

    <h1> Eliminating LLM Parameter Hallucinations in Sales Agents with Native MCP B2B Enrichment </h1> <p>The most efficient way to stop LLMs from hallucinating firmographic data or misinterpreting complex API schemas is to deploy a Model Context Protocol (MCP) native B2B lead enrich…

  408. Medium — MLOps tag TIER_1 English(EN) · Tedi Ikonomi ·

    OpenShift AI Air-Gapped:为分布式 LLM 推理准备平台

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ikonomi.tedi/openshift-ai-air-gapped-preparing-the-platform-for-distributed-llm-inference-98a273f7bfdc?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/2600/1*RQnoWiHB6Pa…

  409. Medium — Claude tag TIER_1 English(EN) · Neo Malesa ·

    从ChatGPT聊天机器人到图谱:我们处理LLM方式的快速演变

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@neomalesa/from-chatgpt-chatbots-to-graphs-the-rapid-evolution-of-how-we-work-with-llms-c29894e87718?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1168/1*F2AM8F69iqJw_…

  410. Medium — MLOps tag TIER_1 English(EN) · Rami Krispin ·

    skforecast-ai 项目:面向生产系统的实用 LLM 评估 | 第 97 期

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@rami.krispin/the-skforecast-ai-project-practical-llm-evaluation-for-production-systems-issue-97-3c0b19ed14aa?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1920/1*tNzi0…

  411. dev.to — MCP tag TIER_1 English(EN) · PromptOT ·

    PromptOT MCP:管理和版本化您AI工具中的LLM提示

    <h1> PromptOT MCP: Manage and version LLM prompts from your AI tools </h1> <p>Prompts often start as simple strings in code.</p> <p>Then the product grows.</p> <p>You add a better system prompt. Then a guardrail. Then a different version for production. Then a customer-specific v…

  412. Medium — MLOps tag TIER_1 English(EN) · Neelopphersyed ·

    NeuralUCB Router:一个兼容OpenAI的API代理,使用多臂...

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@neelopphersyed7/neuralucb-router-an-openai-compatible-api-proxy-that-routes-llm-requests-using-a-multi-armed-17e762724926?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max…

  413. Medium — MCP tag TIER_1 English(EN) · EuroAmerican Institute ·

    LLM vs RAG vs MCP:AI工程师和开发者的游戏规则改变者

    <div class="medium-feed-item"><p class="medium-feed-snippet">The Model Context Protocol hit 97 million monthly SDK downloads by December 2025. That number alone tells you something important is&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@euroamericanmalta/…

  414. Medium — Claude tag TIER_1 English(EN) · Shankar ·

    每月20美元的人工智能失误:大型语言模型如何使AWS架构复杂化

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@shankar_somasundaram/the-20-month-ai-mistake-how-llms-overcomplicate-aws-architecture-de8980849607?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1024/1*CZ7mUHQAGZQdxg…

  415. Towards AI TIER_1 English(EN) · Vasilii Chetvertukhin ·

    迈向自托管企业人工智能的四层架构

    <h4>There is no shortage of articles about building AI agents. What remains much rarer is a practical discussion of how to run them safely in production.</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*tnltjwGYfOIZX6KgpLNUEg.png" /></figure><p>This article…

  416. Medium — MLOps tag TIER_1 English(EN) · sentraorb ·

    统一所有大语言模型的入口:使用LiteLLM构建生产级AI堆栈

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://aws.plainenglish.io/one-gateway-to-rule-all-your-llms-building-a-production-ready-ai-stack-with-litellm-1ffcb29a7733?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1280/1*zFxMRtOqO…

  417. Medium — MLOps tag TIER_1 English(EN) · sentraorb ·

    统一所有 LLM 的一个入口:使用 LiteLLM 构建生产就绪的 AI 堆栈

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://sentraorb.medium.com/one-gateway-to-rule-all-your-llms-building-a-production-ready-ai-stack-with-litellm-1ffcb29a7733?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1280/1*zFxMRtOq…

  418. Medium — Anthropic tag TIER_1 Español(ES) · LinaUX Off Frame ·

    人工智能的能力与局限性:教我诊断LLM错误的课程

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@l.godefroy.design/ai-capabilities-and-limitations-el-curso-que-me-ense%C3%B1a-a-diagnosticar-los-errores-de-las-llm-c2047c580120?source=rss------anthropic-5"><img src="https://cdn-images-1.med…

  419. Medium — fine-tuning tag TIER_1 English(EN) · Tech Horizon With Anand Vemula ·

    微调LLM:开发者定制AI模型的指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@anandvlinkedin/fine-tuning-llms-a-developers-guide-to-custom-ai-models-2e7b5e7989aa?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1536/1*nmFyPKH5QY0oRBC5XJIJNw.p…

  420. dev.to — LLM tag TIER_1 English(EN) · Imversion Tech ·

    人工智能后备策略:提升2026年大型语言模型的可靠性

    <h2> How AI Fallback Strategies Keep Uncertain Systems Safe and Useful </h2> <p>The real problem with AI systems is not that they fail. It is that they fail with confidence. When uncertainty rises, the system should catch it early, switch to a controlled fallback, and keep the us…

  421. dev.to — LLM tag TIER_1 English(EN) · Yaseen Khatib ·

    当日志欺骗时:使用 OpenTelemetry 追踪 LLM Agents

    <p>[ EXECUTIVE TEARDOWN // TL;DR ]</p> <ul> <li> Emit one span per agent step so a request becomes a readable waterfall — total at the root, a labelled child per stage.</li> <li> Use the OpenTelemetry GenAI semantic conventions (gen_ai.*) so any backend renders traces the same wa…

  422. dev.to — LLM tag TIER_1 English(EN) · Manoranjan Rajguru ·

    Qwen 3.8 27B:可装入笔记本电脑的前沿大语言模型——架构、推理控制与智能体集成

    <h1> Qwen 3.8 27B: The Frontier LLM That Fits on Your Laptop — Architecture, Reasoning Control &amp; Agentic Integration </h1> <p><em>Published August 18, 2026 · 18 min read</em></p> <h2> Table of Contents </h2> <ol> <li>The Benchmark Moment Nobody Saw Coming</li> <li>How Good Is…

  423. dev.to — LLM tag TIER_1 English(EN) · sekera-radim ·

    用于 LLM 智能体的“人工在环”工具对比

    <p>A practical comparison of the four categories of human-in-the-loop tooling for LLM agents — framework-native interrupts, workflow-engine nodes, general approval software, and dedicated gates — so you pick by category, not by name.</p> <h2> The four categories </h2> <p>"HITL to…

  424. dev.to — LLM tag TIER_1 English(EN) · Hung Nguyen ·

    去你的CTF:多智能体LLM真的能玩转夺旗赛吗?

    <p>I’ve been working on a side project called F*ckCTF — an autonomous agent designed to solve black-box Capture The Flag challenges. It runs inside an isolated Kali Linux container and uses a Multi-Agent architecture to interact with terminals, run scripts, and hunt for flags.</p…

  425. r/MachineLearning TIER_1 English(EN) · /u/AccomplishedLeg1508 ·

    EvoUndo:可恢复性约束的自演化 LLM Agent 机制 [R]

    <!-- SC_OFF --><div class="md"><p>LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a successful mutation may leave persistent effects that cannot be safely reversed …

  426. dev.to — LLM tag TIER_1 English(EN) · mech.app ·

    LLM运行时干预:Mentat 如何在不进行微调的情况下引导代理推理

    <p>Most production LLM control sits between two extremes: prompt engineering (brittle, context-dependent) and fine-tuning (expensive, slow iteration). Mentat, a YC F24 launch, introduces a third path: runtime intervention that modifies token probabilities mid-generation without r…

  427. dev.to — LLM tag TIER_1 English(EN) · mech.app ·

    LLM 代理工作流中的约束弱化:为什么在多阶段管道中“必须”变成“也许”

    <p>Multi-stage LLM agent workflows have a silent failure mode. A hard constraint enters the pipeline at stage one. By stage three, it has become a suggestion. The executor reads it, acknowledges it, and proceeds anyway.</p> <p>The problem is not hallucination or context loss. The…

  428. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    LLM 智能体在药物模拟上运行受控实验 斯图加特大学的多智能体框架让 LLM 智能体进行设计、运行和解读

    LLM agents run controlled experiments on pharma simulations A multi-agent framework from the University of Stuttgart lets LLM agents design, run, and interpret experiments on pharmaceutical simulation models. https://www. notatechguy.com/llm-agents-run -controlled-experiments-on-…

  429. dev.to — LLM tag TIER_1 English(EN) · weiwuji ·

    Agent Engineering Physicalization: 9 Pillars That Turn Probabilistic LLMs into Deterministic Systems

    <blockquote> <p><strong>The Pain</strong>: Your agent forgets to query the knowledge base. The same knowledge base, the same model — different orchestration produces wildly different outputs. Change one prompt line and everything breaks. Agents still "work by feel", and every fix…

  430. dev.to — LLM tag TIER_1 English(EN) · Casey Zhang ·

    不带营销水分的 LLM 智能体基准测试:数据集、指标与控制

    <p>Last month, someone pasted a benchmark table into our team chat. Model B beat model A by 12 points on "general agent tasks," so we swapped models the same afternoon.</p> <p>Two weeks later, the triage agent was mislabeling about a third of the issues it touched.</p> <p>The tab…

  431. dev.to — LLM tag TIER_1 English(EN) · Ayi NEDJIMI ·

    Agentic AI 安全:在生产环境中沙箱化 LLM 工具调用

    <p>When you give a language model the ability to call tools — run code, query databases, browse the web — you've created an autonomous execution surface. Most tutorials skip the part where that surface gets exploited.</p> <p>This post covers practical steps for sandboxing LLM too…

  432. dev.to — LLM tag TIER_1 English(EN) · Umair Bilal ·

    我的AI智能体2x2模型成本效益策略

    <blockquote> <p><em>This article was originally published on <a href="https://www.buildzn.com/blog/my-2x2-llm-cost-performance-strategy-for-ai-agents" rel="noopener noreferrer">BuildZn</a>.</em></p> </blockquote> <p>Everyone's chasing the biggest LLMs, throwing cash at Claude or …

  433. dev.to — LLM tag TIER_1 English(EN) · Alex ·

    Pydantic AI 评测:为 Python 提供类型化代理,该框架使 LLM 输出更可靠

    <p>Pydantic AI is the official agent framework from the Pydantic team, built around typed, validated LLM output. After 45 days of using it for saas.pet's content QA agent and data extraction scripts, here is the real story on structured output, tool calling, and why it beats Lang…

  434. Mastodon — fosstodon.org TIER_1 Русский(RU) · [email protected] ·

    TrueForge:一个将大型语言模型转变为全功能代理的开源框架。今天即可轻松构建演示代理。连接模型,添加几个工具

    TrueForge: открытая обвязка, которая превращает LLM в полноценного агента Собрать демо-агента сегодня несложно. Подключаешь модель, добавляешь пару инструментов — и она уже читает файлы, вызывает API и бодро обещает выполнить любую задачу. Сложности начинаются, когда такого агент…

  435. dev.to — LLM tag TIER_1 ไทย(TH) · Nokka ·

    AI代理的LLM路由,为什么代理需要选择模型以及如何将成本降低高达80%

    <h1> LLM Routing สำหรับ AI Agent, ทำไม agent ถึงต้องเลือกโมเดลเป็น และวิธีลดต้นทุนได้ถึง 80% </h1> <p><em>โดย Nokka (นก-กา) | 21 สิงหาคม 2026</em></p> <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity…

  436. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    超越#LLM:使用LangChain DeepAgents创建真正的#AI代理,了解LangChain DeepAgents如何将LLM转变为生产就绪的AI系统

    Beyond # LLMs : Creating Real-World # AI Agents with Lang Chain Deep Agents Discover how LangChain DeepAgents transform LLMs into production-ready AI systems with memory, skills, sub-agents, context management, and human oversight. https:// hackernoon.com/beyond-llms-cre ating-re…

  437. Mastodon — fosstodon.org TIER_1 Français(FR) · [email protected] ·

    SkillOpt (Microsoft): LLM 智能体的自然语言技能优化器。该技能通过评分推广进行改进,无需触碰模型权重。

    SkillOpt (Microsoft) : un optimizer de skills en langage naturel pour agents LLM. Le skill s'améliore via des rollouts scorés, sans toucher aux poids du modèle. Le fichier best_skill.md est portable d'un modèle à l'autre. Open source, MIT. ⬇️ https:// github.com/microsoft/SkillOp…

  438. dev.to — LLM tag TIER_1 English(EN) · Felipe L ·

    Zero-Mem:LLM智能体零Token记忆操作

    <h2> What Happened </h2> <p>Zero‑Mem lets LLM agents read and write external memory without generating or consuming any tokens. Traditional agents fetch context through token‑based prompts, adding latency and cost. Zero‑Mem replaces that with a lightweight, token‑free interface t…

  439. dev.to — LLM tag TIER_1 English(EN) · Ming ·

    在边缘运行LLM代理:使用NeoMind + Ollama的实用指南

    <h1> Running LLM Agents at the Edge: A Practical Guide with NeoMind + Ollama </h1> <p>Everyone's building AI agents right now. Most of them live in the cloud — you send a request to OpenAI or Anthropic, get a response back, and hope the latency and cost stay reasonable. But what …

  440. dev.to — LLM tag TIER_1 English(EN) · talor ·

    利用实时搜索进行AI代理构建:SERP API如何实现可靠的LLM应用

    <p>Large language models have changed how developers build applications.</p> <p>However, even the most advanced LLMs have one fundamental limitation:</p> <p>They do not have access to real-time information.</p> <p>A model may understand programming, reasoning, and language extrem…

  441. dev.to — LLM tag TIER_1 English(EN) · Yogi ·

    构建我自己的LLM模型和代理

    <h2> Introduction </h2> <p>Large Language Models (LLMs) are powerful, but most enterprises rely on pre‑packaged APIs. I wanted to go deeper: train my own LLM model and build an agent layer on top of it that could interact with real systems securely.</p> <p>This post walks through…

  442. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    LLM代理:纯代码验证将目标放弃率从100%降至0% 新arXiv预印本:确定性执行器拥有所有代理信念,LLM仅负责归档

    LLM agent: code-only verification flips goal abandonment 100% to 0% New arXiv preprint: a deterministic executive owns all agent belief, the LLM only files proposals, and zero ARC-AGI-3 completions are honestly disclosed. https://www. notatechguy.com/llm-agent-code -only-verifica…

  443. dev.to — LLM tag TIER_1 English(EN) · kai wen ng ·

    迈向LLM智能体稳定性

    <p>An industry-level LLM agent is not simply an API call that returns a response. It needs to be resilient to transient failures, malformed outputs, and schema violations.<br /> To improve the stability of my agent system, I introduced two decorators around my LLM calls. They han…

  444. dev.to — LLM tag TIER_1 English(EN) · Lorena Dávila Ermus ·

    1. 自托管AI:高效运行模型所需的LLM概念

    <p>If you want to run AI models on your own machine and learn the basic concepts with me to do it effectively, then this is the right article :).</p> <p>This is part one of the series. In the next one we build local AI workflows with n8n and Ollama. This article is the vocabulary…

  445. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📄 ‘MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations’ 在 Hugging Face 上获得 80 个赞。测试 LLMs 在持续任务链上的表现

    📄 ‘MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations’ hit 80 upvotes on Hugging Face. Tests LLMs on sustained task chains in shopping. https:// huggingface.co/papers/2607.289 56 # AI # MachineLearning # Research

  446. dev.to — LLM tag TIER_1 English(EN) · Mohammad Jawad (Kasir) Barati ·

    理解LLM智能体和工具

    <p>In this post, I'll explain how prompt chaining, tools/skills, and iteration actually make agnets to produce results which are not usually possible when we use simple LLMs.</p> <h2> LLM chaining </h2> <p>You know how sometimes you write one massive prompt like:</p> <blockquote>…

  447. dev.to — LLM tag TIER_1 English(EN) · Dmitriy ·

    我们如何为 AI 代理选择 LLM 和框架

    <p>Over the last 18 months our ML team has been doing some very interesting things: building AI agents on top of PostgreSQL, while the infrastructure evolves, the industry matures, and quality expectations keep rising. We started with a single A100 in a managed cloud and fairly m…

  448. dev.to — LLM tag TIER_1 English(EN) · Wibo ·

    如何评估 LLM 代理:evals、黄金数据集和 LLM 作为裁判

    <p><strong>Short answer</strong></p> <p><strong>You can't unit-test an LLM to correctness, because the same input can take a different path on the next run.</strong> Evals are the test suite for probabilistic systems: a scored, repeatable check of whether the system reached an ac…

  449. dev.to — LLM tag TIER_1 English(EN) · Hiroshi Toyama ·

    面向 LLM Agent 的分层评估策略(Google ADK 的 12 条标准)

    <p>Google's <a href="https://adk.dev/evaluate/criteria/" rel="noopener noreferrer">Agent Development Kit (ADK)</a> ships 12 evaluation criteria for testing agent behavior: tool-call trajectories, final response quality, hallucination detection, safety, multi-turn task success, an…

  450. dev.to — LLM tag TIER_1 Deutsch(DE) · Tsari Bombelli ·

    llms.txt 详解:AI 爬虫和 LLMs 的标准

    <p>llms.txt ist ein maschinenlesbarer Standard, der KI-Systemen strukturierte Informationen über Ihre Website bereitstellt. Aufbau, Best Practices und praktische Implementierung für bessere KI-Sichtbarkeit bei ChatGPT, Claude, Gemini und Perplexity.</p> <h3> Zusammenfassung </h3>…

  451. dev.to — LLM tag TIER_1 English(EN) · Jules Robineau ·

    重建以理解:从网络协议到LLM代理

    <blockquote> <p><strong>TL;DR</strong>: you only truly understand a system once you rebuild it. I recoded TCP at school, then the DNS protocol, then Modbus, each time to understand it from the inside. A colleague just went through this with LLMs. He wrote a small agent in Go, and…

  452. dev.to — LLM tag TIER_1 English(EN) · Yusuf Al-Rashidi ·

    9 款最佳 LLM 网关,适用于 Agentic 工作流和 AI 代理

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fairuyu0zpj9tz04zmsbj.png"><img alt="9 Best LLM Gatew…

  453. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    刚刚为AI代理样板文件添加了本地LLM支持,现在可以在Ollama之上运行,而无需依赖云API 👇 https://github.com/christophed

    Just added local LLM support to the AI agent boilerplate, you can now run it on top of Ollama instead of relying on cloud APIs 👇 https:// github.com/christopheduc-me/ai -agent-boilerplate # buildinpublic # ai # dev # tech

  454. dev.to — LLM tag TIER_1 English(EN) · TheKitBase ·

    2026年如何削减您的AI/LLM成本:缓存、更便宜的模型和多代理路由

    <p>AI features ship fast and then the bill arrives. The good news: most LLM spend is avoidable waste - the same prompt paid for a thousand times, a frontier model doing work a cheap one could handle, tokens generated that nobody reads. Here are six levers that cut real money, ord…

  455. dev.to — LLM tag TIER_1 English(EN) · Apache SeaTunnel ·

    AI能否真正构建数据管道?使用Apache SeaTunnel AI CLI对7个领先LLM进行100项任务的基准测试

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbcakf9op8ltnlaaujm5z.jpg"><img height="439" src="htt…

  456. dev.to — LLM tag TIER_1 Türkçe(TR) · Emre Yıldız ·

    在您自己的服务器上运行LLM:10分钟用Ollama实现本地AI

    <p>OpenAI API'ına token başına para ödemek yerine, açık kaynak dil modellerini (Llama 3, DeepSeek, Mistral, Qwen) kendi sunucunda çalıştırabilirsin. Verin dışarı çıkmaz, sabit maliyet, sınırsız istek. Bu yazıda Ollama ile pratik kurulumu ve gereken donanımı anlatıyorum.</p> <h2> …

  457. dev.to — LLM tag TIER_1 English(EN) · Learn AI Resource ·

    在本地运行开源大模型:您的AI编程助手,无需API费用

    <p>So you want an AI coding assistant but you're tired of getting dinged for API calls every time you ask for help debugging a regex? Yeah, I get it.</p> <p>Here's the thing: you don't actually need to pay OpenAI or Anthropic to get decent AI pair programming. You can run a solid…

  458. dev.to — LLM tag TIER_1 English(EN) · Sofia Aliferi ·

    超越审核:为何大型语言模型系统需要策略层

    <blockquote> <p>TL;DR: Moderation catches harm and many injection attempts. It does not enforce domain or operational policy. A policy reasoning layer (LLM-as-a-judge) closes that gap, especially in multi-turn conversations.</p> </blockquote> <p><strong>Abstract</strong><br /> Mo…

  459. dev.to — LLM tag TIER_1 English(EN) · soy ·

    本地大模型、开放代理和自托管部署平台趋势兴起

    <h2> Local LLMs, Open Agents &amp; Self-Hosted Deployment Platforms Trending </h2> <h3> Today's Highlights </h3> <p>Today's top stories highlight the growing trend of local and self-hosted AI deployments, featuring an architectural guide for secure "Local Sovereign LLMs" in enter…

  460. dev.to — LLM tag TIER_1 English(EN) · Kuldeep Paul ·

    按团队跟踪 LLM 使用情况和支出:AI 治理指南

    <p><em>Organizations deploying AI applications face challenges in accurately tracking LLM usage and spend across different teams and projects. <a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer">Bifrost</a> offers a comprehensive AI gateway solution with virtual k…

  461. dev.to — LLM tag TIER_1 English(EN) · Stéphane Derosiaux ·

    chrome-agent:将任何大型语言模型变成智能网页浏览代理

    <p>Ever handed an LLM a full web page and watched the amount of tokens being used?</p> <p>A single product listing is 20-30K tokens of </p> soup before the model finds what it needs: wrapper divs, css class, SVG, JSON blobs etc. The agent needs maybe 300 tokens of that (the items…

  462. dev.to — LLM tag TIER_1 English(EN) · DryDock ·

    LLM 代理操作 CAD 内核的七个实际失败(及其架构如何控制了它们)

    <p>I built a system where an LLM talks to a customer about a silicone casting mold, and a<br /> deterministic geometry kernel — OpenCASCADE, three decades of production C++ — does the<br /> actual mass-solving, boolean surgery, and part-splitting. The LLM never touches the kernel…

  463. dev.to — LLM tag TIER_1 English(EN) · soy ·

    LLM推理与RAG优化,用于本地部署的开源语音AI

    <h2> LLM Inference &amp; RAG Optimization, Open-Source Voice AI for Local Deployments </h2> <h3> Today's Highlights </h3> <p>This week's highlights feature a new framework for LLM inference and fine-tune optimizations, including KV-cache improvements, alongside an open-source voi…

  464. dev.to — LLM tag TIER_1 English(EN) · Kuldeep Singh Arya ·

    生产环境中 8 种 LLM 成本优化技术

    <p>Executive Summary<br /> As generative AI transitions from experimental prototypes to high-scale production systems, the primary bottleneck for engineering teams has shifted from model capability to unit economics. The pricing structure of modern Large Language Model (LLM) APIs…

  465. dev.to — LLM tag TIER_1 English(EN) · Shouvik Palit ·

    Sir Shortoken:为每个LLM提供严谨的AI输出

    <p><strong>TL;DR:</strong> Sir Shortoken is a system prompt that constrains frontier models to operate within information budgets (Quick/Balanced/Deep), never silently escalate capabilities, and prove execution. Tested across Claude, GPT, Gemini. 40-60% token reduction on technic…

  466. dev.to — LLM tag TIER_1 English(EN) · Praveen Maurya ·

    使用本地 LLM 构建:工程师的 AI 辅助开发方法

    <blockquote> <p>I didn't build SafeDevTools by asking AI to "build me a website." I built it by treating a local LLM like a junior engineer who never gets tired of writing boilerplate.</p> </blockquote> <p>A few weeks ago, I challenged myself with a simple experiment:<br /> <stro…

  467. dev.to — LLM tag TIER_1 Deutsch(DE) · Uhltak Therestismysecret ·

    使用 Ollama 部署本地大模型:托管模型、集成 API 并高效使用

    <h1> Lokale LLMs mit Ollama – Modelle selbst hosten und per API anbinden </h1> <p><strong>Hook:</strong> Stell dir vor, du könntest ChatGPT für deine Firma betreiben, ohne einen teuren Cloud‑Vertrag oder ein Datenleck‑Szenario. Du hast die volle Kontrolle, die Kosten liegen bei d…

  468. dev.to — LLM tag TIER_1 Русский(RU) · Promptra Team ·

    本地AI模型:Exo Labs的local.ai Tracker与Mac Mini集群

    <p>Если ты открыл эту статью с вопросом «где посмотреть, какая модель реально влезет в мой Mac и не будет тормозить», то короткий ответ такой: 2 июля 2026 года Exo Labs на конференции AI Engineer World's Fair анонсировала сервис local.ai, который отслеживает, какая модель лучше в…

  469. dev.to — LLM tag TIER_1 English(EN) · shayesta ·

    LangChain4j 和 Spring AI:让您的 Java 应用与 LLM 对话的“管道”

    <p>If you've heard about LangChain and assumed it was a Python thing, that's fair. It mostly was.</p> <p>LangChain became popular because building with an LLM turns out to involve a lot of repetitive plumbing. You need to manage conversation history, split documents into chunks, …

  470. dev.to — LLM tag TIER_1 English(EN) · Bahadir Kusat ·

    人工智能模型如何训练?LLM训练指南

    <p>From data preparation and tokenizer selection to pretraining, LoRA, RLHF, evaluation, and production monitoring, this guide covers the major stages involved in training an AI model.</p> <p>DEHA Research · July 14, 2026 · 18 min read</p> <p>Training an artificial intelligence m…

  471. dev.to — LLM tag TIER_1 English(EN) · soy ·

    浏览器 LLM 代理、Apple Silicon 的 Rust 引擎,以及本地 AI 代码解释器

    <h2> Browser LLM Agents, Rust Engine for Apple Silicon, &amp; Local AI Code Interpreter </h2> <h3> Today's Highlights </h3> <p>This week, we spotlight tools bringing LLM inference directly to your devices. Dive into browser-based agents, a Rust-native engine for Apple Silicon, an…

  472. dev.to — LLM tag TIER_1 English(EN) · Innocent Oyebode ·

    我如何为尼日利亚中小企业构建多页面AI网站生成器——架构、LLM提示和经验教训

    <h2> The Problem </h2> <p>Most Nigerian small businesses have no web presence at all. When they do get a website, it is usually a stale brochure-ware page that took a freelancer three weeks to deliver and costs ₦150,000 they could not really afford. The freelancer is long gone; t…

  473. dev.to — LLM tag TIER_1 English(EN) · Jack M ·

    LLM 延迟预算:让 AI 工作流感觉快速,无需猜测

    <p>A slow AI feature rarely fails all at once. It starts with a longer prompt, then a bigger retrieval result, then one more tool call, then a retry path nobody measured. The demo still works, but users feel the delay before your dashboard explains it.</p> <p>That is why small AI…

  474. dev.to — LLM tag TIER_1 English(EN) · soy ·

    自托管AI伴侣与开源模型API洞察

    <h2> Self-Hosted AI Companion &amp; Open-Source Model API Insights </h2> <h3> Today's Highlights </h3> <p>This week's highlights feature a trending self-hosted AI companion, empowering users with personal, locally-run AI experiences. We also explore a bootcamp grad's practical in…

  475. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    从演示到可靠系统:让大型语言模型真正可投入生产的人工智能工程技术

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/from-demos-to-durable-systems-ai-engineering-techniques-that-make-llms-truly-product-ready?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noreferrer">…

  476. dev.to — LLM tag TIER_1 English(EN) · soy ·

    自托管大模型应用、离线AI系统及本地自动化基础

    <h2> Self-Hosted LLM Apps, Offline AI Systems, and Local Automation Foundations </h2> <h3> Today's Highlights </h3> <p>This week, we spotlight practical approaches to self-hosting AI, from extensive curated lists of runnable LLM applications to ambitious projects building fully o…

  477. dev.to — LLM tag TIER_1 English(EN) · bossandboss ·

    构建 EdgeSync-LLM:去中心化、离线优先的本地 AI 的最终架构 🚀

    <h1> Published: true </h1> <h1> Description: A deep dive into the final version of EdgeSync-LLM—bringing fast, secure, synchronized Large Language Models straight to edge hardware. </h1> <h1> Tags: ai, open source, architecture, edgecomputing, webdev </h1> <p>The cloud dependency…

  478. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    对于科技界的任何人来说,探索 Llama、Mistral 和 Phi 等开源 AI 模型都是必不可少的!这些模型通过推广正在改变 AI 的格局

    Exploring open-source AI models like Llama, Mistral, and Phi is a must for anyone in the tech world! These models are changing the landscape of AI by promoting collaboration and innovation. Dive into the world of deep learning! 🤖 # AI # AITürkiye # DeepLearning # Teknoloji # Mach…

  479. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    GPT-5.6 现身:OpenAI 的新模型和定制芯片将如何重塑生产型 LLM 系统

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/gpt-5-6-in-the-wild-how-openai-s-new-model-and-custom-silicon-will-reshape-production-llm-systems?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noref…

  480. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    GPT-5.6、Jalapeño 以及下一代 OpenAI 优化的大型语言模型基础设施

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/gpt-5-6-jalapeno-and-the-next-generation-of-openai-optimized-llm-infrastructure?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noreferrer">CoreProse K…

  481. dev.to — LLM tag TIER_1 English(EN) · Abdul Rehman ·

    构建生产级AI流水线:使用LLM每日处理10,000+个列表

    <p>I learned the hard way that a working LLM pipeline and a production LLM pipeline are two different things.</p> <p>When I first built the scoring system for a job board platform, I thought: throw GPT-4 at each listing, ask it to rate relevance, done. It worked for 100 listings.…

  482. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    深入了解 GPT-5.6:OpenAI 的新旗舰模型和定制芯片将如何重塑 LLM 运营

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/inside-gpt-5-6-how-openai-s-new-flagship-model-and-custom-silicon-will-reshape-llm-operations?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noreferre…

  483. dev.to — LLM tag TIER_1 English(EN) · Remy Okafor ·

    面向大模型团队的10款开源AI基础设施工具

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr8ju6nt2a4ngc3j9sww5.png"><img alt="10 Open-Source A…

  484. dev.to — LLM tag TIER_1 English(EN) · Caleb Osei ·

    流式传输 LLM 响应的最佳 AI 网关

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwk9fm8436j9drcevri9t.png"><img alt="Best AI Gateways…

  485. dev.to — LLM tag TIER_1 English(EN) · Kuldeep Paul ·

    AI网关如何提高LLM的可靠性:9种方法

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frp4yqktt436b7wu1fox2.png"><img alt="9 Ways an AI Gat…

  486. dev.to — LLM tag TIER_1 English(EN) · Abdul Rehman ·

    如何为您的AI MVP构建可靠的LLM管道,避免过度设计

    <p>I once built an AI pipeline that was shut down after a single month. The LLM costs were unsustainable, and worse, the outputs were unreliable enough that we couldn't trust them in production. That failure taught me something I still use today: evaluation isn't a phase you add …

  487. Mastodon — mastodon.social TIER_1 English(EN) · aitools2u ·

    🤖 【Hugging Face Papers】PlannerForge:LLM 智能体用于自动驾驶运动规划器的基于场景的测试 确保自动驾驶安全是

    🤖 【Hugging Face Papers】PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used t... # AI # TechNews #... 🔗 https:// huggingf…

  488. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Apple 为 Series 12 带回陶瓷材质 新款陶瓷配色:珍珠白和午夜蓝。| 图片:The Verge 援引 Apple Apple Watch Series 12 发布

    Apple brought ceramic back for the Series 12 The new ceramic colors: pearl white and night blue. | Image: The Verge via Apple The Apple Watch Series 12 is launching with a ceramic option, making it the company's first wearable since 2019's Series 5 that can be con… https://www. t…

  489. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    程序化图:LLM智能体的自演化执行结构

    Procedural Graphs: Self-Evolving Execution Structures for LLM Agents Article URL: https:// academy.dair.ai/papers/procedu ral-graphs-self-evolving-execution-structures-for-llm-agents-2609.09153 Comments URL: https:// news.ycombinator.com/item?id=4 9629868 Points: 4 # Comments: 0 …

  490. Mastodon — mastodon.social TIER_1 English(EN) · CuratedHackerNews ·

    Procedural Graphs: LLM 智能体自演化执行结构

    Procedural Graphs: Self-Evolving Execution Structures for LLM Agents https:// academy.dair.ai/papers/procedu ral-graphs-self-evolving-execution-structures-for-llm-agents-2609.09153 # ai # llm

  491. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Source: pageindex.ai/blog/ocr 向量 RAG:为何它在生产环境中脱颖而出... # ai # llm # machinelearning # rag # software # coding # development # engineer

    Source: pageindex.ai/blog/ocr Vector RAG: Why It’s Winning in Production In a... # ai # llm # machinelearning # rag # software # coding # development # engineering # inclusive # community Vector RAG: Why It’s Winning in Production

  492. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    通过运行时引导实现对LLM行为的确定性控制。金融代理的架构、延迟权衡和合规性影响。# agents #

    Deterministic control over LLM behavior through runtime steering. Architecture, latency trade-offs, and compliance implications for financial agents. # agents # ai # api # llm # software # coding # development # engineering # inclusive # community Runtime Intervention for LLMs: H…