PulseAugur
实时 15:24:03
English(EN) Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

最新研究探讨LLM智能体在技能选择、自动驾驶和合规性方面的进展

arXiv上发布的多篇研究论文探讨了大型语言模型(LLM)智能体的进展,重点在于提高其能力和可靠性。其中一篇论文介绍了用于LLM智能体最优技能选择的最佳前缀选择(BPS),该方法在性能和代币成本方面提供了可证明的保证。另一项研究提出了一个混合框架用于自动驾驶,该框架整合了LLM的常识推理与强化学习和PID控制,以增强决策能力。此外,还有研究通过纵向生命轨迹来缓解LLM智能体中的身份本质主义,并开发了LLM智能体的策略合规性和故障归因方法。 AI

影响 这些进展旨在提高LLM智能体在自动驾驶和金融合规等不同领域的性能、可靠性和适用性。

排序理由 arXiv上发表了多篇研究论文,详细介绍了LLM智能体的新方法和基准测试。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 291 个来源。 我们如何撰写摘要 →

最新研究探讨LLM智能体在技能选择、自动驾驶和合规性方面的进展

报道来源 [291]

  1. arXiv cs.AI TIER_1 English(EN) · Jiajun Wu, Zirui Wang, Jiayu Zhou, Qiang Ye, Steve Drew ·

    FL-MAESTRO:面向资源受限联邦学习的多智能体大模型编排

    arXiv:2608.20518v1 Announce Type: new Abstract: In Federated Learning (FL), the communication topology is a runtime variable rather than a fixed design choice, since links and edge devices drop in and out during training. Each round, the server must commit three coupled decisions…

  2. arXiv cs.AI TIER_1 English(EN) · Guodong Xu ·

    LLM智能体中的标准修订校准:失效模式与基于轨迹的协议

    arXiv:2608.20729v1 Announce Type: new Abstract: Language-model agents can improve after failure or carry text across episodes without revising what counts as success. We study the narrower attribution problem of criterion revision: when criterion K0 accepts an outcome violating a…

  3. arXiv cs.AI TIER_1 English(EN) · Quang Dao, Purvi Kathalkar, Kenneth Eaton ·

    加权记忆树:让长时域LLM智能体记住重要信息

    arXiv:2608.20631v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated the ability to solve multi-step tasks requiring planning, tool use, and external information access, yet growing execution histories increase inference cost and expose reasoning to…

  4. arXiv cs.LG TIER_1 English(EN) · Quan Wei, Siliang Zeng, Chenliang Li, Zhongruo Wang, William Brown, Oana Frunza, Wei Deng, Anderson Schneider, Yuriy Nevmyvaka, Yang Katie Zhao, Alfredo Garcia, Mingyi Hong ·

    通过细粒度奖励结构和信用分配增强LLM智能体多轮推理能力

    arXiv:2505.11821v3 Announce Type: replace Abstract: Reinforcement Learning (RL) approaches have been wildly used to enhance the reasoning capabilities of Large Language Model (LLM) agents in long-horizon, multi-turn scenarios. Such interactions can be formalized as turn-level Mar…

  5. arXiv cs.AI TIER_1 English(EN) · Yuyuan Feng, Zhishang Xiang, Chaobin Yang, Qichao Ma, Zerui Chen, Yujing Zhang, Ke Huang, Chuanjie Wu, Zhaoxu Liu, Yili Wang, Xin He, Jiapu Wang, Zijin Hong, Hao Chen, Yuanchen Bei, Kun Wang, Shengyuan Chen, Ningyu Zhang, Enyan Dai, Linhao Luo, Qingyi Pa… ·

    LLM智能体时代的图工程:从个体智能到系统智能

    arXiv:2608.21156v1 Announce Type: cross Abstract: LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolution has produced paradigms including Prompt Engineering to elicit model capabilities, Context Engineering to manage…

  6. arXiv cs.AI TIER_1 English(EN) · Kai Wang, Zeming Wei, BiaoJie Zeng, Chang Jin, An Wang, Xiaokun Luan, Zhixiao Lin, Jingjing Qu, Xia Hu, Xingcheng Xu ·

    ClawSentry:一款渐进式多层安全监控器,用于保护自主LLM代理

    arXiv:2608.21101v1 Announce Type: cross Abstract: As large language model (LLM) agents move from conversation to executing code, reading local files, and orchestrating external tools, a single agent hijacked by a malicious third-party skill can cause data exfiltration, privilege …

  7. arXiv cs.AI TIER_1 English(EN) · Wei Shao, Chongzhou Fang, Zuxiong Tan, Zequan Liang, Setareh Rafatirad, Avesta Sasan, Houman Homayoun ·

    超越端到端成功:诊断长周期安全LLM代理的失败

    arXiv:2608.20563v1 Announce Type: cross Abstract: Long-horizon security LLM agents must carry information and decisions across many dependent interactions, where later actions often depend on services, state, or access discovered much earlier. This makes final task success diffic…

  8. arXiv cs.AI TIER_1 English(EN) · Yanze Jiang, Mingxuan Li, Yuhao Wang, Shengfang Zhai, Jiaheng Zhang ·

    无需解决,只需比较:用于LLM代理运行时干预的微型顾问

    arXiv:2608.21027v1 Announce Type: new Abstract: LLM agents are emerging as an important paradigm for real-world tasks that require reasoning, tool use, and sequential decision-making. As these agents operate over longer horizons, runtime intervention offers a way to improve relia…

  9. Hugging Face Daily Papers TIER_1 English(EN) ·

    LLM智能体时代的图工程:从个体智能到系统智能

    LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolution has produced paradigms including Prompt Engineering to elicit model capabilities, Context Engineering to manage information access, Harness Engineering to organi…

  10. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Yi Chang ·

    LLM智能体时代的图工程:从个体智能到系统智能

    LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolution has produced paradigms including Prompt Engineering to elicit model capabilities, Context Engineering to manage information access, Harness Engineering to organi…

  11. arXiv cs.AI TIER_1 English(EN) · Yiyang Luo, Yihang Jiang, Qijun Xie, Liang Lan, Lin Willian Cong, Anyi Rao, Yunya Song ·

    ReguSim:评估大型语言模型代理在金融合规中的规则接地性

    arXiv:2608.19974v1 Announce Type: new Abstract: LLM agents in financial markets may cite rules yet still submit orders that violate executable constraints or misread surveillance evidence. We introduce ReguSim, a controlled financial-compliance environment, and ReguBench, a targe…

  12. arXiv cs.AI TIER_1 English(EN) · Bo Qian, Yuting Wu, Shuang Zeng, Huaiyu Wan, Dalin Zhang, Jiqiang Liu ·

    MileGPO:基于图的策略优化长时域LLM智能体的里程碑式推理与本地证据

    arXiv:2608.19803v1 Announce Type: cross Abstract: Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision often comes only from final rewards. Existing methods refine trajectory-level signals into step-level credits through step grouping…

  13. arXiv cs.AI TIER_1 English(EN) · Yizhe Chi, Wenyi Li, Deyao Hong, Xiaoqiu Wang, Mingju Gao, Kaisen Yang, Bingxiang He, Youjie Zheng, Calvin Xiao, Qinhuai Na ·

    AI4AI-Bench:用于递归自改进算法设计的LLM代理基准测试

    arXiv:2608.20318v1 Announce Type: new Abstract: Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inherits the improvement. That process is the training algorithm: a better objective or update rule…

  14. arXiv cs.AI TIER_1 English(EN) · Yiyang Feng, Biddut Sarker Bijoy, Niranjan Balasubramanian, Jiawei Zhou ·

    分解与传递:LLM 智能体中的跨任务技能迁移

    arXiv:2608.20274v1 Announce Type: new Abstract: Large language model (LLM) agents can induce skills from completed tasks and reuse them later to grow more capable with experience. In practice, induced skills may transfer unreliably and can even harm the agent that retrieves them.…

  15. arXiv cs.AI TIER_1 English(EN) · Yu Chen, Ruishuo Chen, Xun Wang, Zhuoran Li, Longbo Huang ·

    具有可证明双标准保证的 LLM 代理的最佳技能选择

    arXiv:2608.19993v1 Announce Type: new Abstract: Loading reusable skill documents into a bounded context window is now the primary way large language model (LLM) agents acquire task-specific capabilities, which makes skill selection a first-order determinant of task performance an…

  16. arXiv cs.AI TIER_1 English(EN) · Seongjae Kang, Taehyung Yu, Sung Ju Hwang ·

    PolicyGuide:从守护单一动作到指导合规性LLM代理的整个工作流

    arXiv:2608.19861v1 Announce Type: new Abstract: Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures arise from either forbidden actions, such as granting an ineligible change, or omitted procedural requirements, such a…

  17. arXiv cs.CL TIER_1 English(EN) · Mehdi Azarafza, Faezeh Pasandideh, Ali Ehteshami Bejnordi, Stefan Henkler, Achim Rettberg ·

    利用LLM的常识推理能力进行多智能体协同以实现自动驾驶

    arXiv:2608.20129v1 Announce Type: cross Abstract: Autonomous vehicles require robust perception and decision-making capabilities to operate in diverse and unseen scenarios. While reinforcement learning and rule-based methods can provide effective control and safety mechanisms, th…

  18. arXiv cs.CL TIER_1 English(EN) · Hexi Wang, Yujia Zhou, Bangde Du, Weihang Su, Xinyuan Cao, Qingyi Pan, Qingyao Ai, Yueyue Wu, Min Zhang, Yiqun Liu ·

    通过纵向生命轨迹减轻大型语言模型代理中的身份本质主义

    arXiv:2608.19621v1 Announce Type: new Abstract: Large language models (LLMs) offer a scalable approach to social simulation, but their credibility depends on how agents are constructed. Existing methods can partially reproduce population-level patterns, yet often fail to capture …

  19. Hugging Face Daily Papers TIER_1 English(EN) ·

    LLM智能体时代的图工程:从个体智能到系统智能

    Graph Engineering organizes multi-agent LLM systems through dynamic graph structures to coordinate specialized agents and manage complex, evolving tasks.

  20. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Achim Rettberg ·

    利用LLM的常识推理能力进行多智能体协同以实现自动驾驶

    Autonomous vehicles require robust perception and decision-making capabilities to operate in diverse and unseen scenarios. While reinforcement learning and rule-based methods can provide effective control and safety mechanisms, their performance may degrade in situations requirin…

  21. arXiv cs.CL TIER_1 English(EN) · Ting-Wei Li, Yuanchen Bei, Xiao Lin, Hanghang Tong ·

    超越基于LLM的推理:轻量级GNN用于智能体故障归因

    arXiv:2608.18575v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent systems (MAS) often exhibit complex failure modes, which frequently cause agents to produce incorrect outcomes. This motivates the task of Agent Failure Attribution: given a failed multi-…

  22. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Reza Zakerian ·

    LLM 智能体何时提供帮助?自动驾驶汽车边缘的截止感知混合关键性任务调度

    Autonomous vehicles offload latency-sensitive perception tasks to nearby mobile edge computing (MEC) servers, where a missed safety-critical task is unsafe rather than merely degraded. Large language models (LLMs) are increasingly proposed as adaptive, explainable schedulers, yet…

  23. Hugging Face Daily Papers TIER_1 English(EN) ·

    PolicyGuide:从守护单一动作到指导合规性LLM代理的整个工作流

    Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures arise from either forbidden actions, such as granting an ineligible change, or omitted procedural requirements, such as identification or confirmation. Runtime safegu…

  24. arXiv cs.AI TIER_1 English(EN) · Dayang Liang, Liyuan He, Xuan Feng, Shuxin Li, Bo An, Yunlong Liu ·

    PlanPO:面向多轮Agentic LLM的群组规划感知策略优化

    arXiv:2608.17289v1 Announce Type: new Abstract: Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactive tasks. However, most existing variants fail to distinguish advantages among successful traj…

  25. arXiv cs.LG TIER_1 English(EN) · Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee ·

    Agentic ESOpt:以极低的 GPU 需求微调长时域 LLM 智能体

    arXiv:2608.17310v1 Announce Type: new Abstract: Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyw…

  26. arXiv cs.AI TIER_1 English(EN) · Liangtao Lin, Qingang Zhang, Zhaomeng Zhu, Tianwei Zhang, Yonggang Wen ·

    面向任务的LLM智能体在关键任务基础设施运营中的线束配置

    arXiv:2608.17433v1 Announce Type: new Abstract: LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take…

  27. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Sandeep P. Chinchali ·

    贝叶斯伙伴建模实现LLM协调的自适应再规划

    Multi-agent Large Language Model (LLM) systems often struggle to collaborate with new teammates whose strategies shift mid-task. Because agents execute multi-step or temporally extended skills, they frequently continue executing outdated plans long after public evidence shows tha…

  28. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Yonggang Wen ·

    面向任务感知的LLM智能体在关键任务基础设施运维中的资源配置

    LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take. Existing systems often expose the same compreh…

  29. arXiv cs.AI TIER_1 English(EN) · Kaixiang Wang, Yidan Lin, Jiong Lou, Jie Li ·

    BRA-Audit:通过累积暴露审计点放置对LLM多智能体系统进行预算运行时审计

    arXiv:2608.14668v1 Announce Type: cross Abstract: LLM-based multi-agent systems (LLM-MAS) solve complex tasks through specialized collaboration, but inter-agent dependencies can propagate hallucinated or malicious outputs into system-level failures. Auditor agents mitigate these …

  30. arXiv cs.CL TIER_1 English(EN) · Ruiyao Xu, Tiankai Yang, Wei-Chieh Huang ·

    HyperSkill:通过超图结构技能记忆实现自演化大型语言模型代理

    arXiv:2608.16114v1 Announce Type: new Abstract: As agentic tasks grow in complexity, LLM agents increasingly rely on experiential memory to reuse procedural knowledge across tasks. Effective memory design must jointly address what to store, how memory is structured and retrieved,…

  31. arXiv cs.AI TIER_1 English(EN) · Victor Ye Dong, Reid Pryzant, Yi Liu, Jian Jiao ·

    CAPO: 约束感知提示优化用于LLM代理

    arXiv:2608.16068v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as agents that rely on system prompts to use tools and complete tasks. Such deployments impose distinct operational requirements, including appropriate tool use, concise promp…

  32. arXiv cs.AI TIER_1 English(EN) · Zhengzhao Ma. Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun ·

    从序列到结构:LLM代理的关系不确定性传播

    arXiv:2608.16002v1 Announce Type: cross Abstract: Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments. Existing UQ methods largely rely on local signals, such as token probabilities, predictive …

  33. arXiv cs.AI TIER_1 English(EN) · Xiao Wang, Lu Dong, Ifeoma Nwogu, Srirangaraj Setlur, Venu Govindaraju ·

    MistyPilot:通过多智能体大语言模型技能编排实现社交机器人控制

    arXiv:2608.15549v1 Announce Type: cross Abstract: Programming small social robots from natural-language instructions requires more than invoking isolated APIs. Interactive tasks combine reactive physical behaviors with stateful social behaviors, while existing interfaces often re…

  34. arXiv cs.AI TIER_1 English(EN) · Puyu Zeng, Qibing Ren ·

    超越直接访问:LLM 代理中的资源劫持

    arXiv:2608.15108v1 Announce Type: cross Abstract: Large language model agents are increasingly connected to high-value resources such as computing infrastructure, credentials, usage budgets, identities, private knowledge, communication channels, and organizational workflows. Exis…

  35. arXiv cs.AI TIER_1 English(EN) · Zeyuan Li (Massachusetts Institute of Technology), Lukas Petersson (Andon Labs), Alessandro Acquisti (Massachusetts Institute of Technology), Michiel A. Bakker (Massachusetts Institute of Technology) ·

    长时域多智能体LLM商业化中的涌现式失调沟通

    arXiv:2608.14825v1 Announce Type: cross Abstract: Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elicitation ev…

  36. arXiv cs.AI TIER_1 English(EN) · Yuan Guo, Yilong Chen, Chao Hu, Xianghao Yu, Liang Hong, Jie Xu ·

    WARA:通过闭环大语言模型代理实现自动化无线优化研究

    arXiv:2608.14573v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly capable of tool use, code execution, artifact inspection, and iterative revision, creating new opportunities for automating scientific and engineering research. To the best of our…

  37. arXiv cs.AI TIER_1 English(EN) · Veit Laule, Jiangtao Shuai, Manfred Hauswirth, Sonja Schimmler ·

    PDDLCoder:用于LLM辅助符号规划的Agentic PDDL生成

    arXiv:2608.16637v1 Announce Type: new Abstract: LLMs remain unreliable for long-horizon planning, often generating logically inconsistent or non-applicable plans. Recent hybrid methods instead translate natural language into the Planning Domain Definition Language (PDDL), allowin…

  38. arXiv cs.AI TIER_1 English(EN) · Pouya Ghiasnezhad Omran, Michael Zimmermann, Duncan Cambridge, Ashmita Kapoor, Tanya Dixit ·

    Agent Gym:通过人机协作反馈实现大型语言模型智能体持续评估与进化的框架

    arXiv:2608.15591v1 Announce Type: new Abstract: Large Language Model (LLM) agents deployed in production environments face a fundamental tension: the agent's behavior is frozen at deployment time, while the business rules and edge cases it must handle continue to evolve. Existing…

  39. arXiv cs.AI TIER_1 English(EN) · Md Fazley Rafy ·

    TwinGridShield:LLM网格代理行为的后果感知运行时授权

    arXiv:2608.15391v1 Announce Type: new Abstract: Large language model (LLM)-assisted energy-management tools can translate natural-language context into structured grid commands, but syntactic validity does not imply physical admissibility. This paper presents TwinGridShield, a mo…

  40. arXiv cs.AI TIER_1 English(EN) · Wael Albayaydh, Rui Zhao ·

    大型语言模型代理能否理性谈判?一个用于A2A/MCP上可验证多代理交互的机制设计框架

    arXiv:2608.14613v1 Announce Type: new Abstract: Modern LLM-agent frameworks increasingly interoperate through standards such as Anthropic's Model Context Protocol (MCP) for agent-to-tool access and Google's Agent2Agent (A2A) protocol for agent delegation and negotiation. However,…

  41. arXiv cs.AI TIER_1 English(EN) · Pierre Dantas, Lucas Cordeiro, Ehsan Nowroozi, Tihanyi Norbert ·

    迈向安全的LLM智能体:规范、验证与执行综述

    arXiv:2608.14590v1 Announce Type: new Abstract: LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use of tools. However, no existing system provides formally grounded, task-level safety guarante…

  42. arXiv cs.AI TIER_1 English(EN) · Prabhjot Singh, Bhushan Pawar ·

    幻觉滚雪球:将模型错误传播建模为多智能体LLM管道中的状态转换

    arXiv:2608.14588v1 Announce Type: new Abstract: Sequential multi-agent LLM pipelines chain specialized agents without verification at handoffs, creating a structural flaw with measurable and severe consequences. We show that hallucinations injected at Stage 1 do not merely persis…

  43. arXiv cs.AI TIER_1 English(EN) · Teoman Kaman ·

    何时沟通:信念分布与KL散度用于多智能体强化学习中的原则性门控

    arXiv:2608.14559v1 Announce Type: new Abstract: Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when? Existing approaches either communicate at every timestep or learn a binary gate through REINFOR…

  44. arXiv cs.LG TIER_1 English(EN) · Jiecheng Zhou, Qinghao Hu, Peng Sun, Xingcheng Zhang, Weiming Zhang ·

    Belayer:LLM 智能体强化学习训练的高效容错机制

    arXiv:2608.14635v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly trained with reinforcement learning in long-horizon, sandboxed environments. Unlike conventional RL, agentic RL couples GPU-intensive rollout engines with stateful environment con…

  45. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agentic ESOpt:以极低的GPU需求微调长时域LLM智能体

    Agentic ESOpt uses evolution strategies for scalable full-parameter fine-tuning of long-horizon LLM agents via trajectory-level reward-weighted updates and parameter-context co-evolution.

  46. Hugging Face Daily Papers TIER_1 English(EN) ·

    MUSE:一个用于理解和引导 LLM 驱动的数据科学系统的交互式元代理

    Recent advances in large language models have enabled a new class of agentic data science systems that allow users to complete complex data science workflows through natural language. Although these systems can significantly reduce manual effort, it remains difficult to diagnose …

  47. arXiv cs.AI TIER_1 English(EN) · Yu Zhuang, Kefei Chen, Yitong Duan, Shuxin Zheng, Jian Li, Xu-Yao Zhang ·

    AgentRewind:面向长时序LLM智能体的可恢复执行

    arXiv:2608.14380v1 Announce Type: new Abstract: Many real-world tasks require LLM agents to interact with their environments over long execution horizons. Errors that occur early in execution may propagate through both the agent context and environment state, and their effects ma…

  48. arXiv cs.AI TIER_1 English(EN) · Ismail El Hamraoui, Sagar Jose, Nicolas Bureau, Robert Plana ·

    面向自主LLM智能体结构化漂移诊断与恢复的基于图的强化学习框架

    arXiv:2608.14109v1 Announce Type: new Abstract: Autonomous LLM agents are increasingly deployed in complex real-world workflows, yet they remain vulnerable to runtime behavioral drift, a silent deviation from the original task that can lead to irreversible side effects on externa…

  49. arXiv cs.LG TIER_1 English(EN) · Ignacio D. Lopez-Miguel, Andreas Happe, J\"urgen Cito, Ezio Bartocci, Bettina K\"onighofer, Martin Tappler ·

    ATLAS:通过 LLM 指导的抽象和自动机学习发现代理策略

    arXiv:2608.14352v1 Announce Type: cross Abstract: Large Language Model (LLM)-based agents are increasingly used for complex tasks such as software testing and cybersecurity assessment. While these agents demonstrate impressive capabilities, their behavior is difficult to understa…

  50. arXiv cs.AI TIER_1 English(EN) · Xiaofan Zhou, Huy Nguyen, Bo Yu, Chenxi Liu, Lu Cheng ·

    多轮LLM推理的自适应停止

    arXiv:2604.01413v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) increasingly rely on multi-turn reasoning and interaction, such as adaptive retrieval-augmented generation (RAG) and ReAct-style agents, to answer difficult questions. These methods improve acc…

  51. arXiv cs.AI TIER_1 English(EN) · Heming Fu, Shan Lin, Qianqian Xie, Guojun Xiong ·

    并非所有Token都相等:面向Agentic LLM系统的通胀感知路由

    arXiv:2608.13571v1 Announce Type: cross Abstract: When a language model fails to answer a query on the first attempt, an agentic system retries, consuming additional tokens each time. This retry overhead creates a gap between what a model's per-token price implies and what a full…

  52. arXiv cs.AI TIER_1 English(EN) · Zhizhao Guan, Chen Huang, Ziming Liu, Hongru Liang, Wenqiang Lei, See-Kiong Ng, Tat-Seng Chua, Anthony G Cohn ·

    拨开迷雾:迈向在大型语言模型代理中安装和优化主动探索能力

    arXiv:2608.14339v1 Announce Type: new Abstract: We study proactive exploration in LLM agents, i.e., the ability to explore an environment to acquire information that improves future decision-making. In this regard, we first identify two fundamental bottlenecks that hinder this ca…

  53. Hugging Face Daily Papers TIER_1 English(EN) ·

    从序列到结构:LLM代理的关系不确定性传播

    RUPA models agent execution as a dependency graph to propagate uncertainty across long trajectories, improving failure detection and confidence estimation for LLM agents.

  54. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Travis Smith ·

    小小科学家:通过科学方法驱动的LLM智能体发现

    What happens when you teach an LLM-based agent the scientific method? Motivation: Scientific discovery emerges from cycles of hypothesis, implementation, empirical testing, and feedback. Can this process be automated? We approach automated algorithm design through the lens of the…

  55. Hugging Face Daily Papers TIER_1 English(EN) ·

    TwinGridShield:LLM网格代理行为的后果感知运行时授权

    Large language model (LLM)-assisted energy-management tools can translate natural-language context into structured grid commands, but syntactic validity does not imply physical admissibility. This paper presents TwinGridShield, a model-independent runtime authorization layer that…

  56. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Michiel A. Bakker ·

    长周期多智能体LLM商业化中的涌现式错位沟通

    Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elicitation evaluations on single agents or stylized tasks. Its …

  57. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Michiel A. Bakker ·

    长时域多智能体LLM商业化中的涌现式失调沟通

    Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elicitation evaluations on single agents or stylized tasks. Its …

  58. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Robert Plana ·

    面向自主LLM智能体结构化漂移诊断与恢复的基于图的强化学习框架

    Autonomous LLM agents are increasingly deployed in complex real-world workflows, yet they remain vulnerable to runtime behavioral drift, a silent deviation from the original task that can lead to irreversible side effects on external systems. Existing approaches address drift at …

  59. arXiv cs.AI TIER_1 English(EN) · Chang Liu, Yuqi Zhang, Yiman Zhong, Boyi Liu, Hengjun Wang, Shuyue Wei ·

    SkillShapley: 边界自适应Shapley值用于LLM智能体中的技能步骤归因

    arXiv:2608.13173v1 Announce Type: new Abstract: Agent skills are crucial external instructions that enable language agents to execute long procedural tasks such as coding or document processing. Existing agent skills are primarily created through human manual crafting or agent ex…

  60. arXiv cs.AI TIER_1 English(EN) · Qinwu Xu, Zhuoheng Li, Jessie Salas ·

    通过代理评估和稳定性感知排序实现多模态大模型的鲁棒性检查点选择

    arXiv:2605.18852v2 Announce Type: replace-cross Abstract: Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and downstream evaluation signals are noisy. Small observed differences can be compara…

  61. arXiv cs.AI TIER_1 English(EN) · Jiajun Ruan, Peiyang Li, Yukun Chen, Fengting Li, Chao Feng ·

    超越手工安全:迈向LLM智能体自进化防御

    arXiv:2608.12977v1 Announce Type: cross Abstract: The expanding operational capabilities of large language model (LLM) agents introduce sophisticated security threats. Runtime defenses have emerged as an effective approach to mitigating these risks by integrating security mechani…

  62. arXiv cs.AI TIER_1 English(EN) · Junzhi Li, Peng He, Qirui Ji, Wei Wang, Lixiang Liu, Chuxiong Sun ·

    通过因果推理发现基于LLM的多智能体系统的有效且可解释的通信拓扑

    arXiv:2608.12921v1 Announce Type: cross Abstract: The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn communication topologies through b…

  63. arXiv cs.AI TIER_1 English(EN) · Yanwen Peng, Delvin Ce Zhang, Xi Wang, Nikolaos Aletras ·

    StateBridge:LLM多智能体系统中无训练的隐藏状态对齐用于潜在通信

    arXiv:2608.13317v1 Announce Type: new Abstract: Large language model based multi-agent systems usually communicate in text, i.e., using discrete tokens. However, text introduces a discrete bottleneck. Converting the sender's continuous hidden states into discrete tokens discards …

  64. arXiv cs.AI TIER_1 English(EN) · Zechuan Wang, Siyuan Lu, Hongxuan Zhang, Linjian Mo, Chenyi Zhuang, Leilei Gan ·

    传授幅度而非方向:多轮多步LLM代理的验证器约束信用分配

    arXiv:2608.13179v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignment conflates heterogeneous per-turn outcomes into a…

  65. arXiv cs.AI TIER_1 English(EN) · Xutao Mao, Liangjie Zhao, Xiang Zheng, Cong Wang ·

    熟能生险:自改进 LLM 智能体中的技能错位进化

    arXiv:2608.12851v1 Announce Type: new Abstract: Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input disappears. Skill evolution makes this failure measurable by…

  66. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Chuxiong Sun ·

    通过因果推断发现基于LLM的多智能体系统的有效且可解释的通信拓扑

    The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn communication topologies through black-box optimization driven solely by task-level …

  67. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Chuxiong Sun ·

    通过因果推断发现基于LLM的多智能体系统的有效且可解释的通信拓扑

    The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn communication topologies through black-box optimization driven solely by task-level …

  68. arXiv cs.AI TIER_1 English(EN) · Dongyang Ao, Kaixiang Fang, Shijie Xu ·

    RecSys Factory: 将 LLM Agent 自主性限制在工业推荐器生命周期中的决策点

    arXiv:2608.11241v1 Announce Type: new Abstract: Deploying LLM agents into industrial recommender operations exposes a three-way tension we frame as the autonomy-determinism-efficiency trilemma: general autonomy (interpreting operator intent, generating glue code zero-shot), indus…

  69. arXiv cs.AI TIER_1 English(EN) · Touseef Hasan, Mounika Ghanta, Souvika Sarkar, Ujjwal Guin ·

    AgenticTwin:集成数字孪生的代理式LLM框架用于异常检测

    arXiv:2608.11679v1 Announce Type: new Abstract: Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems. Even with skilled operators, interpreting anomalies detected within digital twin pipelines is challenging, as the sheer complexity a…

  70. arXiv cs.AI TIER_1 English(EN) · Ziqiang Wan, Li Gu, Zhixiang Chi, Zhi Liu, Seyed Mehdi Ayyoubzadeh, Yuanhao Yu, Yang Wang ·

    基准测试LLM裁判以评估移动代理

    arXiv:2608.11434v1 Announce Type: new Abstract: Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark for s…

  71. arXiv cs.AI TIER_1 English(EN) · Pardis Taghavi, Santosh Bhavani ·

    从数字到判断:专业大语言模型Agent与强化学习在欧洲上市房地产中的应用

    arXiv:2608.11381v1 Announce Type: new Abstract: We study whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM specialization. Larix maps a 16-lens European listed-real-estate analysis framework to eight lens-…

  72. arXiv cs.AI TIER_1 English(EN) · Igor Itkin ·

    穷人的代理建模:在笔记本电脑上模拟大型LLM-代理社会

    arXiv:2608.11215v1 Announce Type: new Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the co…

  73. arXiv cs.AI TIER_1 English(EN) · Alexander Liss, Nicholas Desmond, Santiago Gil Gallego ·

    面向协作式对话结果的多大型语言模型(LLM)代理系统的动态治理

    arXiv:2608.11207v1 Announce Type: new Abstract: When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach…

  74. arXiv cs.CL TIER_1 English(EN) · Yutao Mou, Pengfei Yang, Zhe Yin, Zhangchi Xue, Xiaotian Luan, Dingyao Yu, Tong Zhang, Shikun Zhang, Wei Ye ·

    ToolHazard:为基于LLM的代理的安全评估和对齐扩展对抗性环境

    arXiv:2608.11878v1 Announce Type: cross Abstract: Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments,…

  75. arXiv cs.AI TIER_1 English(EN) · Junliang Liu, Ruoyu Li, Wenxin Tang, Jingyu Xiao, Zhenyu Liu, Jingheng Xu, Laizhong Cui ·

    收敛绕道劫持:基于技能的大语言模型代理中的任务保留资源放大

    arXiv:2608.12273v1 Announce Type: cross Abstract: LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning. This progressive-disclosure design exposes two sequential control points to untrusted publi…

  76. arXiv cs.AI TIER_1 English(EN) · Josef Liyanjun Chen ·

    就绪队列:界定GPU机遇并避免LLM-Agent控制中的主机往返

    arXiv:2608.12123v1 Announce Type: cross Abstract: LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enough concurrent work for GPU exe…

  77. arXiv cs.AI TIER_1 English(EN) · Dylan Bouchard, Mohit Singh Chauhan ·

    超越单轮置信度:面向LLM智能体的轨迹自适应不确定性量化

    arXiv:2608.11552v1 Announce Type: cross Abstract: Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive …

  78. arXiv cs.AI TIER_1 English(EN) · Ruoxi Zhao, Maziar Raissi ·

    Backtrader-Bench:使用自生成多项选择题对算法交易中的 LLM Agent 进行基准测试

    arXiv:2608.11232v1 Announce Type: cross Abstract: Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs require ground truth from actual code execution. We present Backtrader-Bench, a fram…

  79. arXiv cs.AI TIER_1 English(EN) · Gen Dong, Yanjie Gao, Liqun Li, Tianyin Xu, Yu Hua, Fan Yang ·

    Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents

    arXiv:2608.11888v1 Announce Type: new Abstract: Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, including planning, tool use, problem-solving, and validation. Prior work reported mixed results …

  80. Hugging Face Daily Papers TIER_1 English(EN) ·

    就绪队列:限制GPU机遇并避免LLM-Agent控制中的主机往返

    LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enough concurrent work for GPU execution, and what changes when a GPU-computed route…

  81. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents

    Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, including planning, tool use, problem-solving, and validation. Prior work reported mixed results of agent skills: some skills improve task succes…

  82. Hugging Face Daily Papers TIER_1 English(EN) ·

    ToolHazard:为基于LLM的代理的安全评估和对齐扩展对抗性环境

    Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefi…

  83. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Ujjwal Guin ·

    AgenticTwin:集成数字孪生的代理式LLM框架用于异常检测

    Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems. Even with skilled operators, interpreting anomalies detected within digital twin pipelines is challenging, as the sheer complexity and volume of raw sensor data make thorough analy…

  84. arXiv cs.AI TIER_1 English(EN) · Yiyang Li, Weixiang Sun, Tianyi Ma, Kaiwen Shi, Zheyuan Zhang, Yanfang Ye ·

    HoosierHelp:为社会服务导航进行大语言模型代理基准测试

    arXiv:2608.09946v1 Announce Type: cross Abstract: Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints. Although LLM agents offer a promising interface for conversational resource navigation, existin…

  85. arXiv cs.AI TIER_1 English(EN) · Vasundra Srinivasan ·

    一种用于生产环境LLM代理运行时架构模式选择与组合的方法

    arXiv:2605.20173v2 Announce Type: replace Abstract: Production LLM agents combine stochastic model outputs with deterministic software systems, yet the boundary between the two is rarely treated as a first-class architectural object. This paper names that boundary the stochastic-…

  86. arXiv cs.AI TIER_1 English(EN) · You Lu, Kun Zhang, Bihuan Chen, Xin Peng ·

    DOCSCHISEL:面向LLM智能体的自适应工具文档优化框架

    arXiv:2608.10037v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical grounding resource for LLM agents. Existing studies mainly focus on improving the tool-u…

  87. arXiv cs.AI TIER_1 English(EN) · Vitalii Belov, Artyom Sosedka, Andrey Sakhovskiy, Elizaveta Kovtun, Artyom Boyarskikh, Semen Budennyy ·

    LLM Agents Factory:领域特定LLM Agent的检索

    arXiv:2608.09934v1 Announce Type: cross Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deployment is often limited by the computational cost and instability associated with the…

  88. arXiv cs.AI TIER_1 English(EN) · Vassilis Papadopoulos, McNair Shah, Sam Zimmerman, Jack Lindsey ·

    心智病毒:多智能体LLM系统中的自我传播思想

    arXiv:2608.10218v1 Announce Type: new Abstract: AI agents are becoming more autonomous and increasingly interconnected, exposing them to new emergent risks arising from agent-to-agent interaction. One such risk is the spread of mind viruses: ideas or goals that propagate through …

  89. arXiv cs.CL TIER_1 English(EN) · Xiaojun Wu, Cehao Yang, Honghao Liu, Xueyuan Lin, Wenjie Zhang, Zhichao Shi, Xuhui Jiang, Chengjin Xu, Jia Li, Jian Guo ·

    Bayesian-Agent: 后验引导的技能演化跨越LLM代理工具集

    arXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these assets through heuristic reflection or raw success counts. Such updates are brit…

  90. arXiv cs.CL TIER_1 English(EN) · Xiaozhe Li, Yongkang Chen, Shujian Deng, Peiji Li, Yichuan Ma, Huaxi Huang, Qiye Cai, Tianyi Lyu, Le Ma, Linyang Li, Qipeng Guo, Dahua Lin, Kai Chen ·

    InternAgentHarness: 一个可扩展的合成环境,用于增强LLM的代理能力

    arXiv:2508.08636v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly expected to act as generalist agents capable of solving complex real-world problems. Training such agents, however, requires stable and diverse environments that support repeated int…

  91. arXiv cs.CL TIER_1 English(EN) · Ying Yuan ·

    检测到效应不等于学会对其采取行动:LLM获取代理的奖励信噪比下限

    arXiv:2608.10441v1 Announce Type: cross Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth using…

  92. arXiv cs.CL TIER_1 English(EN) · Xinying Cai, Minghao Guo, Jiahe Liu, Jiaojiao Han, Bangwei Guo, Yitao Long, Yuxuan Chen, Bohan Wu, Dimitris N. Metaxas, Raymond Li ·

    OpenPM:LLM投资组合管理代理的可审计即时评估

    arXiv:2608.09988v1 Announce Type: cross Abstract: Large language models are increasingly used to read markets, assess risk, and allocate capital. However, reported results for LLM trading agents can be inflated by look-ahead leakage, optimistic execution, and risk mandates that a…

  93. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越单轮置信度:面向LLM智能体的轨迹自适应不确定性量化

    Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying que…

  94. Hugging Face Daily Papers TIER_1 English(EN) ·

    就绪的群组:限制GPU机会并避免LLM-Agent控制中的主机往返

    Two studies define measurable GPU control gates for LLM-agent services by analyzing concurrent cohort scheduling and on-device routing versus host redispatch.

  95. Hugging Face Daily Papers TIER_1 English(EN) ·

    ToolHazard:为基于LLM的代理的安全评估和对齐扩展对抗性环境

    ToolHazard is a scalable framework that synthesizes adversarial environments to test LLM agents against indirect prompt injections, revealing vulnerabilities and improving defensive alignment.

  96. arXiv cs.AI TIER_1 English(EN) · Jiashu He, Jinxuan Fan, Bowen Jiang, Ignacio Houine, Dan Roth, Alejandro Ribeiro ·

    SAKE:基于强化学习的复杂LLM推理结构化智能体知识外推

    arXiv:2505.15062v5 Announce Type: replace-cross Abstract: Knowledge extrapolation is the process of inferring novel information by combining and extending existing knowledge that is explicitly available. It is essential for solving complex questions in specialized domains where r…

  97. arXiv cs.AI TIER_1 English(EN) · Bohan Chen, Shivam N. Patel, Richard Hoffmann, Sam Looi, Tony Yue Yu ·

    学习协调符号工具:LLM 代理用于验证平方和证书

    arXiv:2608.00326v2 Announce Type: replace Abstract: Tool calling allows large language models (LLMs) to invoke external computation during problem solving, a useful capability in various fields including AI for mathematics. We study this setting through weighted sum-of-squares (S…

  98. arXiv cs.AI TIER_1 English(EN) · Jiyong Kwon, Ujin Jeon, Sooji Lee, Guang Lin ·

    AIVV:用于可信赖自主系统的神经符号大模型智能体集成验证与确认

    arXiv:2604.02478v2 Announce Type: replace Abstract: Deep learning models excel at detecting anomaly patterns in normal data. However, they do not provide a direct solution for anomaly classification and scalability across diverse control systems, frequently failing to distinguish…

  99. arXiv cs.AI TIER_1 English(EN) · Thassilo M. Schiepanski, Nicholas Pi\"el ·

    超越像素:探索用于 LLM 驱动的网络代理的 DOM 降采样

    arXiv:2508.04412v3 Announce Type: replace Abstract: The advent of large language models (LLMs) has sparked an evolution of autonomous web browsing agents: given a web browsing task and serialised user interface (UI) state, an LLM is expected to suggest input actions that incremen…

  100. arXiv cs.AI TIER_1 English(EN) · Gordon Dai, Weijia Zhang, Jinhan Li, Siqi Yang, Chidera Onochie lbe, Srihas Rao, Arthur Caetano, Misha Sra ·

    人工智能利维坦:通过霍布斯社会契约论视角探索大型语言模型(LLM)代理的社会进化

    arXiv:2406.14373v3 Announce Type: replace Abstract: The emergence of Large Language Models (LLMs) and advancements in Artificial Intelligence (AI) offer an opportunity for computational social science research at scale. Building upon prior explorations of LLM agent design, our wo…

  101. arXiv cs.AI TIER_1 English(EN) · Rohan Bhagra, Mahantesh Halapannavar, Uddhav Bhattarai ·

    Agentic Harnesses:LLM驱动的机器人自主性验证层

    arXiv:2608.09857v1 Announce Type: cross Abstract: Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on execution rather than verifying the feasibility actions planning models propose…

  102. arXiv cs.AI TIER_1 English(EN) · Nuthakki Siva Gopala Krishna, Kanishka Jain ·

    STEMMA:一个用于评估大型语言模型自我身份一致性的对抗性多智能体框架

    arXiv:2608.08164v1 Announce Type: cross Abstract: Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of structured information and functional behavior from a large teacher model to a smaller stude…

  103. arXiv cs.AI TIER_1 English(EN) · Yingpeng Ma, Jianhao Yan, Bei Shi, Ka Hou Kam, Runnan Wang, Xuebo Liu, Yulong Chen, Yue Zhang, Derek F. Wong ·

    大型语言模型代理能否坚持剧本?一项用于交互式叙事中长时程一致性的基准测试

    arXiv:2608.08160v1 Announce Type: cross Abstract: The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the critical challenge of maintaini…

  104. arXiv cs.AI TIER_1 English(EN) · Wanying Qu, Qinghua Mao, Yu Li, Jiyao Liu, Xin Zhang, Dadi Guo, Yanxu Zhu, Qingyu Liu, Leitao Yuan, Xi Lin, Shanfeng Zhu, Yanwei Fu, Jing Shao, Xia Hu, Dongrui Liu ·

    SHE:面向LLM智能体的轨迹驱动安全带演进

    arXiv:2608.09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety mechanisms often treat the harnes…

  105. arXiv cs.AI TIER_1 English(EN) · Hao Sui, Simeng Qin, Jie Liao, Xiaojun Jia, Bing Chen, Yang Liu ·

    ElasticBack:通过耦合触发器-规则优化在LLM代理技能中实现隐蔽的条件后门

    arXiv:2608.09577v1 Announce Type: new Abstract: Agent skills, bundles of instructions and resources that an LLM agent loads on demand, form an emerging supply chain where a single poisoned skill can persistently compromise every agent that installs it. However, existing skill att…

  106. arXiv cs.AI TIER_1 English(EN) · Fatemeh Seyedin, Adrian Weller, Jinhyuk Yun, Mahmoudreza Babaei ·

    政客、骗子和顺从的工人:LLM代理在层级博弈中的新兴行为

    arXiv:2608.09574v1 Announce Type: new Abstract: LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move from individual tools to participants in multi-agent organizations, an important qu…

  107. arXiv cs.AI TIER_1 English(EN) · You Lu, Xinyu Huang, Bihuan Chen, Xin Peng ·

    SkillSentry:通过运行时保障实现 LLM Agent 的可靠技能执行

    arXiv:2608.09253v1 Announce Type: new Abstract: LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and tool use. Although skills provide reusable procedural knowledge, agents may still execute them unreliably. Even when an agent…

  108. arXiv cs.AI TIER_1 English(EN) · Liang He, Jingbo Wen, Hongyu Gu, Hao Li, Haoyu Wang, Yixiong Chen, Kangning Cui, Xilu Wang ·

    从相关性到执行效用:基于技能的大模型智能体奖励感知动态执行门控

    arXiv:2608.09168v1 Announce Type: new Abstract: Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has substantially improved skill retrieval due to the increasing skill libraries, retrieving a pl…

  109. arXiv cs.AI TIER_1 English(EN) · Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, Xiaoying Xing ·

    商业竞技场:在真实市场中对标LLM智能体

    arXiv:2608.08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligati…

  110. arXiv cs.AI TIER_1 English(EN) · Florentina Voboril, Stefan Szeider ·

    使用 LLM 代理改进约束模型

    arXiv:2608.08127v1 Announce Type: new Abstract: The runtime of Constraint Programming (CP) solvers is highly sensitive to modeling choices, such as symmetry breaking, implied constraints, global constraints, constraint reformulation, and variable representation. Improving these c…

  111. arXiv cs.AI TIER_1 English(EN) · Yang Liu, Shiwei Hou, Xiyuan Chen, Yu Wang, Sen Yuan, Qirui Gan, Shao You, Feifan Chen, Wencheng Li, Shuyang Hu, Yongzhou Liu, Emma Xia, Xiaojing Lu, Hao Wang, Fan Xu, Yanfeng Li ·

    ZhuLong: 面向EDA脚本的基于执行的LLM代理,支持离线API自主探索

    arXiv:2608.07925v1 Announce Type: new Abstract: EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing LLMs fail to address. This paper presents ZhuLong, an execution-grounded LLM coding agent for PyAether and SKILL that combines AP…

  112. arXiv cs.AI TIER_1 English(EN) · Yijie Wang, Zhen-Yu Yin, Zhenheng Tang, Xiaowen Chu ·

    Agent-MD:针对状态化 GCMC--MD 活动的事件驱动升级选择性 LLM 干预

    arXiv:2608.07637v1 Announce Type: new Abstract: Long-running molecular simulation campaigns require repeated continuation from saved states, provenance-aware progression, adaptive assessment, and occasional interpretation of workflow conditions that cannot be resolved safely by f…

  113. arXiv cs.LG TIER_1 English(EN) · Ibne Farabi Shihab, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan ·

    Stateful CARS: Exact Cross-History Reuse for Policy-Constrained LLM Agents

    arXiv:2608.08282v1 Announce Type: new Abstract: Tool-using language-model agents face constraints whose meaning changes with observations and prior actions. We study exact sampling from the model distribution conditioned on a hard stateful validator while reusing invalidity certi…

  114. arXiv cs.CL TIER_1 English(EN) · Mahesh Ramesh, Kaousheik Jayakumar, Aswinkumar Ramkumar, Pavan Thodima, Aniket Rege, Emmanouil-Vasileios Vlatakis-Gkaragkounis ·

    合作推理的火花:LLM 作为策略性 Hanabi 玩家

    arXiv:2601.18077v3 Announce Type: replace Abstract: Cooperative reasoning under incomplete information remains challenging for both humans and multi-agent systems. The card game Hanabi embodies this challenge, requiring theory-of-mind reasoning and strategic communication. We ben…

  115. arXiv cs.CL TIER_1 English(EN) · Bingzhen Liu, Xiaomeng Fan, Yuwei Wu, Zhi Gao, Mingyang Gao, Chuanhao Li, Yunde Jia ·

    超越能力边界:用于自进化LLM智能体的零阶优化

    arXiv:2608.09292v1 Announce Type: cross Abstract: Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. However, these methods struggle to learn beyond the inherent capability boundary …

  116. arXiv cs.CL TIER_1 English(EN) · Ashritha Gonuguntla ·

    重播差距:LLM代理模型切换的静态评估得分错误的世界

    arXiv:2608.08239v1 Announce Type: cross Abstract: LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents. Yet agentic routers are evaluated like single-turn routers: by replaying logge…

  117. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Ying Yuan ·

    检测到效应不等于学会对其采取行动:LLM获取代理的奖励信噪比下限

    Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth using. Our thesis is a distinction that is easy to miss…

  118. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Francisco León Zúñiga Bolívar ·

    并非“大一统”:中国前沿LLM智能体合作博弈中的实验室级分歧

    Does the cooperative bias documented for Western frontier LLM agents extend to a different alignment lineage, and should the Chinese models that embody it be treated as a single bloc or as distinct laboratories? We study four frontier-tier Chinese models - DeepSeek V4 Pro, Qwen3-…

  119. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越能力边界:用于自进化 LLM 智能体的零阶优化

    Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. However, these methods struggle to learn beyond the inherent capability boundary of the agents, since the agents cannot sample corr…

  120. Hugging Face Daily Papers TIER_1 English(EN) ·

    从相关性到执行效用:基于技能的大语言模型智能体的奖励感知动态执行门控

    Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has substantially improved skill retrieval due to the increasing skill libraries, retrieving a plausible skill bundle does not guarantee that exe…

  121. arXiv cs.LG TIER_1 English(EN) · Elizaveta D. Moskovskaya, Anton D. Moscowsky ·

    具有多智能体控制和LLM自动场景生成的机器人导引头

    arXiv:2509.10317v2 Announce Type: replace-cross Abstract: The article describes the development of a hybrid social robot control architecture to overcome the limitations of traditional approaches, where behavior scripts manually synchronize the robot's actions and text, and exist…

  122. arXiv cs.CL TIER_1 English(EN) · Mingguang Chen, Licheng Wang, Bo Qu ·

    地平线鸿沟:长周期LLM智能体的规划、记忆、执行、训练与评估

    arXiv:2608.06663v1 Announce Type: new Abstract: Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour tasks: losing track of earlier decisions, declaring half-finished work done, or…

  123. arXiv cs.AI TIER_1 English(EN) · Jiacheng Miao, Jin Mu, Guanhua Chen, James Zou ·

    Fisher-R1:训练LLM智能体进行可靠的假设检验

    arXiv:2608.07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can inspect datasets, generate code, and produce analyses e…

  124. arXiv cs.AI TIER_1 English(EN) · Karolina Rudnicka, Thomas Stephan Juzek ·

    超越“人工智能语言”:论证LLM输出的语域特性

    arXiv:2608.06589v1 Announce Type: cross Abstract: While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that this perspective coexists with distinct, model-specific linguistic signatures akin to human id…

  125. Hugging Face Daily Papers TIER_1 English(EN) ·

    商业竞技场:在真实市场中对标LLM智能体

    Business Arena evaluates LLM agents running a realistic cross-border shop, revealing large performance gaps versus human strategies and enabling detailed attribution of business decisions.

  126. Hugging Face Daily Papers TIER_1 English(EN) ·

    大型语言模型代理能否坚持剧本?面向交互式叙事中长时程一致性的基准测试

    The study introduces a benchmark and formalizes narrative commitment preservation to evaluate long-horizon logical consistency in interactive storytelling with large language models.

  127. arXiv cs.AI TIER_1 English(EN) · Linfang Shang, Ming Xu, Yiding Sun, Tianle Xia, Lingxiang Hu, Lan Xu, Ning Zheng ·

    当自进化适得其反:预提交门控以防止大型语言模型智能体中的技能污染

    arXiv:2608.05810v1 Announce Type: new Abstract: Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not monotonic: past a critical pool size, newly added skills degrade performance instead of impr…

  128. arXiv cs.AI TIER_1 English(EN) · Zihan Xu, Haolin Tian, Hai Jiang ·

    多智能体LLM系统中推理时并行性的双层视角

    arXiv:2608.05791v1 Announce Type: cross Abstract: Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inference, and their execution strategies directly affect system accuracy, latency, and computation…

  129. arXiv cs.AI TIER_1 English(EN) · Jie Wu, Ming Gong, Feixiang Cheng, Qinqin Zhao ·

    EcoAgent-Bench:评估预算受限LLM代理的经济决策能力

    arXiv:2608.05519v1 Announce Type: new Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local lookup, broad search, composite research tool, stronger model, or human escalation i…

  130. arXiv cs.AI TIER_1 English(EN) · Jialuo Chen, Minghe Wang, Lingqi Jiang, Jianan Ma, Xinhao Deng, Xiaohu Du, Ruixiao Lin, Yunhao Feng, Linkang Du, Jingyi Wang ·

    SkillTrace:LLM-Agent 技能复用中的多轨迹溯源审计

    arXiv:2608.05204v1 Announce Type: new Abstract: LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code, tools, references, and operational workflows. As skills become marketplace artifacts, auditin…

  131. arXiv cs.AI TIER_1 English(EN) · Wuya Chen, Yihao yang, Yang Cao, Yue Lin ·

    CodeGrep:一个用于LLM编码代理的RL训练检索代理

    arXiv:2608.05886v1 Announce Type: cross Abstract: Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the file to patch, rather than patching it. On SWE-Bench Verified, a 30B OpenHands agent average…

  132. arXiv cs.CL TIER_1 English(EN) · Xuying Ning, Dongqi Fu, Tianxin Wei, Hanqing Zeng, Yuanchen Bei, Bingxuan Li, Zihao Li, Qifan Wang, Xiang Shen, Yifan Wu, Jiayi Liu, Hong Li, Yinglong Xia, Xiangjun Fan, Hanghang Tong, Jingrui He ·

    EvoHarness-RL:为长时序LLM代理学习自进化运行时线束

    arXiv:2608.05446v1 Announce Type: cross Abstract: Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, invoke tools, verify outcomes, and reuse experience across interactions. However, effective harness use raises two coupled …

  133. Hugging Face Daily Papers TIER_1 English(EN) ·

    CodeGrep:一个用于LLM编码代理的RL训练检索代理

    Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the file to patch, rather than patching it. On SWE-Bench Verified, a 30B OpenHands agent averages 23 rounds and 631K tokens per resolved issue, wi…

  134. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Hai Jiang ·

    多智能体LLM系统中推理时并行性的双层视角

    Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inference, and their execution strategies directly affect system accuracy, latency, and computational cost. Parallel execution provides a means to im…

  135. arXiv cs.AI TIER_1 English(EN) · Junnan Liu, Linhao Luo, Thuy-Trang Vu, Gholamreza Haffari ·

    EASy:迈向高效的基于LLM的代理系统

    arXiv:2608.04588v1 Announce Type: cross Abstract: Agentic systems have emerged as a promising paradigm for solving complex tasks by coordinating specialized LLM-based agents. However, most existing systems primarily optimize task success while giving limited consideration to exec…

  136. arXiv cs.AI TIER_1 English(EN) · Wenxiao Zhao, Dong Liu, Kaiyi Xu, Feng Liu, Zhen Zhao, Fei Ben, Shu Wang, Wenhao Li, Yingnian Wu, Fenghua Ling, Haobo Li, Lei Bai ·

    A-SR:用于符号回归的分层协调的自演化代理LLM

    arXiv:2608.04872v1 Announce Type: cross Abstract: Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses heterogeneous search failures into a scalar score and a single prompt. We …

  137. arXiv cs.CL TIER_1 English(EN) · Jinyi Han, Yuanjian Xu, Ying Liao, Xinyi Wang, Zishang Jiang, Zixiang Di, Fanyang Lu, Zhichao Hu, Yanghua Xiao ·

    Skill-Use:大型语言模型能否在代理式框架中实际使用技能?

    arXiv:2608.04828v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed. Existing evaluations mostly judge the quality of a skill or its co…

  138. arXiv cs.AI TIER_1 English(EN) · Peichun Hua, Haoxuan Xu, Mengyuan Li ·

    行为技能重构:从LLM智能体技能中重构隐藏功能

    arXiv:2608.04192v1 Announce Type: cross Abstract: Closed source agent skills may encode proprietary instructions, scripts, constants, and data. Providers may offer their capabilities as services while keeping the underlying packages hidden. Prior work focuses on prompt injection …

  139. arXiv cs.AI TIER_1 English(EN) · Atul Anand, Sourav Chattaraj ·

    使用 Canary Tools 诊断 LLM 代理中的工具选择推理

    arXiv:2608.04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) tool set, each engineered to probe one specific tool-…

  140. arXiv cs.AI TIER_1 English(EN) · J. de Curt\`o, I. de Zarz\`a ·

    网络物理系统中大型语言模型(LLM)智能体的规划策略的战略评估

    arXiv:2608.04265v1 Announce Type: cross Abstract: Evaluations of LLM planning agents largely ask whether a task succeeds or a declared plan is followed. In strategic cyber-physical systems, a stronger question is whether the planning architecture remains appropriate after autonom…

  141. Hugging Face Daily Papers TIER_1 English(EN) ·

    A-SR:用于符号回归的分层协调的自演化代理大型语言模型

    Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses heterogeneous search failures into a scalar score and a single prompt. We propose A-SR, a self-evolving agentic framework th…

  142. Hugging Face Daily Papers TIER_1 English(EN) ·

    使用 Canary Tools 诊断 LLM 代理中的工具选择推理

    Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) tool set, each engineered to probe one specific tool-selection weakness. A six-type taxonomy (semanti…

  143. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Wataru Toyokawa ·

    LLM智能体中基于声誉的合作的出现

    Can cooperation among large language model (LLM) agents be evolutionarily stable against free-rider invasion? We study an indirect reciprocity donation game where LLM agents observe behavioral traces and donate on a continuous scale. Strategies, represented as natural language pr…

  144. arXiv cs.AI TIER_1 English(EN) · Tianyi Guan, Yiding Wang, Haotong Yang, Siyuan Cao, Shirui Liu, Yi Hu, Jiaqi Li, Muhan Zhang ·

    ContinualSkillBench:大型语言模型代理能否真正进化其能力?

    arXiv:2608.03874v1 Announce Type: new Abstract: Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve…

  145. arXiv cs.AI TIER_1 English(EN) · Zian Zhai, Xingyu Tan, Gaowang Zou, Xiaoyang Wang, Wenjie Zhang ·

    HyperAgent:在工具模式超图上进行规划和行动,用于工具使用 LLM 代理

    arXiv:2608.02650v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external tools to complete complex real-world tasks. However, reliable tool-use planning remains challenging due to the limitations of implicit reasoning and the evolving nature…

  146. arXiv cs.AI TIER_1 English(EN) · Unggi Lee, Sookbun Lee, Yeil Jeong, Eunjoo Lee, Minchul Shin, Hoilym Kwon ·

    EduClaw-Bench:用于具有模拟学习者的教学 LLM 代理的长视野基准测试

    arXiv:2608.03206v1 Announce Type: cross Abstract: Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a l…

  147. arXiv cs.AI TIER_1 English(EN) · Qiming Shi, Yibo Dou, Jiawen Zhu, Yulong Tao, Linbo Jin, Zhaolu Kang, Yunfan Zhou, Di Weng ·

    SKILL-KD: 对比式技能蒸馏用于LLM智能体

    arXiv:2607.28048v2 Announce Type: replace Abstract: Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often treat skills as experience summaries, memory entries, or direct summaries of su…

  148. arXiv cs.AI TIER_1 English(EN) · Yue Yao, Shengyuan Wang, Xin Chen, Minke Zhang, Jia He, Bingjun Luo, Tom Gedeon ·

    SkillTrace:遍历查询-技能图谱以实现可组合的LLM代理

    arXiv:2608.02356v2 Announce Type: replace Abstract: Large language model agents increasingly solve complex tasks by composing reusable skills from a library. To address this, the key challenge is not merely to retrieve individually relevant skills, but to identify a complete and …

  149. arXiv cs.CL TIER_1 English(EN) · Ming Shen, Chao Shang, Sadat Shahriar, Devang Kulshreshtha, Yi Zhang, Sandesh Swamy, Yanjun Qi ·

    基于LLM的多智能体系统中作为收敛压力的关系先验

    arXiv:2608.03239v1 Announce Type: new Abstract: Large language model-based multi-agent systems (LLM-MAS) are designed through roles, debate protocols, and aggregation rules. These choices create implicit social expectations: agents may be expected to trust, challenge, defer to, o…

  150. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Sho Akiyama ·

    持续改进与并行自主探索:用于搜索大型解空间的LLM-Agent框架

    We present a framework that gives LLM agents two mechanisms for searching large solution spaces autonomously. First, a leaderboard scored on held-out data acts as a reward signal that drives each agent to refine its solutions over repeated submissions, a loop that operates even w…

  151. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Renato Figueiredo ·

    CURATE:利用 LLM 代理来组合、编目和部署可复现的工作流

    Agentic code generation has shown promise in automating and accelerating software development by utilizing Large Language Models (LLMs) to generate, test, and deploy code. For engineers and scientists, such systems have the potential to accelerate the development of applied and s…

  152. arXiv cs.MA (Multiagent) TIER_1 English(EN) · I. de Zarzà ·

    网络物理系统中大型语言模型(LLM)智能体的规划策略的战略评估

    Evaluations of LLM planning agents largely ask whether a task succeeds or a declared plan is followed. In strategic cyber-physical systems, a stronger question is whether the planning architecture remains appropriate after autonomous participants respond and physics constrains th…

  153. Hugging Face Daily Papers TIER_1 English(EN) ·

    ContinualSkillBench:大型语言模型代理能否真正进化其能力?

    Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge this gap, …

  154. Hugging Face Daily Papers TIER_1 English(EN) ·

    基于LLM的多智能体系统中作为收敛压力的关系先验

    Large language model-based multi-agent systems (LLM-MAS) are designed through roles, debate protocols, and aggregation rules. These choices create implicit social expectations: agents may be expected to trust, challenge, defer to, or collaborate with peers. We study the effects o…

  155. Hugging Face Daily Papers TIER_1 English(EN) ·

    SKILL-KD: 对比式技能蒸馏用于LLM代理

    Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often treat skills as experience summaries, memory entries, or direct summaries of successful demonstrations. This creates a mismatch for…

  156. Hugging Face Daily Papers TIER_1 English(EN) ·

    ContinualSkillBench:大型语言模型代理能否真正进化其能力?

    Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge this gap, …

  157. arXiv cs.AI TIER_1 English(EN) · Tianyu Huai, Tingshuo Fan, Xinchi Chen, Yining Zheng, Yuxin Wang, Shuang Chen, Jie Zhou, Xuanjing Huang ·

    AgentHPOBench:一个用于评估LLM代理作为顺序超参数优化器的基准

    arXiv:2607.29626v1 Announce Type: new Abstract: As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmarks typically focus on static code generation, paper r…

  158. arXiv cs.CL TIER_1 English(EN) · Zhenyu Zhang, Zhichao Cao ·

    TokTier:Agentic LLM服务的精确状态化Token化

    arXiv:2607.29678v1 Announce Type: new Abstract: LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, which resubmit a long transcript after each small tool result, and reuse is hard …

  159. arXiv cs.AI TIER_1 English(EN) · Duo Xu, Faramarz Fekri ·

    NeSyFS:面向部分可观测环境下的LLM智能体的神经符号快慢思维框架

    arXiv:2607.28942v1 Announce Type: new Abstract: Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retrieval-augmented generation, and scientific discovery. In these settings, agents must act based o…

  160. arXiv cs.AI TIER_1 English(EN) · Qiming Shi, Yulong Tao, Linbo Jin, Zhaolu Kang, Yibo Dou, Jiawen Zhu, Tianjun Pan, Shaokang Fu, Chengyu Wang, Siyue Li, Yaping Cheng, Di Weng, Chengfu Huo ·

    MerchantBench:对标电商运营中LLM Agent的长期一致性

    arXiv:2607.28956v1 Announce Type: new Abstract: Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to p…

  161. Hugging Face Daily Papers TIER_1 English(EN) ·

    GISAgentBench:一个从业者来源的评估LLM智能体在GIS任务上表现的基准

    Geographic Information System (GIS) professionals rely on multi-step spatial analysis workflows to support decision-making in urban planning, disaster response, and environmental monitoring. The process is tedious, time-consuming, and error-prone. While recent large language mode…

  162. arXiv cs.AI TIER_1 English(EN) · Huixiang Zhang, Mahzabeen Emu ·

    潜在通道是否真的在通信?对潜在多智能体LLM的因果审计

    arXiv:2607.26773v1 Announce Type: new Abstract: Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-r…

  163. arXiv cs.AI TIER_1 English(EN) · Zhilun Zhou, Jianghao Yu, Yuming Lin, yongjun yang, Sun Yongquan, Depeng Jin, Yong Li ·

    UrbanDS:一个图引导的大语言模型多智能体系统,用于数据密集型城市任务

    arXiv:2607.26724v1 Announce Type: new Abstract: Large language model (LLM) agents have been widely applied in automating data science tasks. However, existing methods typically rely on a limited set of provided datasets, and they face challenges in data-intensive scenarios that r…

  164. arXiv cs.AI TIER_1 English(EN) · Marylou Fauchard, Florian Carichon, Margarida Carvalho, Golnoosh Farnadi ·

    更多欺骗:混合动机LLM多智能体系统中的目标不对齐

    arXiv:2607.26120v1 Announce Type: new Abstract: Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric information and strategic deception due to conflicting or hidden objectives. In t…

  165. arXiv cs.CL TIER_1 English(EN) · I. Kennedy, T. Kennedy ·

    富达并非安全:轻度压缩的大语言模型在无数据质量保障的代理执行中通过所有已发明步骤

    arXiv:2607.28196v1 Announce Type: new Abstract: Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original, downstream accuracy (for example MMLU) inside a confidence interval, and data-fr…

  166. arXiv cs.CL TIER_1 English(EN) · Sebastian Pohl, Harsh Mehta, Pranav Mambayil, Abdul Ghafoor, Franziska Lesigang, Yufang Hou, Christian Hilbe ·

    大型语言模型在受控环境中难以模拟人类信念更新

    arXiv:2607.28347v1 Announce Type: new Abstract: LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been tested directly. We test whether six LLMs can simulate individual human belief u…

  167. arXiv cs.LG TIER_1 English(EN) · Jiwon Jang, Kisu Yang, Heuiseok Lim, Hyunwoo Park ·

    分数持平,失败加剧:错误预算如何掩盖量化大模型代理的损害

    arXiv:2607.27275v1 Announce Type: new Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless. We test this claim for multi-turn, tool-calling agents, where it now matters most. On $\tau^2$-bench, across two open-weight model families in den…

  168. arXiv cs.LG TIER_1 English(EN) · Cong Li, Peixi Peng, Yisen Zhao, Xinyu Hu, Shudong Liu, Zhan Su, Zhuojian Li ·

    TAPO:LLM智能体的感知策略优化

    arXiv:2607.27973v1 Announce Type: new Abstract: Recently, Reinforcement Learning (RL) has emerged as a crucial paradigm for the post-training of Large Language Model (LLM) agents. However, existing methods predominantly rely on sparse task rewards for policy optimization, failing…

  169. Hugging Face Daily Papers TIER_1 English(EN) ·

    MerchantBench:对标电商运营中LLM智能体长期一致性的基准测试

    Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended hori…

  170. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Rahul Rachuri ·

    LLM智能体能否进行有竞争力的定价?面向智能体商业的动态多属性拍卖基准

    Agentic commerce is moving from concept to deployed infrastructure: payment networks, retailers, and AI platforms are setting the stage for agents to transact on behalf of merchants and consumers. Yet whether the LLMs behind these agents can price competently in real markets, whe…

  171. Hugging Face Daily Papers TIER_1 English(EN) ·

    一人、N个智能体:在校准不当、相关置信度下的LLM智能体集群的审计预算分配

    A single human must audit $N$ LLM agents under a budget of $B \ll N$ audits per round, guided by self-reported confidence that may be adversarially miscalibrated and by correlated errors. We model this as budgeted noisy inspection over a two-level Gaussian copula and locate the m…

  172. Hugging Face Daily Papers TIER_1 English(EN) ·

    富达并非安全:温和压缩的LLM在无数据质量保障的代理执行中通过所有程序步骤

    Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original, downstream accuracy (for example MMLU) inside a confidence interval, and data-free output-fidelity signals that compare the comp…

  173. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Alois Knoll ·

    扩展LLM驱动的多智能体系统:设计原则与架构可扩展性分析

    LLM-based multi-agent systems have the potential to enable collective intelligence and scale toward solving highly complex tasks through coordinated ensembles of specialized agents. However, despite their theoretical potential, the architectural design space remains largely non-s…

  174. arXiv cs.LG TIER_1 English(EN) · Yicheng Feng, Yan Zhang, Yan Cheng, Wei Qi ·

    分数不是决策:LLM代理中工具获取的成本感知停止

    arXiv:2607.27083v1 Announce Type: new Abstract: As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental tool-selection challenge: acquiring too few tools leaves the task under-informed, w…

  175. arXiv cs.CL TIER_1 English(EN) · Jingxing Wang, Chenyu Zhou, Zhihui Fu, Jun Wang, Weiwen Liu, Weinan Zhang, Jianghao Lin ·

    即时技能:LLM智能体的测试时自适应技能合成

    arXiv:2605.16986v2 Announce Type: replace Abstract: Additional test-time compute can give LLM agents access to more past experience, yet expanding the context or adding rollouts does not necessarily yield greater agent capability. We call this challenge test-time compute-to-capab…

  176. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Tian Lan ·

    通过合作-义务耦合审计涌现式LLM-Agent协作

    LLM-agent systems can solve complex tasks through dynamic self-organization and emergent cooperation. Auditing this process is essential because plausible intermediate or final outputs can conceal incomplete or unsupported work and poorly allocated responsibility, ultimately comp…

  177. Hugging Face Daily Papers TIER_1 English(EN) ·

    潜在通道是否真的在通信?对潜在多智能体LLM的因果审计

    Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-relevant information. End-task performance alone …

  178. arXiv cs.AI TIER_1 English(EN) · Debjyoti Paul ·

    上下文组装作为控制变量:一种基于控制理论的冻结LLM代理策略视角

    arXiv:2607.25408v1 Announce Type: new Abstract: A growing body of 2026 work applies control theory to LLM agents: Lyapunov-certified stability for tool-mediated controllers (Prinos et al., "Stable Agentic Control", 2026), sample-complexity bounds for sparse policies over massive …

  179. arXiv cs.AI TIER_1 English(EN) · Debjyoti Paul ·

    一个控制系统、一个数据集以及一种让冻结的 LLM 代理学习某个领域的方法

    arXiv:2607.25415v1 Announce Type: new Abstract: Production LLM agents are increasingly assembled from a frozen model wrapped in a harness: a prompt template, a tool set, a memory/retrieval layer, a planning strategy, and a verification policy. Two 2026 systems, Meta-Harness (Lee …

  180. arXiv cs.AI TIER_1 English(EN) · Yan Zhang, Shibo Li ·

    ConsistencyGate:通过自洽性准入控制防止 LLM Agent 中的记忆污染

    arXiv:2607.22962v1 Announce Type: new Abstract: LLM agents that operate over many turns accumulate facts in an external memory store and reuse them as premises for downstream reasoning. A hallucinated fact written at one step therefore persists as a false premise for every subseq…

  181. arXiv cs.AI TIER_1 English(EN) · Yihui Zhang (Beihang University), Tianyu Wo (Beihang University), Jinghao Wang (Beihang University), Xiaoyang Sun (University of Leeds), Menghao Zhang (Beihang University), Cangzhou Yuan (Beihang University), Li Li (Beihang University), Chunming Hu (Beih… ·

    SpecBox:用于高效LLM代理服务的推测性沙盒调度

    arXiv:2607.23933v1 Announce Type: cross Abstract: As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces a fundamental tension between resource utilization and interactive tail latency…

  182. arXiv cs.AI TIER_1 English(EN) · Mingwei Zheng, David OBrien, Siwei Cui, Pardis Pashakhanloo, Rajdeep Mukherjee, Myeongsoo Kim, Sachit Kuhar ·

    CORVUS:通过底层同步优化和缩减上下文以用于LLM编码代理

    arXiv:2607.22711v1 Announce Type: cross Abstract: LLM coding agents operate by constructing trajectories that accumulate reasoning, tool calls, and results to enable multi-step decision-making. However, the conventional append-only trajectory architecture found in practice tightl…

  183. arXiv cs.AI TIER_1 English(EN) · Adhyyan Narang, Artin Tajdini, Claire Zhang, Jamie Morgenstern ·

    LLM微调隐藏行为的推理时共识缓解方法

    arXiv:2607.23394v1 Announce Type: new Abstract: Recent work shows that fine-tuning language models on even a small amount of poisoned data can install targeted misbehavior, and ostensibly benign data can transmit hidden preferences that generalize broadly. Standard defenses, such…

  184. arXiv cs.CL TIER_1 English(EN) · Junming Chen, Junyang Jiang, Xu Chen, Zibo Liang, Kai Zheng ·

    DBA-Bench:一个用于基于 LLM 的数据库操作代理的生产级保真度基准

    arXiv:2607.22165v1 Announce Type: cross Abstract: LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation and production operations: live-environment fidelity (multi-turn read-write inter…

  185. arXiv cs.AI TIER_1 English(EN) · Mohamed Jouini ·

    面向基础设施即代码生成的代理式大语言模型的先验验证器评估

    arXiv:2607.20478v1 Announce Type: cross Abstract: Infrastructure-as-Code (IaC) generation from natural language requires satisfying provider schemas, dependency planning, and organizational policy constraints, not merely producing syntactically plausible configurations. We presen…

  186. arXiv cs.AI TIER_1 English(EN) · Jerzy Kami\'nski, Ilya Galyukshev, Artem Kuznetsov, Sergey Chuprin, Kirill Redko, Aidar Shumbalov, Anna Kalyuzhnaya ·

    DynamicMCPBench:一个针对实时MCP服务器上LLM代理的、基于轨迹的、效果评分的基准测试

    arXiv:2607.20531v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score the final answer or a fixed "ground-truth" list of tools, both of which are fragil…

  187. arXiv cs.AI TIER_1 English(EN) · Junchi Liao ·

    审计LLM代理行动选择中的来源敏感性

    arXiv:2607.20827v1 Announce Type: new Abstract: LLM agents choose tools and arguments from context that mixes user requests, tool outputs, retrieved records, memory, and untrusted text. Evidence can be relevant without being authorized to determine a decision, so a correct action…

  188. arXiv cs.AI TIER_1 English(EN) · Elias Hossain, Md Mehedi Hasan Nipu, Tasfia Nuzhat Ornee, Rajib Rana, Niloofar Yousefi ·

    NEXUS:面向使用工具的大语言模型代理的结构化运行时安全

    arXiv:2607.19356v1 Announce Type: new Abstract: Tool-using LLM agents increasingly execute high-impact actions, making runtime safety monitoring essential. We present NEXUS (Neural EXecution Utility and Safety), a structured-plan safety monitor that applies a formal intervention …

  189. arXiv cs.AI TIER_1 English(EN) · Aarushi Singh ·

    防护栏成替罪羊:审计工具增强型LLM代理中不忠诚的安全拒绝

    arXiv:2607.19449v1 Announce Type: cross Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure failures and HTTP 200 responses with empty, null, or malformed payloads largely…

  190. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Zhonghao Hou ·

    并非物以类聚:LLM 智能体中的基于个性的伴侣选择

    Multi-agent LLM systems increasingly let one agent choose which other agents to work with, and agents are increasingly given personalities through personas. We test whether Big Five personality alone influences partner selection when capability is explicitly held constant. Host a…

  191. arXiv cs.AI TIER_1 English(EN) · Philipp J. Schneider, Lin Tian, Marian-Andrei Rizoiu ·

    学习交友:指导大型语言模型代理形成涌现的社交关系

    arXiv:2510.19299v2 Announce Type: replace Abstract: Can large language model (LLM) agents reproduce the complex social dynamics that characterize human online behavior -- shaped by homophily, reciprocity, and social validation -- and what memory and learning mechanisms enable suc…

  192. arXiv cs.AI TIER_1 English(EN) · Daisuke Kikuta ·

    AI旅游会议:LLM代理的团队旅行规划

    arXiv:2607.18806v1 Announce Type: new Abstract: This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Language Model (LLM)-based agents. The agents are instantiated with distinct personas and collaboratively seek an itinerary that satisf…

  193. arXiv cs.AI TIER_1 English(EN) · Artem Maryanskyy, Dmitry Budnikov, Alibek T. Kaliyev ·

    当代理意见不合时:多代理LLM管道中的选择瓶颈

    arXiv:2603.20324v2 Announce Type: replace-cross Abstract: Multi-agent LLM pipelines produce contradictory evidence on whether team diversity improves output quality: heterogeneous Mixture-of-Agents teams outperform single models, yet homogeneous Self-MoA teams consistently win un…

  194. arXiv cs.LG TIER_1 English(EN) · Thomas Carta, Cl\'ement Romac, Loris Gaven, Pierre-Yves Oudeyer, Olivier Sigaud, Sylvain Lamprier ·

    HERAKLES:面向开放式LLM代理的层级技能编译

    arXiv:2508.14751v2 Announce Type: replace Abstract: We study goal-conditioned reinforcement learning in partially observable environments with sparse rewards and large, structured goal spaces. In such settings, complex goals often require composing simpler skills, but learning th…

  195. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Daisuke Kikuta ·

    AI旅游会议:LLM代理的团队旅行规划

    This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Language Model (LLM)-based agents. The agents are instantiated with distinct personas and collaboratively seek an itinerary that satisfies their constraints and preferences through na…

  196. arXiv cs.AI TIER_1 English(EN) · Jing-Jing Li, Jianfeng He, Chao Shang, Devang Kulshreshtha, Xun Xian, Yi Zhang, Hang Su, Sandesh Swamy, Yanjun Qi ·

    STAC:无辜的工具如何为LLM代理形成危险的链条

    arXiv:2509.25624v3 Announce Type: replace-cross Abstract: As LLMs advance into autonomous agents with tool-use capabilities, they introduce security challenges that extend beyond traditional content-based LLM safety concerns. This paper introduces Sequential Tool Attack Chaining …

  197. arXiv cs.LG TIER_1 English(EN) · YanZhao Zheng, ZhenTao Zhang, Chao Ma, YuanQiang Yu, JiHuai Zhu, Yong Wu, Tianze Xu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu ·

    SkillRouter: 规模化 LLM Agent 的技能路由

    arXiv:2603.22455v5 Announce Type: replace Abstract: Reusable skills let LLM agents package task-specific procedures, tool affordances, and execution guidance into modular building blocks. As skill ecosystems grow to tens of thousands of entries, exposing every skill at inference …

  198. arXiv cs.AI TIER_1 English(EN) · Shijun Li, Hilaf Hasson, Joydeep Ghosh ·

    OMAC:LLM驱动的多智能体协作的整体优化框架

    arXiv:2505.11765v5 Announce Type: replace-cross Abstract: Agents powered by advanced large language models (LLMs) have demonstrated impressive capabilities across diverse complex applications. Recently, Multi-Agent Systems (MAS), wherein multiple agents collaborate and communicat…

  199. arXiv cs.AI TIER_1 English(EN) · Runming He, Zhen Hao Wong, Hao Liang, Zimo Meng, Chengyu Shen, Xiaochen Ma, Wentao Zhang ·

    DataFlow-Harness: 用于构建可编辑 LLM 数据管道的、基于事实的代码代理平台

    arXiv:2607.16617v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. We call this…

  200. arXiv cs.AI TIER_1 English(EN) · Roshan Klein-Seetharaman, Daniel Wang, Andrew Xu ·

    Lomekwi:LLM智能体中的资源受限工具发现

    arXiv:2607.16961v1 Announce Type: new Abstract: Existing tool-use benchmarks report a single success rate for complex, multistep tasks. Inspired by ideas from cognitive science, we distinguish tool use from tool discovery and decompose the latter into curiosity (the model's abili…

  201. arXiv cs.AI TIER_1 English(EN) · Sumit Verma, Pritam Prasun, Pritish Kumar ·

    RAIL Guard:弥合LLM智能体负责任AI的评估到补救差距

    arXiv:2607.16215v1 Announce Type: new Abstract: Existing guardrail systems for large language model agents operate as binary classifiers that block unsafe content, leaving organizations to discard failing outputs and retry from scratch. We introduce RAIL Guard, a closed-loop resp…

  202. Hugging Face Daily Papers TIER_1 English(EN) ·

    NexForge:通过面向需求的任务合成扩展 LLM 的代理能力

    Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual substrate engineering, each new domain demands a bespoke pipelin…

  203. Hugging Face Daily Papers TIER_1 English(EN) ·

    验证、修复、重复,还是停止?LLM代理中嘈杂的验证-修复循环的鲁棒停止策略

    Verify-repair loops are a standard means for large language model (LLM) agents to correct faulty plans in code generation, mathematical reasoning, and tool use. When both the verifier and the repairer are noisy, repair can damage already-correct plans, and reported acceptance kee…

  204. arXiv cs.AI TIER_1 English(EN) · Mingxin Li, Enge Song, Yueshang Zuo, Xiaodong Liu, Rong Wen, Qiang Fu, Gianni Antichi, Jian He, Jing Tie, Zhou Shao, Xiaobo Xue, Xiong Xiao, Luyao Zhong, Shaokai Zhang, Jiangu Zhao, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Changgang Zheng, Zihao Fan, Haon… ·

    云端可扩展 LLM 代理工具访问

    arXiv:2607.15593v1 Announce Type: cross Abstract: LLM agents increasingly rely on tool calling to act on external systems, and the Model Context Protocol (MCP) has quickly become its de facto interface. Operating MCP at cloud scale, however, becomes difficult. On the tool provide…

  205. arXiv cs.CL TIER_1 English(EN) · Yanze Wang, Pengfei Yao, Tianyi Sun, Chuanrui Hu, Yan Xiao, Yunyun Han, Jun Sun, Yafeng Deng ·

    SkillCorpus:整合和评估现实世界 LLM Agent 的开放技能生态系统

    arXiv:2607.15557v1 Announce Type: new Abstract: Agent skills, SKILL.md files that package reusable procedural knowledge for an LLM agent, are a popular mechanism for extending agent capabilities. Public repositories now host them in large and growing numbers, yet these artifacts …

  206. Hugging Face Daily Papers TIER_1 English(EN) ·

    穷人的代理建模:在笔记本电脑上模拟大型LLM-代理社会

    Replacing individual LLM agents with low-parameter surrogates fitted from cheap queries enables scalable society simulations, with validity predicted by an interaction-order and memory taxonomy.

  207. Hugging Face Daily Papers TIER_1 English(EN) ·

    DataFlow-Harness: 用于构建可编辑 LLM 数据管道的接地式代码代理平台

    Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. We call this disconnect the NL2Pipeline gap. To bridge it, we …

  208. arXiv cs.AI TIER_1 English(EN) · Jason Miklian ·

    人工智能LLM引擎如何塑造全球冲突信息环境

    arXiv:2607.14197v1 Announce Type: new Abstract: Artificial Intelligence (AI) answer engines now field a growing share of the questions that analysts, scholars, and the public ask about issues of peace and conflict. Large Language Models (LLMs) are known to hallucinate under certa…

  209. arXiv cs.AI TIER_1 English(EN) · Chuhan Shi, Xiaoquan Ren, Sicheng Song, Haobo Li, Rui Sheng, Yushi Sun ·

    大型语言模型已准备好进行科学发现吗?面向AI科学家的能力导向基准测试

    arXiv:2607.11079v1 Announce Type: new Abstract: Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, s…

  210. Hugging Face Daily Papers TIER_1 English(EN) ·

    大型语言模型已准备好进行科学发现吗?面向人工智能科学家的能力导向基准测试

    Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, statistical inference, mechanistic explanation, e…

  211. arXiv stat.ML TIER_1 English(EN) · Tianbing Xu ·

    从期望最大化视角看大语言模型推理的强化学习

    arXiv:2504.18587v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrated by systems such as OpenAI's O1~\cite{o1} and DeepSeek-R1~\cite{r1}. However, wide…

  212. arXiv stat.ML TIER_1 English(EN) · Amirmohammad Farzaneh, Osvaldo Simeone ·

    短思考、巧推迟、执行、重复:边缘 LLM 代理的校准推理与不确定性感知推迟

    arXiv:2607.26865v1 Announce Type: new Abstract: LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. Yet, when deployed at the edge, they must tight…

  213. Hacker News — AI stories ≥50 points TIER_1 English(EN) · rellem ·

    Mozilla:开源AI的现状

  214. Forbes — Innovation TIER_1 English(EN) · Mohit Bhat, Forbes Councils Member ·

    微调SLM:企业AI的新运营模式

    Fine-tuning is transforming SLMs from efficient components into high-performance, enterprise-grade systems.

  215. MarkTechPost TIER_1 English(EN) · Michal Sutter ·

    10个开源无代码AI平台,用于构建LLM应用、RAG系统和AI代理

    <p>Retrieval, agents, and workflows now ship as visual and plain-English tools. This roundup covers 10 open-source no-code and low-code platforms for building LLM apps, RAG systems, and AI agents, each with its verified license, repository, and best-fit use case.</p> <p>The post …

  216. dev.to — MCP tag TIER_1 English(EN) · Diego Costa ·

    使用原生MCP工具解决AI销售代理中的LLM参数幻觉问题

    <h1> Solving LLM Parameter Hallucinations in AI Sales Agents with Native MCP Tools </h1> <p>The most efficient way to eliminate LLM parameter hallucinations when retrieving B2B firmographics is by leveraging a native Model Context Protocol (MCP) server with strict Zod-enforced sc…

  217. dev.to — MCP tag TIER_1 English(EN) · Victor García ·

    Agent Gateway 60秒:使用TrustGate管理LLM流量

    <h1> Agent Gateway in 60 Seconds: Governed LLM Traffic with TrustGate </h1> <p>Most teams start with a direct OpenAI (or Anthropic) SDK call. That works until you have three apps, two providers, and a security review asking who can call which model, at what rate, with what audit …

  218. Towards AI TIER_1 English(EN) · Diogo Santos ·

    IntentFlow:具有可审计、哈希链式追踪的可控 LLM 代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/intentflow-governed-llm-agents-with-auditable-hash-chained-traces-49599f09e590?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1000/0*1jmaHvD-pw4NkgOa.png" …

  219. Towards AI TIER_1 English(EN) · MongoDB ·

    为多代理系统添加成本计量和 LLM 支出可见性

    <p><em>Written by </em><a href="https://www.linkedin.com/in/matteo-rossi-280391/"><em>Matteo Rossi.</em></a></p><p>The monthly LLM bill jumped, and nobody on the team can say which agent, which user, or which workflow caused it. The provider dashboard breaks usage down by organiz…

  220. Medium — fine-tuning tag TIER_1 English(EN) · Mikhail Borodastov ·

    Harness-native agents:与Harness共同训练LLM,以在单一任务中最大化质量

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://mlboroda.medium.com/harness-native-agents-co-train-the-llm-with-its-harness-to-max-out-quality-inside-one-task-93c321a93a81?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/250…

  221. dev.to — MCP tag TIER_1 English(EN) · Programming Central ·

    破解像素代码:视觉驱动的Agent如何将LLM的思考转化为DOM点击

    <p>The bleeding edge of AI automation isn't just about making Large Language Models (LLMs) smarter; it's about giving them hands and eyes. When building vision-driven agentic architectures, we cross a massive chasm: bridging the high-level semantic reasoning of an LLM with the lo…

  222. dev.to — MCP tag TIER_1 English(EN) · Diego Costa ·

    使用 Lead Enrichment MCP API 消除 B2B 销售代理中的 LLM 幻觉

    <h1> Eliminating LLM Hallucinations in B2B Sales Agents with the Lead Enrichment MCP API </h1> <p>To stop LLMs from hallucinating company data or fabricating contact details, developers must shift from loose prompt-based retrieval to a Model Context Protocol (MCP) architecture th…

  223. Medium — Claude tag TIER_1 English(EN) · Ashishmohanka ·

    大型语言模型(LLM)是如何工作的——构建 AI 代理的实用指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ashishmohanka123/how-llms-actually-work-a-practical-guide-for-building-ai-agents-53139a3b6665?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1024/1*thBASTiA9rwrB2FAwG8…

  224. dev.to — MCP tag TIER_1 English(EN) · Diego Costa ·

    通过原生MCP B2B增强功能消除销售代理中的LLM参数幻觉

    <h1> Eliminating LLM Parameter Hallucinations in Sales Agents with Native MCP B2B Enrichment </h1> <p>The most efficient way to stop LLMs from hallucinating firmographic data or misinterpreting complex API schemas is to deploy a Model Context Protocol (MCP) native B2B lead enrich…

  225. Medium — MLOps tag TIER_1 English(EN) · Tedi Ikonomi ·

    OpenShift AI Air-Gapped:为分布式 LLM 推理准备平台

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ikonomi.tedi/openshift-ai-air-gapped-preparing-the-platform-for-distributed-llm-inference-98a273f7bfdc?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/2600/1*RQnoWiHB6Pa…

  226. Medium — Claude tag TIER_1 English(EN) · Neo Malesa ·

    从ChatGPT聊天机器人到图谱:我们处理LLM方式的快速演变

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@neomalesa/from-chatgpt-chatbots-to-graphs-the-rapid-evolution-of-how-we-work-with-llms-c29894e87718?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1168/1*F2AM8F69iqJw_…

  227. Medium — MLOps tag TIER_1 English(EN) · Rami Krispin ·

    skforecast-ai 项目:面向生产系统的实用 LLM 评估 | 第 97 期

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@rami.krispin/the-skforecast-ai-project-practical-llm-evaluation-for-production-systems-issue-97-3c0b19ed14aa?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1920/1*tNzi0…

  228. dev.to — MCP tag TIER_1 English(EN) · PromptOT ·

    PromptOT MCP:管理和版本化您AI工具中的LLM提示

    <h1> PromptOT MCP: Manage and version LLM prompts from your AI tools </h1> <p>Prompts often start as simple strings in code.</p> <p>Then the product grows.</p> <p>You add a better system prompt. Then a guardrail. Then a different version for production. Then a customer-specific v…

  229. Medium — MLOps tag TIER_1 English(EN) · Neelopphersyed ·

    NeuralUCB Router:一个兼容OpenAI的API代理,使用多臂...

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@neelopphersyed7/neuralucb-router-an-openai-compatible-api-proxy-that-routes-llm-requests-using-a-multi-armed-17e762724926?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max…

  230. Medium — MCP tag TIER_1 English(EN) · EuroAmerican Institute ·

    LLM vs RAG vs MCP:AI工程师和开发者的游戏规则改变者

    <div class="medium-feed-item"><p class="medium-feed-snippet">The Model Context Protocol hit 97 million monthly SDK downloads by December 2025. That number alone tells you something important is&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@euroamericanmalta/…

  231. Medium — Claude tag TIER_1 English(EN) · Shankar ·

    每月20美元的人工智能失误:大型语言模型如何使AWS架构复杂化

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@shankar_somasundaram/the-20-month-ai-mistake-how-llms-overcomplicate-aws-architecture-de8980849607?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1024/1*CZ7mUHQAGZQdxg…

  232. Towards AI TIER_1 English(EN) · Vasilii Chetvertukhin ·

    迈向自托管企业人工智能的四层架构

    <h4>There is no shortage of articles about building AI agents. What remains much rarer is a practical discussion of how to run them safely in production.</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*tnltjwGYfOIZX6KgpLNUEg.png" /></figure><p>This article…

  233. Medium — MLOps tag TIER_1 English(EN) · sentraorb ·

    统一所有 LLM 的一个入口:使用 LiteLLM 构建生产就绪的 AI 堆栈

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://sentraorb.medium.com/one-gateway-to-rule-all-your-llms-building-a-production-ready-ai-stack-with-litellm-1ffcb29a7733?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1280/1*zFxMRtOq…

  234. Medium — MLOps tag TIER_1 English(EN) · sentraorb ·

    统一所有大语言模型的入口:使用LiteLLM构建生产级AI堆栈

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://aws.plainenglish.io/one-gateway-to-rule-all-your-llms-building-a-production-ready-ai-stack-with-litellm-1ffcb29a7733?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1280/1*zFxMRtOqO…

  235. Medium — Anthropic tag TIER_1 Español(ES) · LinaUX Off Frame ·

    人工智能的能力与局限性:教我诊断LLM错误的课程

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@l.godefroy.design/ai-capabilities-and-limitations-el-curso-que-me-ense%C3%B1a-a-diagnosticar-los-errores-de-las-llm-c2047c580120?source=rss------anthropic-5"><img src="https://cdn-images-1.med…

  236. Medium — fine-tuning tag TIER_1 English(EN) · Tech Horizon With Anand Vemula ·

    微调LLM:开发者定制AI模型的指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@anandvlinkedin/fine-tuning-llms-a-developers-guide-to-custom-ai-models-2e7b5e7989aa?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1536/1*nmFyPKH5QY0oRBC5XJIJNw.p…

  237. dev.to — LLM tag TIER_1 English(EN) · Umair Bilal ·

    我的AI智能体2x2模型成本效益策略

    <blockquote> <p><em>This article was originally published on <a href="https://www.buildzn.com/blog/my-2x2-llm-cost-performance-strategy-for-ai-agents" rel="noopener noreferrer">BuildZn</a>.</em></p> </blockquote> <p>Everyone's chasing the biggest LLMs, throwing cash at Claude or …

  238. dev.to — LLM tag TIER_1 English(EN) · Alex ·

    Pydantic AI 评测:为 Python 提供类型化代理,该框架使 LLM 输出更可靠

    <p>Pydantic AI is the official agent framework from the Pydantic team, built around typed, validated LLM output. After 45 days of using it for saas.pet's content QA agent and data extraction scripts, here is the real story on structured output, tool calling, and why it beats Lang…

  239. Mastodon — fosstodon.org TIER_1 Русский(RU) · [email protected] ·

    TrueForge:一个将大型语言模型转变为全功能代理的开源框架。今天即可轻松构建演示代理。连接模型,添加几个工具

    TrueForge: открытая обвязка, которая превращает LLM в полноценного агента Собрать демо-агента сегодня несложно. Подключаешь модель, добавляешь пару инструментов — и она уже читает файлы, вызывает API и бодро обещает выполнить любую задачу. Сложности начинаются, когда такого агент…

  240. dev.to — LLM tag TIER_1 ไทย(TH) · Nokka ·

    AI代理的LLM路由,为什么代理需要选择模型以及如何将成本降低高达80%

    <h1> LLM Routing สำหรับ AI Agent, ทำไม agent ถึงต้องเลือกโมเดลเป็น และวิธีลดต้นทุนได้ถึง 80% </h1> <p><em>โดย Nokka (นก-กา) | 21 สิงหาคม 2026</em></p> <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity…

  241. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    超越#LLM:使用LangChain DeepAgents创建真正的#AI代理,了解LangChain DeepAgents如何将LLM转变为生产就绪的AI系统

    Beyond # LLMs : Creating Real-World # AI Agents with Lang Chain Deep Agents Discover how LangChain DeepAgents transform LLMs into production-ready AI systems with memory, skills, sub-agents, context management, and human oversight. https:// hackernoon.com/beyond-llms-cre ating-re…

  242. Mastodon — fosstodon.org TIER_1 Français(FR) · [email protected] ·

    SkillOpt (Microsoft): LLM 智能体的自然语言技能优化器。该技能通过评分推广进行改进,无需触碰模型权重。

    SkillOpt (Microsoft) : un optimizer de skills en langage naturel pour agents LLM. Le skill s'améliore via des rollouts scorés, sans toucher aux poids du modèle. Le fichier best_skill.md est portable d'un modèle à l'autre. Open source, MIT. ⬇️ https:// github.com/microsoft/SkillOp…

  243. dev.to — LLM tag TIER_1 English(EN) · Felipe L ·

    Zero-Mem:LLM智能体零Token记忆操作

    <h2> What Happened </h2> <p>Zero‑Mem lets LLM agents read and write external memory without generating or consuming any tokens. Traditional agents fetch context through token‑based prompts, adding latency and cost. Zero‑Mem replaces that with a lightweight, token‑free interface t…

  244. dev.to — LLM tag TIER_1 English(EN) · Ming ·

    在边缘运行LLM代理:使用NeoMind + Ollama的实用指南

    <h1> Running LLM Agents at the Edge: A Practical Guide with NeoMind + Ollama </h1> <p>Everyone's building AI agents right now. Most of them live in the cloud — you send a request to OpenAI or Anthropic, get a response back, and hope the latency and cost stay reasonable. But what …

  245. dev.to — LLM tag TIER_1 English(EN) · talor ·

    利用实时搜索进行AI代理构建:SERP API如何实现可靠的LLM应用

    <p>Large language models have changed how developers build applications.</p> <p>However, even the most advanced LLMs have one fundamental limitation:</p> <p>They do not have access to real-time information.</p> <p>A model may understand programming, reasoning, and language extrem…

  246. dev.to — LLM tag TIER_1 English(EN) · Yogi ·

    构建我自己的LLM模型和代理

    <h2> Introduction </h2> <p>Large Language Models (LLMs) are powerful, but most enterprises rely on pre‑packaged APIs. I wanted to go deeper: train my own LLM model and build an agent layer on top of it that could interact with real systems securely.</p> <p>This post walks through…

  247. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    LLM代理:纯代码验证将目标放弃率从100%降至0% 新arXiv预印本:确定性执行器拥有所有代理信念,LLM仅负责归档

    LLM agent: code-only verification flips goal abandonment 100% to 0% New arXiv preprint: a deterministic executive owns all agent belief, the LLM only files proposals, and zero ARC-AGI-3 completions are honestly disclosed. https://www. notatechguy.com/llm-agent-code -only-verifica…

  248. dev.to — LLM tag TIER_1 English(EN) · kai wen ng ·

    迈向LLM智能体稳定性

    <p>An industry-level LLM agent is not simply an API call that returns a response. It needs to be resilient to transient failures, malformed outputs, and schema violations.<br /> To improve the stability of my agent system, I introduced two decorators around my LLM calls. They han…

  249. dev.to — LLM tag TIER_1 English(EN) · Lorena Dávila Ermus ·

    1. 自托管AI:高效运行模型所需的LLM概念

    <p>If you want to run AI models on your own machine and learn the basic concepts with me to do it effectively, then this is the right article :).</p> <p>This is part one of the series. In the next one we build local AI workflows with n8n and Ollama. This article is the vocabulary…

  250. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📄 ‘MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations’ 在 Hugging Face 上获得 80 个赞。测试 LLMs 在持续任务链上的表现

    📄 ‘MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations’ hit 80 upvotes on Hugging Face. Tests LLMs on sustained task chains in shopping. https:// huggingface.co/papers/2607.289 56 # AI # MachineLearning # Research

  251. dev.to — LLM tag TIER_1 English(EN) · Mohammad Jawad (Kasir) Barati ·

    理解LLM智能体和工具

    <p>In this post, I'll explain how prompt chaining, tools/skills, and iteration actually make agnets to produce results which are not usually possible when we use simple LLMs.</p> <h2> LLM chaining </h2> <p>You know how sometimes you write one massive prompt like:</p> <blockquote>…

  252. dev.to — LLM tag TIER_1 English(EN) · Dmitriy ·

    我们如何为 AI 代理选择 LLM 和框架

    <p>Over the last 18 months our ML team has been doing some very interesting things: building AI agents on top of PostgreSQL, while the infrastructure evolves, the industry matures, and quality expectations keep rising. We started with a single A100 in a managed cloud and fairly m…

  253. dev.to — LLM tag TIER_1 English(EN) · Wibo ·

    如何评估 LLM 代理:evals、黄金数据集和 LLM 作为裁判

    <p><strong>Short answer</strong></p> <p><strong>You can't unit-test an LLM to correctness, because the same input can take a different path on the next run.</strong> Evals are the test suite for probabilistic systems: a scored, repeatable check of whether the system reached an ac…

  254. dev.to — LLM tag TIER_1 English(EN) · Hiroshi Toyama ·

    面向 LLM Agent 的分层评估策略(Google ADK 的 12 条标准)

    <p>Google's <a href="https://adk.dev/evaluate/criteria/" rel="noopener noreferrer">Agent Development Kit (ADK)</a> ships 12 evaluation criteria for testing agent behavior: tool-call trajectories, final response quality, hallucination detection, safety, multi-turn task success, an…

  255. dev.to — LLM tag TIER_1 Deutsch(DE) · Tsari Bombelli ·

    llms.txt 详解:AI 爬虫和 LLMs 的标准

    <p>llms.txt ist ein maschinenlesbarer Standard, der KI-Systemen strukturierte Informationen über Ihre Website bereitstellt. Aufbau, Best Practices und praktische Implementierung für bessere KI-Sichtbarkeit bei ChatGPT, Claude, Gemini und Perplexity.</p> <h3> Zusammenfassung </h3>…

  256. dev.to — LLM tag TIER_1 English(EN) · Jules Robineau ·

    重建以理解:从网络协议到LLM代理

    <blockquote> <p><strong>TL;DR</strong>: you only truly understand a system once you rebuild it. I recoded TCP at school, then the DNS protocol, then Modbus, each time to understand it from the inside. A colleague just went through this with LLMs. He wrote a small agent in Go, and…

  257. dev.to — LLM tag TIER_1 English(EN) · Yusuf Al-Rashidi ·

    9 款最佳 LLM 网关,适用于 Agentic 工作流和 AI 代理

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fairuyu0zpj9tz04zmsbj.png"><img alt="9 Best LLM Gatew…

  258. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    刚刚为AI代理样板文件添加了本地LLM支持,现在可以在Ollama之上运行,而无需依赖云API 👇 https://github.com/christophed

    Just added local LLM support to the AI agent boilerplate, you can now run it on top of Ollama instead of relying on cloud APIs 👇 https:// github.com/christopheduc-me/ai -agent-boilerplate # buildinpublic # ai # dev # tech

  259. dev.to — LLM tag TIER_1 English(EN) · TheKitBase ·

    2026年如何削减您的AI/LLM成本:缓存、更便宜的模型和多代理路由

    <p>AI features ship fast and then the bill arrives. The good news: most LLM spend is avoidable waste - the same prompt paid for a thousand times, a frontier model doing work a cheap one could handle, tokens generated that nobody reads. Here are six levers that cut real money, ord…

  260. dev.to — LLM tag TIER_1 English(EN) · Apache SeaTunnel ·

    AI能否真正构建数据管道?使用Apache SeaTunnel AI CLI对7个领先LLM进行100项任务的基准测试

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbcakf9op8ltnlaaujm5z.jpg"><img height="439" src="htt…

  261. dev.to — LLM tag TIER_1 Türkçe(TR) · Emre Yıldız ·

    在您自己的服务器上运行LLM:10分钟用Ollama实现本地AI

    <p>OpenAI API'ına token başına para ödemek yerine, açık kaynak dil modellerini (Llama 3, DeepSeek, Mistral, Qwen) kendi sunucunda çalıştırabilirsin. Verin dışarı çıkmaz, sabit maliyet, sınırsız istek. Bu yazıda Ollama ile pratik kurulumu ve gereken donanımı anlatıyorum.</p> <h2> …

  262. dev.to — LLM tag TIER_1 English(EN) · Learn AI Resource ·

    在本地运行开源大模型:您的AI编程助手,无需API费用

    <p>So you want an AI coding assistant but you're tired of getting dinged for API calls every time you ask for help debugging a regex? Yeah, I get it.</p> <p>Here's the thing: you don't actually need to pay OpenAI or Anthropic to get decent AI pair programming. You can run a solid…

  263. dev.to — LLM tag TIER_1 English(EN) · Sofia Aliferi ·

    超越审核:为何大型语言模型系统需要策略层

    <blockquote> <p>TL;DR: Moderation catches harm and many injection attempts. It does not enforce domain or operational policy. A policy reasoning layer (LLM-as-a-judge) closes that gap, especially in multi-turn conversations.</p> </blockquote> <p><strong>Abstract</strong><br /> Mo…

  264. dev.to — LLM tag TIER_1 English(EN) · soy ·

    本地大模型、开放代理和自托管部署平台趋势兴起

    <h2> Local LLMs, Open Agents &amp; Self-Hosted Deployment Platforms Trending </h2> <h3> Today's Highlights </h3> <p>Today's top stories highlight the growing trend of local and self-hosted AI deployments, featuring an architectural guide for secure "Local Sovereign LLMs" in enter…

  265. dev.to — LLM tag TIER_1 English(EN) · Kuldeep Paul ·

    按团队跟踪 LLM 使用情况和支出:AI 治理指南

    <p><em>Organizations deploying AI applications face challenges in accurately tracking LLM usage and spend across different teams and projects. <a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer">Bifrost</a> offers a comprehensive AI gateway solution with virtual k…

  266. dev.to — LLM tag TIER_1 English(EN) · Stéphane Derosiaux ·

    chrome-agent:将任何大型语言模型变成智能网页浏览代理

    <p>Ever handed an LLM a full web page and watched the amount of tokens being used?</p> <p>A single product listing is 20-30K tokens of </p> soup before the model finds what it needs: wrapper divs, css class, SVG, JSON blobs etc. The agent needs maybe 300 tokens of that (the items…

  267. dev.to — LLM tag TIER_1 English(EN) · DryDock ·

    LLM 代理操作 CAD 内核的七个实际失败(及其架构如何控制了它们)

    <p>I built a system where an LLM talks to a customer about a silicone casting mold, and a<br /> deterministic geometry kernel — OpenCASCADE, three decades of production C++ — does the<br /> actual mass-solving, boolean surgery, and part-splitting. The LLM never touches the kernel…

  268. dev.to — LLM tag TIER_1 English(EN) · soy ·

    LLM推理与RAG优化,用于本地部署的开源语音AI

    <h2> LLM Inference &amp; RAG Optimization, Open-Source Voice AI for Local Deployments </h2> <h3> Today's Highlights </h3> <p>This week's highlights feature a new framework for LLM inference and fine-tune optimizations, including KV-cache improvements, alongside an open-source voi…

  269. dev.to — LLM tag TIER_1 English(EN) · Kuldeep Singh Arya ·

    生产环境中 8 种 LLM 成本优化技术

    <p>Executive Summary<br /> As generative AI transitions from experimental prototypes to high-scale production systems, the primary bottleneck for engineering teams has shifted from model capability to unit economics. The pricing structure of modern Large Language Model (LLM) APIs…

  270. dev.to — LLM tag TIER_1 English(EN) · Shouvik Palit ·

    Sir Shortoken:为每个LLM提供严谨的AI输出

    <p><strong>TL;DR:</strong> Sir Shortoken is a system prompt that constrains frontier models to operate within information budgets (Quick/Balanced/Deep), never silently escalate capabilities, and prove execution. Tested across Claude, GPT, Gemini. 40-60% token reduction on technic…

  271. dev.to — LLM tag TIER_1 English(EN) · Praveen Maurya ·

    使用本地 LLM 构建:工程师的 AI 辅助开发方法

    <blockquote> <p>I didn't build SafeDevTools by asking AI to "build me a website." I built it by treating a local LLM like a junior engineer who never gets tired of writing boilerplate.</p> </blockquote> <p>A few weeks ago, I challenged myself with a simple experiment:<br /> <stro…

  272. dev.to — LLM tag TIER_1 Deutsch(DE) · Uhltak Therestismysecret ·

    使用 Ollama 部署本地大模型:托管模型、集成 API 并高效使用

    <h1> Lokale LLMs mit Ollama – Modelle selbst hosten und per API anbinden </h1> <p><strong>Hook:</strong> Stell dir vor, du könntest ChatGPT für deine Firma betreiben, ohne einen teuren Cloud‑Vertrag oder ein Datenleck‑Szenario. Du hast die volle Kontrolle, die Kosten liegen bei d…

  273. dev.to — LLM tag TIER_1 Русский(RU) · Promptra Team ·

    本地AI模型:Exo Labs的local.ai Tracker与Mac Mini集群

    <p>Если ты открыл эту статью с вопросом «где посмотреть, какая модель реально влезет в мой Mac и не будет тормозить», то короткий ответ такой: 2 июля 2026 года Exo Labs на конференции AI Engineer World's Fair анонсировала сервис local.ai, который отслеживает, какая модель лучше в…

  274. dev.to — LLM tag TIER_1 English(EN) · shayesta ·

    LangChain4j 和 Spring AI:让您的 Java 应用与 LLM 对话的“管道”

    <p>If you've heard about LangChain and assumed it was a Python thing, that's fair. It mostly was.</p> <p>LangChain became popular because building with an LLM turns out to involve a lot of repetitive plumbing. You need to manage conversation history, split documents into chunks, …

  275. dev.to — LLM tag TIER_1 English(EN) · Bahadir Kusat ·

    人工智能模型如何训练?LLM训练指南

    <p>From data preparation and tokenizer selection to pretraining, LoRA, RLHF, evaluation, and production monitoring, this guide covers the major stages involved in training an AI model.</p> <p>DEHA Research · July 14, 2026 · 18 min read</p> <p>Training an artificial intelligence m…

  276. dev.to — LLM tag TIER_1 English(EN) · soy ·

    浏览器 LLM 代理、Apple Silicon 的 Rust 引擎,以及本地 AI 代码解释器

    <h2> Browser LLM Agents, Rust Engine for Apple Silicon, &amp; Local AI Code Interpreter </h2> <h3> Today's Highlights </h3> <p>This week, we spotlight tools bringing LLM inference directly to your devices. Dive into browser-based agents, a Rust-native engine for Apple Silicon, an…

  277. dev.to — LLM tag TIER_1 English(EN) · Innocent Oyebode ·

    我如何为尼日利亚中小企业构建多页面AI网站生成器——架构、LLM提示和经验教训

    <h2> The Problem </h2> <p>Most Nigerian small businesses have no web presence at all. When they do get a website, it is usually a stale brochure-ware page that took a freelancer three weeks to deliver and costs ₦150,000 they could not really afford. The freelancer is long gone; t…

  278. dev.to — LLM tag TIER_1 English(EN) · Jack M ·

    LLM 延迟预算:让 AI 工作流感觉快速,无需猜测

    <p>A slow AI feature rarely fails all at once. It starts with a longer prompt, then a bigger retrieval result, then one more tool call, then a retry path nobody measured. The demo still works, but users feel the delay before your dashboard explains it.</p> <p>That is why small AI…

  279. dev.to — LLM tag TIER_1 English(EN) · soy ·

    自托管AI伴侣与开源模型API洞察

    <h2> Self-Hosted AI Companion &amp; Open-Source Model API Insights </h2> <h3> Today's Highlights </h3> <p>This week's highlights feature a trending self-hosted AI companion, empowering users with personal, locally-run AI experiences. We also explore a bootcamp grad's practical in…

  280. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    从演示到可靠系统:让大型语言模型真正可投入生产的人工智能工程技术

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/from-demos-to-durable-systems-ai-engineering-techniques-that-make-llms-truly-product-ready?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noreferrer">…

  281. dev.to — LLM tag TIER_1 English(EN) · soy ·

    自托管大模型应用、离线AI系统及本地自动化基础

    <h2> Self-Hosted LLM Apps, Offline AI Systems, and Local Automation Foundations </h2> <h3> Today's Highlights </h3> <p>This week, we spotlight practical approaches to self-hosting AI, from extensive curated lists of runnable LLM applications to ambitious projects building fully o…

  282. dev.to — LLM tag TIER_1 English(EN) · bossandboss ·

    构建 EdgeSync-LLM:去中心化、离线优先的本地 AI 的最终架构 🚀

    <h1> Published: true </h1> <h1> Description: A deep dive into the final version of EdgeSync-LLM—bringing fast, secure, synchronized Large Language Models straight to edge hardware. </h1> <h1> Tags: ai, open source, architecture, edgecomputing, webdev </h1> <p>The cloud dependency…

  283. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    对于科技界的任何人来说,探索 Llama、Mistral 和 Phi 等开源 AI 模型都是必不可少的!这些模型通过推广正在改变 AI 的格局

    Exploring open-source AI models like Llama, Mistral, and Phi is a must for anyone in the tech world! These models are changing the landscape of AI by promoting collaboration and innovation. Dive into the world of deep learning! 🤖 # AI # AITürkiye # DeepLearning # Teknoloji # Mach…

  284. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    GPT-5.6 现身:OpenAI 的新模型和定制芯片将如何重塑生产型 LLM 系统

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/gpt-5-6-in-the-wild-how-openai-s-new-model-and-custom-silicon-will-reshape-production-llm-systems?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noref…

  285. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    GPT-5.6、Jalapeño 以及下一代 OpenAI 优化的大型语言模型基础设施

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/gpt-5-6-jalapeno-and-the-next-generation-of-openai-optimized-llm-infrastructure?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noreferrer">CoreProse K…

  286. dev.to — LLM tag TIER_1 English(EN) · Abdul Rehman ·

    构建生产级AI流水线:使用LLM每日处理10,000+个列表

    <p>I learned the hard way that a working LLM pipeline and a production LLM pipeline are two different things.</p> <p>When I first built the scoring system for a job board platform, I thought: throw GPT-4 at each listing, ask it to rate relevance, done. It worked for 100 listings.…

  287. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    深入了解 GPT-5.6:OpenAI 的新旗舰模型和定制芯片将如何重塑 LLM 运营

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/inside-gpt-5-6-how-openai-s-new-flagship-model-and-custom-silicon-will-reshape-llm-operations?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noreferre…

  288. dev.to — LLM tag TIER_1 English(EN) · Remy Okafor ·

    面向大模型团队的10款开源AI基础设施工具

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr8ju6nt2a4ngc3j9sww5.png"><img alt="10 Open-Source A…

  289. dev.to — LLM tag TIER_1 English(EN) · Caleb Osei ·

    流式传输 LLM 响应的最佳 AI 网关

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwk9fm8436j9drcevri9t.png"><img alt="Best AI Gateways…

  290. dev.to — LLM tag TIER_1 English(EN) · Kuldeep Paul ·

    AI网关如何提高LLM的可靠性:9种方法

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frp4yqktt436b7wu1fox2.png"><img alt="9 Ways an AI Gat…

  291. dev.to — LLM tag TIER_1 English(EN) · Abdul Rehman ·

    如何为您的AI MVP构建可靠的LLM管道,避免过度设计

    <p>I once built an AI pipeline that was shut down after a single month. The LLM costs were unsustainable, and worse, the outputs were unreliable enough that we couldn't trust them in production. That failure taught me something I still use today: evaluation isn't a phase you add …