PulseAugur
实时 20:41:20
English(EN) Computer-Using Agent

OpenAI 在科学和安全领域推进 AI 代理;Google DeepMind 资助多智能体研究

OpenAI 正通过多项举措推进科学计算和 AI 安全。该公司发布了一个新的基准 GeneBench-Pro,以评估 AI 代理处理复杂生物数据的能力。OpenAI 还在欧洲为可信赖 AI 的共享标准制定做出贡献,并与伦敦证券交易所集团等组织合作,将 AI 整合到业务运营中。与此同时,Google DeepMind 正在向多智能体 AI 安全研究投资 1000 万美元,旨在理解和减轻与交互式 AI 系统相关的风险。 AI

影响 专注于用于科学发现的 AI 代理和多智能体安全研究,突显了未来 AI 发展和风险缓解的关键领域。

排序理由 详细介绍了与 AI 代理和安全相关的多项研究计划和基准。

在 OpenAI News 阅读 →

AI 生成摘要 · Google Gemini · 来自 2810 个来源。 我们如何撰写摘要 →

OpenAI 在科学和安全领域推进 AI 代理;Google DeepMind 资助多智能体研究

报道来源 [2810]

  1. X — Meta AI TIER_1 English(EN) · AIatMeta ·

    推出 Muse Glimmer,一款专为本地、始终在线的代理工作流程优化的开源 30B 参数模型。

    Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows. Muse Glimmer delivers strong performance on key agentic use cases and benchmarks compared with leading models in its size category, and is designed to run entirely on htt…

  2. OpenAI News TIER_1 English(EN) ·

    Agentic AI 时代的科学计算

    A new field report shows how scientists use AI coding agents to modernize scientific computing, accelerating software development and discovery in genomics and beyond.

  3. X — OpenAI TIER_1 English(EN) · OpenAI ·

    我们正在推出 GeneBench-Pro,一个用于更困难的 AI 进展的研究级基准:代理在处理混乱的生物数据、选择正确

    We’re introducing GeneBench-Pro, a research-level benchmark for a harder kind of AI progress: how well agents can navigate messy biological data, choose the right analysis path, and make judgment calls that real computational research depends on. https://t.co/AsilnnSxnE

  4. OpenAI News TIER_1 English(EN) ·

    助力构建先进人工智能的共享标准

    OpenAI helps build shared standards for advanced AI, supporting evaluation frameworks, safety practices, and global cooperation through the Appia Foundation.

  5. X — Google DeepMind TIER_1 English(EN) · GoogleDeepMind ·

    当数百万个AI代理相互交互时,可能会出现新的集体行为。🌐

    When millions of AI agents interact with each other, new collective behaviors can emerge. 🌐 Together with @schmidtsciences, @coop_ai, @ARIA_research and supported by @GoogleOrg, we’re launching a $10M research fund to help understand how AI systems behave as a group. → https://t…

  6. OpenAI News TIER_1 English(EN) ·

    支持欧洲建立值得信赖的AI生态系统

    OpenAI supports the EU Code of Practice on AI content transparency, advancing provenance standards and tools to help people understand AI-generated content.

  7. Google DeepMind TIER_1 English(EN) ·

    投资多智能体AI安全研究

    Google DeepMind and partners announce a $10M funding call for multi-agent safety research.

  8. OpenAI News TIER_1 English(EN) ·

    从数据到决策:LSEG 如何扩展可信赖的 AI

    See how LSEG uses OpenAI to scale trusted AI across its global business, accelerating insights, shrinking release cycles, and empowering 4,000 employees.

  9. Google AI / Research TIER_1 English(EN) ·

    利用 Gemini Enterprise Agent Platform 的 Agentic RAG 实现可靠响应

    Data Management

  10. OpenAI News TIER_1 English(EN) ·

    Endava如何围绕AI代理重新设计软件交付

    Learn how Endava is using AI agents, ChatGPT Enterprise, and Codex to accelerate software delivery, automate workflows, and build an AI-native culture across the enterprise.

  11. Meta AI blog TIER_1 English(EN) ·

    两年四颗MTIA芯片:为数十亿用户扩展AI体验

    Serving a wide range of AI models on a global scale, while maintaining the lowest possible costs, is one of the most demanding infrastructure challenges in the industry.

  12. Meta AI blog TIER_1 English(EN) ·

    扩展我们构建和测试最先进AI的方式

    As we build more capable, personalized AI, reliability, security, and user protections are more important than ever.

  13. OpenAI News TIER_1 English(EN) ·

    为更安全、更透明的AI生态系统推进内容溯源

    OpenAI advances AI content provenance with Content Credentials, SynthID, and a verification tool to help people identify and trust AI-generated media.

  14. OpenAI News TIER_1 English(EN) ·

    Sea 对 Codex 在代理软件开发未来应用的看法

    Sea Limited's CPO explains why the company is deploying Codex across engineering teams to accelerate AI-native software development in Asia.

  15. Google DeepMind TIER_1 English(EN) ·

    Co-Scientist:加速研究的多智能体AI伙伴

    Introducing Co-Scientist, a collaborative AI partner built with Gemini to help researchers accelerate scientific breakthroughs.

  16. Google AI / Research TIER_1 English(EN) ·

    TurboQuant:以极致压缩重新定义AI效率

    Algorithms & Theory

  17. OpenAI News TIER_1 English(EN) ·

    Harness工程:在以Agent为先的世界中利用Codex

    By Ryan Lopopolo, Member of the Technical Staff

  18. Google AI / Research TIER_1 English(EN) ·

    迈向智能体系统科学:智能体系统何时以及为何有效

    Generative AI

  19. Google AI / Research TIER_1 English(EN) ·

    探索一种基于太空的、可扩展的AI基础设施系统设计

    General Science

  20. Google DeepMind TIER_1 English(EN) ·

    推出CodeMender:一款用于代码安全的AI代理

    Using advanced AI to fix critical software vulnerabilities

  21. Google AI / Research TIER_1 English(EN) ·

    Coral NPU:面向边缘AI的全栈平台

    Generative AI

  22. OpenAI News TIER_1 English(EN) ·

    推出 AgentKit、新的 Evals 和 RFT 以用于代理

    Today, we’re releasing new tools to help developers go from prototype to production faster: AgentKit, expanded evals capabilities, and reinforcement fine-tuning for agents.

  23. Google AI / Research TIER_1 English(EN) ·

    人工智能作为研究伙伴:用AlphaEvolve推进理论计算机科学

    Algorithms & Theory

  24. Google DeepMind TIER_1 English(EN) ·

    AlphaEvolve:一个由Gemini驱动的编码代理,用于设计高级算法

    New AI agent evolves algorithms for math and practical applications in computing by combining the creativity of large language models with automated evaluators

  25. OpenAI News TIER_1 English(EN) ·

    使用电脑的代理

  26. Microsoft Research TIER_1 Nederlands(NL) · Akshay Nambi, Yash Pandya, Sahil Gupta, Sarthak Harne, Archana Yadav, Kavyansh Chourasia, Yash Lara, Ahmed Awadallah, Ece Kamar ·

    Echoverse:计算机使用代理的深度、演进式环境

    <p>Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the tasks, tests, and environments evolve.</p> <p>The post <a …

  27. Hugging Face Blog TIER_1 (CA) ·

    用于 Agent 的数据

  28. Apple Machine Learning Research TIER_1 English(EN) ·

    Weblica:视觉化网络代理的可扩展且可复现的训练环境

    The web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection attempts remain limited to offline trajectories for supervised fine-tuning or a handful of simulated environments for RL training,…

  29. Hugging Face Blog TIER_1 English(EN) ·

    ScarfBench:为企业 Java 框架迁移进行 AI Agent 性能基准测试

  30. Microsoft Research TIER_1 Norsk(NO) · Yifan Yang, Xuemei Gao, Qi Dai, Bei Liu, Kai Qiu, Dongdong Chen, Chong Luo ·

    SkillOpt: Agent skills as trainable parameters

    <p>AI agents often fail because their instructions, or skills, are manually modified with no guarantee of improvement. Learn how SkillOpt turns skill editing into a training process, making agent behavior more reliable without changing model weights.</p> <p>The post <a href="http…

  31. Hugging Face Blog TIER_1 English(EN) ·

    欢迎 NVIDIA Cosmos 3:首个用于物理人工智能推理和行动的开放式全能模型

  32. Microsoft Research TIER_1 English(EN) · Ken Archer, Harald Wiltsche ·

    通过人工智能拓展人类智能

    <p>Understanding AI as an extension of human intelligence—not a replacement for it—offers a more grounded path for building trustworthy AI systems.</p> <p>The post <a href="https://www.microsoft.com/en-us/research/blog/extending-human-intelligence-through-ai/">Extending Human Int…

  33. Hugging Face Blog TIER_1 English(EN) ·

    Harness、Scaffold 和值得正确理解的 AI Agent 术语

  34. Microsoft Research TIER_1 English(EN) · Microsoft Research AI Frontiers ·

    MagenticLite, MagenticBrain, Fara1.5:为小型模型优化的智能体体验

    <p>MagenticLite is an agentic system for small models that works across the browser and local file system in a single workflow. It combines specialized models and orchestration to support efficient agentic performance on everyday tasks.</p> <p>The post <a href="https://www.micros…

  35. Qwen tech blog TIER_1 Nederlands(NL) · QwenTeam ·

    Qwen3.7:智能体前沿

    Today we introduce Qwen3.7-Max, our latest proprietary model designed for the agent era. Qwen3.7-Max is built to be a versatile agent foundation — equally capable of writing and debugging code, automating office workflows, and sustaining autonomous execution across hundreds or th…

  36. Qwen tech blog TIER_1 English(EN) · QwenTeam ·

    Qwen3.6-Plus:迈向真实世界代理

    Following the release of the Qwen3.5 series in February, we are thrilled to announce the official launch of Qwen3.6-Plus. Available immediately via our API, this release represents a massive capability upgrade over its predecessor. Most notably, we have drastically enhanced the m…

  37. Hugging Face Blog TIER_1 English(EN) ·

    Python中的微型代理:一个约70行代码的MCP驱动代理

  38. Hugging Face Blog TIER_1 English(EN) ·

    Tiny Agents:一个由 MCP 驱动的 50 行代码代理

  39. Hugging Face Blog TIER_1 English(EN) ·

    推出 smolagents:用代码编写动作的简单代理。

  40. arXiv cs.LG TIER_1 English(EN) · Yuhang Yao, Zeyu Wang, Wanyi Chen, Tongyun Yang, Yuhang Han, Jie Xiao, Chengke Bao, Tianyi Zhao, Lynn Ai, Eric Yang, Tianyu Shi ·

    MERA:面向大规模智能体系统的模型演进与技能自适应路由

    arXiv:2608.10333v1 Announce Type: new Abstract: LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured steps such as formatting or tool-argument construction. Prior routing methods ex…

  41. arXiv cs.AI TIER_1 English(EN) · Giuseppe Destefanis, Daniel Graziotin, Matteo Vaccargiu, Marco Ortu ·

    GitSkills: GitHub上的Agent技能数据集

    arXiv:2608.10906v1 Announce Type: cross Abstract: An agent skill is a folder containing a SKILL.md file with instructions for a language-model agent, optionally accompanied by scripts and reference files. The agent loads the skill when it judges that a task matches the skill desc…

  42. arXiv cs.AI TIER_1 English(EN) · Scott E. Frias ·

    相似性门控批准逆转:代理系统中嵌入余弦阈值的有效性审计

    arXiv:2608.10216v1 Announce Type: cross Abstract: Agent frameworks ship quality gates that compare text blocks by embedding-cosine similarity and decide at a fixed cutoff. Deduplication filters, semantic caches, drift guards, and answer grader gates deploy to answer the question:…

  43. arXiv cs.AI TIER_1 English(EN) · Hanrong Zhang (Steve), Shicheng Fan (Steve), Henry Peng Zou (Steve), Yankai Chen (Steve), Zhenting Wang (Steve), Jiayu Zhou (Steve), Chengze Li (Steve), Wei-Chieh Huang (Steve), Yifei Yao (Steve), Kening Zheng (Steve), Xue (Steve), Liu, Xiaoxiao Li, Ph… ·

    CoEvoSkills:通过协同进化验证实现自主进化的智能体技能

    arXiv:2604.01687v3 Announce Type: replace Abstract: Anthropic proposes the concept of skills for LLM agents to tackle multi-step professional tasks that simple tool invocations cannot address. A tool is a single, self-contained function, whereas a skill is a structured bundle of …

  44. arXiv cs.AI TIER_1 English(EN) · Bhaskar Gurram ·

    工具使用语言代理中的自动化评估、错误传播和运行时缓解审计

    arXiv:2604.16706v2 Announce Type: replace Abstract: Automated evaluation of tool-using large language model (LLM) agents is widely assumed to be reliable, yet this assumption is rarely validated against human annotation. We present AgentProp-Bench, a diagnostic benchmark of 14,75…

  45. arXiv cs.LG TIER_1 English(EN) · Shuo Hao, You Lu, Bihuan Chen, Xin Peng ·

    FlowScout:从执行反馈到可靠的工具使用代理工作流

    arXiv:2608.10039v1 Announce Type: new Abstract: Agentic workflows have become an important abstraction for building reliable LLM-based automation systems by organizing large language models (LLMs), tools, and control logic into explicit execution structures. However, constructing…

  46. arXiv cs.AI TIER_1 English(EN) · Charles L. Wang, Keir Dorchen, Peter Jin ·

    论自改进智能体的统计学极限

    arXiv:2510.04399v3 Announce Type: replace Abstract: We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this framework, we prove a sharp boundary: under standard i.i.d. assumptions, distribution-fre…

  47. arXiv cs.LG TIER_1 English(EN) · Yanyu Ren, Xizheng Wang, Xiao Liu, Bowen Lv, Hanchen Zhang, Shudan Zhang, Hanyu Lai, Shuai Wang, Li Chen, Dan Li, Jie Tang ·

    TideRL:通过就绪感知调度提升智能体 RL 的 Goodput

    arXiv:2608.10402v1 Announce Type: new Abstract: Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks repeatedly pause for external environments, resume with growing contexts, and finish at highly variable times. …

  48. arXiv cs.AI TIER_1 English(EN) · Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub, Md Tahmid Rahman Laskar, Shafiq Joty, Enamul Hoque Prince ·

    DSAgentBench:代理能否在真实的计算机环境中自动化端到端数据科学工作流?

    arXiv:2608.10366v1 Announce Type: new Abstract: Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and databases…

  49. arXiv cs.AI TIER_1 English(EN) · Au Kwok Chun, Abhigyan Acherjee, Amrutha Rao, Zaiqian Chen, Kazem Meidani, C. Bayan Bruss, Micah Goldblum ·

    在自动研究代理中恢复浪费的计算资源

    arXiv:2608.10424v1 Announce Type: new Abstract: A slew of recent works develop agents for solving research problems end-to-end, a paradigm increasingly referred to as autoresearch. Such agents have inspired large industry investment, motivated by their potential to automate time-…

  50. arXiv cs.AI TIER_1 English(EN) · Jung Hwan Lee, Kyu Ho Lee, Gwang Hoon Yoo ·

    MEGA:通过智慧图谱实现自演化智能体优化基础设施

    arXiv:2608.10504v1 Announce Type: new Abstract: As coding agents increasingly handle implementation, the central challenge shifts from building individual agents to building an infrastructure that systematically improves them. Current approaches optimize agent systems without acc…

  51. arXiv cs.AI TIER_1 English(EN) · Fanqi Zhou, Qiaosheng Chen, Zixian Huang, Gong Cheng ·

    Agentic Instruction Data Selection: Let DataMaster Interpret Your Intent

    arXiv:2608.10579v1 Announce Type: new Abstract: Although existing instruction data selection methods have introduced various metrics, the inherent complexity of real-world datasets makes it impractical for any single metric to generalize across all scenarios. Developers are thus …

  52. arXiv cs.CL TIER_1 English(EN) · Qing Zong, Jiayu Liu, Junhao Shen, Zecong Tang, Linsi Wu, Yuxuan Liu, Rui Wang, Zhaowei Wang, Weiqi Wang, Cheng Qian, Xiusi Chen, Yangqiu Song ·

    Agentic系统中的协同进化:迈向超越人类设计的自主进化

    arXiv:2608.10299v1 Announce Type: new Abstract: Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often bounded by a static learning context, such as fixed tasks and feedback. This survey focuses on co-evolution in agentic …

  53. arXiv cs.AI TIER_1 English(EN) · Zixing Chen, Xingyuan Liu, Jie Zhu, Huaixia Dou, Shuo Jiang, Junhui Li, Lifan Guo, Feng Chen, Chi Zhang ·

    REDAgentBench:LLM Agent系统的可执行红队测试与忠实测量

    arXiv:2608.10669v1 Announce Type: new Abstract: Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interactions between the agent and its environment, causing the agent to violate safety …

  54. arXiv cs.AI TIER_1 English(EN) · Xiaofan Bai, Hongqiang Lin, Chao Liu, Yantao Zhang, Xuan Jin, Xipeng Cao, Yuhong Li ·

    SkillZip:通过发现可重用结构实现无需评估的技能压缩,用于自进化智能体

    arXiv:2608.11079v1 Announce Type: new Abstract: Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is often restated in several branches, examples, and warnings, while common action sequences are c…

  55. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Tianwei Zhang ·

    ASCon:一种面向方向的互惠代理-步长上下文模型,用于多智能体系统中的故障归因

    Failure attribution in LLM-based multi-agent systems (MAS) aims to answer who caused failures, when they occurred, and why by identifying responsible targets including faulty agents, erroneous steps, and failure modes. Existing methods have primarily focused on developing dedicat…

  56. arXiv cs.AI TIER_1 English(EN) · Dongyi Lv, Fushun E, Aichen Cai, Liang Huang, Ya Zhang, Qiuyu Ding, Canhui Wu, Zhi Wang, Yuesong Zhang, Jiaqi Wang, Nan Duan ·

    FailForge:从持续失败中提炼程序化能力并将其转化为代码代理

    arXiv:2608.08570v1 Announce Type: new Abstract: Rejection sampling fine-tuning (RFT) is widely used to train code agents by generating trajectories on verifiable software engineering tasks, retaining those that pass the tests, and fine-tuning on the successful rollouts. However, …

  57. arXiv cs.AI TIER_1 English(EN) · Zejun Xu, Taiyi Chen, Jin Li, Yongtong Gu, Qi Cheng, Aixuan Lv, Shuai Zhu, Pengfei Zhu, Kaichen Yang, Boyu Sun, Yixian Yang, Mulong Xie, Xin Liu, Dagang Li, Xiaoteng Ma, Hongru Wang ·

    CAP:一个用于评估具有复杂动作和感知能力的跨站点浏览器代理的可扩展基准

    arXiv:2608.08392v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has been driven by benchmarks that evaluate end-to-end task success, these evaluations largely o…

  58. arXiv cs.AI TIER_1 English(EN) · Zhengyang Shan, Xu Qian, Jiayun Xin, Kun Li, Yue Zhang, Minghui Xu ·

    OBLIVION:已部署Agent的工作流级操作技能遗忘

    arXiv:2608.08264v1 Announce Type: new Abstract: Large language model agents are becoming operational interfaces to files, memories, registries, and external tools. This deployment shift creates a new skill revocation problem: after a skill is removed from an explicit registry, an…

  59. arXiv cs.AI TIER_1 English(EN) · Xinle Jiang, Remy Xie, Ming Tang ·

    SkillSmith:通过自动技能构建和演进增强本地部署的代理

    arXiv:2608.08037v1 Announce Type: new Abstract: LLM-based agent frameworks now act as personal assistants for multi-step tasks. Existing agent frameworks such as OpenClaw commonly follow the Cloud Agent depolyment mode using closed-source cloud LLMs as backbone model, which may e…

  60. arXiv cs.AI TIER_1 English(EN) · Chi Zhang, Yimin Liu, Xinze Chen, Ping Ji ·

    是什么阻碍了 Agent Skills 的可重用性?来自 138K SKILL.md 文件的证据

    arXiv:2608.08453v1 Announce Type: new Abstract: Under the current standard, Agent Skills are SKILL.md files that combine instructions with supporting resources, enabling Large Language Model (LLM) agents to reuse procedures beyond a single conversation. Yet many public skills app…

  61. arXiv cs.AI TIER_1 English(EN) · Donghong Jiang, Endian Lin, Luoping Cui, Hanqing Liu, Mingjie Liu, Fan Yang, Hong Wang, Zhao Yang, Chuang Zhu ·

    SkillReason: 面向隐式用户请求的增强推理的代理技能检索

    arXiv:2608.08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the appropriate skill from a large- scale library remains challenging because realistic…

  62. arXiv cs.AI TIER_1 English(EN) · Marc Alier Forment, Mar\'ia Jos\'e Casa\~n Guerrero, Francisco Jos\'e Garc\'ia-Pe\~nalvo, Juanan Pereira ·

    脚手架比界面更重要:对七种代理脚手架、五种语言模型和一项软件任务的 MCP 和 CLI 工具使用的对照比较

    arXiv:2608.08654v1 Announce Type: new Abstract: How much an AI coding agent costs to run can depend more on the agent scaffolding that drives it than on the interface through which it reaches its tools. We set out to measure the cost of tool use over the Model Context Protocol (M…

  63. arXiv cs.AI TIER_1 English(EN) · Siqi Wang, Xinlin Li, Zhenglin Li, Li Li ·

    OpenLoopEvolve:用于长时域复杂任务中循环策略的可验证自演化框架

    arXiv:2608.09380v1 Announce Type: new Abstract: Long-horizon complex tasks require agents to repeatedly observe states, formulate plans, invoke tools, verify results, and recover from failures in continuously changing environments. However, such control experience often remains c…

  64. arXiv cs.AI TIER_1 English(EN) · Neel Tushar Shah, Manglam Kartik, Akshat Karkar ·

    能力并非倾向:衡量公民LLM代理的压力鲁棒合作行为

    arXiv:2608.09485v1 Announce Type: new Abstract: Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strategic omission, false consensus, and manipulative framing. We argue that Cooperative AI evaluat…

  65. arXiv cs.AI TIER_1 English(EN) · Hui Xue, Fan Yang ·

    重新思考自主进化代理:我们还需要预设的优化流程吗?

    arXiv:2608.09629v1 Announce Type: new Abstract: Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a persistent artifact, select candidates, and stop. We ask whether this task-specific procedure re…

  66. arXiv cs.AI TIER_1 English(EN) · Tailin Zhou ·

    分层自我改进:一种用于任务特定可进化代理的框架

    arXiv:2608.08466v1 Announce Type: new Abstract: Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the \emph{harness}---is typically treated as a fixed artifact after deployment. This work…

  67. arXiv cs.AI TIER_1 English(EN) · Hanye Zhao, Muning Wen, Yong Yu, Weinan Zhang ·

    MARA:面向计算资源高效学习的流匹配引导多智能体资源分配

    arXiv:2608.09130v1 Announce Type: cross Abstract: Allocating limited computation among concurrent learning tasks is difficult when each task must reach a target loss before a deadline but its required training effort is unknown. Existing approaches combine online loss prediction …

  68. arXiv cs.AI TIER_1 English(EN) · Puzhen Yuan, Angyuan Ma, Yunchao Yao, Huaxiu Yao, Masayoshi Tomizuka, Mingyu Ding ·

    REMAC:用于长时域机器人操作的自反思、自演化多智能体协作

    arXiv:2503.22122v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that require a holistic understanding of the environment for task decomposition. Existing met…

  69. arXiv cs.AI TIER_1 English(EN) · Xinze Chen, Chi Zhang, Ping Ji, Yimin Liu ·

    SkillsMetric:静态分析对恶意代理技能检测边界的映射

    arXiv:2608.08468v1 Announce Type: cross Abstract: Agent Skills---structured packages of instructions and scripts that augment LLM-based agents---are rapidly proliferating, yet their security properties remain under-explored. We present \textsc{SkillsMetric}, a five-stage static a…

  70. arXiv cs.AI TIER_1 English(EN) · Peiwen Li, Shiyang Zhang, Yangtian Zhang, Sizhuang He, David van Dijk, Rex Ying ·

    MoRSE:具有角色-子任务专家混合的任务导向型多智能体系统

    arXiv:2608.09251v1 Announce Type: cross Abstract: Large language model-based multi-agent systems have recently shown strong potential for complex, long-horizon tasks. However, existing methods mainly rely on coarse prompt-level differentiation without parameter adaptation for div…

  71. arXiv cs.AI TIER_1 English(EN) · Libin Qiu, Zhirong Gao, Junfu Chen, Yuhang Ye, Liangyu Li, Weizhi Huang, Xiaobo Xue, Wenkai Qiu, Shuo Tang ·

    AutoRefine:将轨迹编译成经过验证的类型化代理工件

    arXiv:2601.22758v2 Announce Type: replace Abstract: Large language model agents repeatedly encounter related tasks, yet systems that learn from trajectories commit every lesson to one predefined artifact form. A local constraint, a reusable procedure, and a delegated objective re…

  72. arXiv cs.AI TIER_1 English(EN) · Jiaru Bai, Abdulrahman Aldossary, Thomas Swanick, Marcel M\"uller, Yeonghun Kang, Changhyeok Choi, Naruki Yoshikawa, Zijian Zhang, Jin Won Lee, Tsz Wai Ko, Aiwei Yin, Mohammad Ghazi Vakili, Chris Crebolder, Varinia Bernales, Al\'an Aspuru-Guzik ·

    El Agente Gráfico: 科学代理的语义执行运行时

    arXiv:2602.17902v2 Announce Type: replace Abstract: Large language models (LLMs) can plan scientific workflows and generate code, but these capabilities do not specify how scientific state is validated, transferred and recorded across heterogeneous computational and experimental …

  73. arXiv cs.AI TIER_1 English(EN) · Ziyu Ma, Shidong Yang, Yuxiang Ji, Xucong Wang, Yong Wang, Yiming Hu, Tongwen Huang, Xiangxiang Chu ·

    SkillClaw:让技能随Agentic Evolver集体进化

    arXiv:2604.08377v2 Announce Type: replace Abstract: Large language model (LLM) agents such as OpenClaw rely on reusable skills to perform complex tasks, yet these skills remain largely static after deployment. As a result, similar workflows, tool usage patterns, and failure modes…

  74. arXiv cs.AI TIER_1 English(EN) · Chaofan Meng, Yuhang Zheng, Yingnan Zhou, Sihan Xu ·

    SkillConsist:通过双向图对齐检测智能体技能中的不一致性

    arXiv:2608.07639v1 Announce Type: cross Abstract: Agent Skills provide reusable capabilities to LLM agents. Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wrong Skill selection. Recent Agent Skill research has increasingly examined Agent Skill cons…

  75. arXiv cs.AI TIER_1 English(EN) · Keyang Zhong, Junlin Xie, Hefeng Wu, Haofeng Li, Guanbin Li ·

    为增强推理不完美信息谋杀之谜游戏而进行的协作多智能体脚本生成

    arXiv:2604.11741v2 Announce Type: replace Abstract: Vision-language models (VLMs) have shown impressive capabilities in perceptual tasks, yet they degrade in complex multi-hop reasoning under multiplayer game settings with imperfect and deceptive information. In this paper, we st…

  76. arXiv cs.AI TIER_1 English(EN) · Neharika Jali, Anupam Nayak, Gauri Joshi ·

    并非所有转折都同样困难:用于智能体高效多轮推理的自适应思考预算

    arXiv:2604.05164v3 Announce Type: replace-cross Abstract: As LLM reasoning performance plateaus, improving inference-time compute efficiency is crucial to mitigate overthinking and long thinking traces even for simple queries. Prior approaches including length regularization, ada…

  77. arXiv cs.CL TIER_1 English(EN) · Xueping Gao ·

    面向异构编码代理的证据校准运行时重构以提升代理技能

    arXiv:2608.08793v1 Announce Type: new Abstract: Agent Skills package reusable instructions and assets for tool-using language-model agents. Progressive loading creates failure boundaries poorly represented by session-, model-, or tool-centric traces: a Skill can be discovered but…

  78. arXiv cs.CL TIER_1 English(EN) · Zihao Deng, Yining Zhu, Leiming Wang, Jingfei Lu, Junbo Wang, Chuncheng Ran, Yu Yang, Dixuan Yang, Jikun Shen ·

    Tree-of-Experience: 层次化经验管理助力自进化智能体

    arXiv:2608.09044v1 Announce Type: new Abstract: Continual self-evolution requires LLM agents to transform environmental interactions into reliable and reusable experience. Existing methods typically refine individual trajectories or abstract shared knowledge from related trajecto…

  79. arXiv cs.CL TIER_1 English(EN) · Lisheng Huang, Chen Yang, Hao Zhou, Huatong Song, Zongchao Chen, Ran Le, Yang Song, Wayne Xin Zhao, Tao Zhang ·

    Evo-Bench:语言模型能否改进Agent Harness?

    arXiv:2608.09096v1 Announce Type: new Abstract: Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize…

  80. arXiv cs.AI TIER_1 English(EN) · Hongwei Yao, Yiming Liu, Meihui Chen, Jieling Chen, Zikun Chen, Yiling He, Wangze Ni, Cong Wang, Kui Ren ·

    ActBench:协作代理行为安全性的自演进基准

    arXiv:2608.09476v1 Announce Type: cross Abstract: Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavioral safety and introduce ActBench, a self-evolving benchmark that evaluates such…

  81. arXiv cs.AI TIER_1 English(EN) · Hanlin Jiang, Jionghao Huang, Shaofei Li, Bojia Yu, Peng Jiang, Yuxin Ren, Ning Jia, Yao Guo, Ding Li ·

    STAIR:使用端到端代理规划框架实现有效的事件响应

    arXiv:2608.09524v1 Announce Type: cross Abstract: Incident response planning is critical for restoring compromised software systems after cyberattacks. Common practice relies on expert-driven playbooks that encode fixed response procedures, but these static workflows struggle to …

  82. Hugging Face Daily Papers TIER_1 English(EN) ·

    DSAgentBench:代理能否在真实的计算机环境中自动化端到端的数据科学工作流?

    DSAgentBench evaluates autonomous agents on complete, multi-tool data-science workflows in real computing environments and reveals major performance gaps.

  83. Hugging Face Daily Papers TIER_1 English(EN) ·

    SkillZip:通过发现可重用结构实现无需评估的技能压缩,用于自进化智能体

    SkillZip compresses self-evolving agent skills by finding a minimal faithful structural explanation that shares repeated rules and procedures while preserving rare exceptions, without requiring evaluation rollouts.

  84. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Rex Ying ·

    MoRSE:具有角色-子任务专家混合的任务导向型多智能体系统

    Large language model-based multi-agent systems have recently shown strong potential for complex, long-horizon tasks. However, existing methods mainly rely on coarse prompt-level differentiation without parameter adaptation for diverse subtasks, resulting in insufficient inter-age…

  85. Hugging Face Daily Papers TIER_1 English(EN) ·

    MoRSE:具有角色-子任务专家混合的任务导向型多智能体系统

    Large language model-based multi-agent systems have recently shown strong potential for complex, long-horizon tasks. However, existing methods mainly rely on coarse prompt-level differentiation without parameter adaptation for diverse subtasks, resulting in insufficient inter-age…

  86. arXiv cs.AI TIER_1 English(EN) · Daniel Koh Ji Yang, Yannic Noller, Corina S. Pasareanu, Youcheng Sun ·

    用于符号执行的代理规划

    arXiv:2608.06397v1 Announce Type: cross Abstract: Symbolic execution seeks to explore feasible program paths, yet a practical run may exhaust its resources while much program behaviour remains unreached. We investigate a complementary way of extending its practical reach by reaso…

  87. arXiv cs.AI TIER_1 (CA) · Mingxuan Zheng, Yujin Zhou, Chuxue Cao, Boqin Yin, Yuyao Zhang, Jiapeng Sun, Shuaishuai Gong, Sirui Han, Yike Guo ·

    SkillProx:通过近端文本梯度下降实现自主进化代理技能

    arXiv:2608.07449v1 Announce Type: new Abstract: LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These skills are lightweight, reusable textual artifacts that are loaded into the agent's context without weight updates. Recent method…

  88. arXiv cs.AI TIER_1 English(EN) · Chao Fei, Qingyi Si, Kaihua Liang, Yanghua Xiao, Panos Kalnis, Hongcheng Guo ·

    EMAS:通过证据指导的修订稳定多智能体系统演化

    arXiv:2608.07196v1 Announce Type: new Abstract: Many methods for automated multi-agent system design optimize prompts and topologies during an initial design stage and then deploy the resulting system unchanged on subsequent samples. Experience from these samples is rarely consol…

  89. arXiv cs.AI TIER_1 English(EN) · Haoning Wang, Mingxun Zhang, Chenyue Yu, Yingjun Shang, Xia Hu, Guanchu Wang, Na Zou ·

    端到端代理审计引擎

    arXiv:2608.07346v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains. The fast-evolving harness ecosystem has also made rigorous capability ev…

  90. arXiv cs.AI TIER_1 English(EN) · Xiao Zhang, Yusheng Wang, Yuhao Fei, Dongyuan Li, Zian Liang, Liuyu Xiang, Hongxun Gu, Zhaofeng He ·

    HarnessSafe:评估代理马具中持久载体的安全性

    arXiv:2608.06984v1 Announce Type: cross Abstract: Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills, tools, and shared artifacts. However, this capability creates delayed safety risks: attacker-influenced content can cr…

  91. arXiv cs.AI TIER_1 English(EN) · Lekang Jiang, Bohan Tang, Stephan Goetz, Yiwen Guo ·

    ADIAS:交互式Agentic系统的自动化设计

    arXiv:2608.06410v1 Announce Type: new Abstract: Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization. Existing methods are largely candidate-centric: cross-round experience is organized around candidate agents, which l…

  92. arXiv cs.AI TIER_1 English(EN) · Jing Chen, Yang Sun, Li Zhang, Lin Xu, Jie Shi ·

    长时域智能体轨迹归因:统一基准与细粒度标注框架

    arXiv:2608.06909v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external observations, and memory. Existing benchmarks primarily evaluate behavioral outcomes but provid…

  93. arXiv cs.AI TIER_1 English(EN) · Shuyang Liu, Saman Dehghan, Ji Young Kim, Jatin Ganhotra, Martin Hirzel, Reyhaneh Jabbarvand ·

    在线监控和纠正性引导编程代理

    arXiv:2608.06701v1 Announce Type: cross Abstract: Fixing GitHub issues in large-scale projects is a long-horizon task, especially when a fix requires changes across multiple locations or the issue description lacks the information needed to localize and repair it. As a result, ag…

  94. arXiv cs.AI TIER_1 English(EN) · Yingtao Tian ·

    CEDAR:面向目标导向的复杂系统优化代理编排树搜索

    arXiv:2608.06871v1 Announce Type: new Abstract: Complex systems, core objects of study in artificial life, model diverse phenomena through nonlinear, feedback-driven interactions that produce emergent behavior, with applications from population dynamics and biology to economic po…

  95. arXiv cs.AI TIER_1 English(EN) · Jiahui Han, Qinuo Li, Ziheng Peng, Haotian Wu, Haoze Liu, Danfeng Shan, Guanchu Wang, Huiqi Deng, Ninghao Liu ·

    SkillEval:将智能体技能质量分解为可解释的信号

    arXiv:2608.06891v1 Announce Type: new Abstract: Agent skills provide reusable procedural knowledge that helps agents solve specialized tasks. As their use expands, evaluating skill quality becomes increasingly important. Existing evaluations often measure skill quality by testing…

  96. arXiv cs.AI TIER_1 English(EN) · Junbo Li, Boyi Liu, Canwen Xu, Yite Wang, Yuxiong He, Zhangyang Wang, Qiang Liu, Zhewei Yao ·

    优化器即智能体:跨越提示、程序和机器学习工作流的驱动式搜索

    arXiv:2608.06714v1 Announce Type: new Abstract: Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods. We ask a fundamentally different question: how mu…

  97. Hugging Face Daily Papers TIER_1 English(EN) ·

    Evo-Bench:语言模型能否改进Agent Harness?

    Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However, systematica…

  98. Hugging Face Daily Papers TIER_1 English(EN) ·

    Evo-Bench:语言模型能否改进Agent Harness?

    Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However, systematica…

  99. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agentic 系统中的协同进化:迈向超越人类设计的自主进化

    Agentic systems can achieve open-ended improvement through multi-component co-evolution that progressively removes fixed human constraints across agents, environments, and evolution mechanisms.

  100. arXiv cs.AI TIER_1 English(EN) · Xi Wang, Kun Li, Xianyao Ling, Gang Yin, Liang Zhang, Jiang Wu, Wenbo Lei, Jun Xu, Annie Wang, Fu Zhang, Weizhe Wang ·

    Agentic Nesting:现有企业应用集成与服务的新方法论

    arXiv:2608.05159v1 Announce Type: new Abstract: Enterprise operations extensively rely on multiple heterogeneous business systems and information applications, which also result in severe data silos and process fragmentation. Enterprises have invested considerable financial and m…

  101. arXiv cs.AI TIER_1 English(EN) · Indivara Kolluru, Nathan Sportsman ·

    大型技能库中的代理检索的比较方法

    arXiv:2608.06196v1 Announce Type: new Abstract: Agents backed by large skill libraries must decide which skills to load and in what order. Loading the entire library into context is expensive and provides no structure for autonomous sequencing. We study two systems for this probl…

  102. arXiv cs.AI TIER_1 English(EN) · Bo Deng (Beihang University, Qwen DianJin Team, Alibaba Cloud Computing), Kang Zhou (Qwen DianJin Team, Alibaba Cloud Computing), Lifan Guo (Qwen DianJin Team, Alibaba Cloud Computing), Chongyang Tao (Beihang University), Xuanren Chen (Beihang University… ·

    FinEvo-Bench:专业金融工作流中自进化代理的纵向基准测试

    arXiv:2608.06144v1 Announce Type: new Abstract: Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks. Existing self-evolution benchmarks do not jointly cover professional workflows, open-ended deliverables, …

  103. arXiv cs.AI TIER_1 English(EN) · Jiacheng Wei, Zhaoxin Fan, Xin Wen, Yuqin Lan, Dongrun Li, Wenjun Wu, Faguo Wu, Xiao Zhang ·

    ChainClaw:一个用于可靠链上执行的分层代理框架

    arXiv:2608.05790v1 Announce Type: new Abstract: General-purpose large language model agents have achieved strong performance on tool-augmented tasks, yet they rely on assumptions break down in blockchain environments. On-chain execution is stateful, adversarial, and economically …

  104. arXiv cs.AI TIER_1 English(EN) · Weihong Lin, Lin Sun, Xiangzheng Zhang ·

    提示端智能体剧本何时可迁移?智能体部署中的准确性、成本和运行时变化

    arXiv:2608.05778v1 Announce Type: new Abstract: Prompt-side playbooks can improve tool-using language agents without retraining, but their portability beyond the source setting is unclear. We study frozen playbook transfer under a shared distill--validate--transfer protocol. On A…

  105. arXiv cs.AI TIER_1 English(EN) · Xinshuang Liu, Runfa Blark Li, Shaoxiu Wei, Xin Lin, Truong Nguyen ·

    统一代理:跨设备管理交互

    arXiv:2608.05729v1 Announce Type: new Abstract: As capabilities rapidly increase, AI agents can move from running inside one app to acting across a user's devices over time. Yet existing agent systems still fall short in this scenario. This is because observations are scattered a…

  106. arXiv cs.AI TIER_1 English(EN) · Yuru Feng, Yaoqi Chen, Beidi Zhao, Qianxi Zhang, Xinjiang Wang, Jianan Lu, Zhirui Wang, Shusen Xu, Zengzhong Li, Qi Chen ·

    SkillHEX:通过假设驱动的自主探索与利用来提升智能体技能

    arXiv:2608.05628v1 Announce Type: new Abstract: Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment. Real-world deployments thus require autonomous, on-demand skill evolution at test time…

  107. arXiv cs.AI TIER_1 English(EN) · Zhi Han, Chenxi Zeng, Liuhaichen Yang, Zihan Guo, Ming Zhou, Yang Li ·

    SkillTV-Bench:评测法官在技能增强代理执行方面的表现

    arXiv:2608.05573v1 Announce Type: new Abstract: LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executions. For skill-augmented agents, verification additi…

  108. arXiv cs.AI TIER_1 English(EN) · Yidian Chen, Yingzi Gu, Natan Vidra, Spurthi Setty, Sharon Zheng ·

    OrchestraBench:评估多智能体编排失败模式、恢复和分解质量

    arXiv:2608.05263v1 Announce Type: new Abstract: Multi-agent orchestration frameworks are moving from demos to production, yet benchmarks typically report task accuracy without diagnosing why a pipeline failed, where a cascade began, or which routing decision caused the breakdown.…

  109. arXiv cs.AI TIER_1 English(EN) · Zhixiang Liang, Yifei Liu, Yidan Huang, Haozhe Zhao, Beichen Huang, Jiaqi Wang, Nan Duan, Qiong Cao ·

    SearchAuditor: 审计和归因长时域搜索代理中的失败

    arXiv:2608.05212v1 Announce Type: new Abstract: Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long, noisy trajectories into fluent but incorrect answe…

  110. arXiv cs.AI TIER_1 English(EN) · Jingzhi Gong, Ruizhen Gu, Zhiwei Fei, Yazhuo Cao, Lukas Twist, Alina Geiger, Shuo Han, Dominik Sobania, Federica Sarro, Jie M. Zhang ·

    SkillMOO:软件工程中智能体技能的多目标优化

    arXiv:2604.09297v3 Announce Type: replace-cross Abstract: Agent skills are increasingly used to configure coding agents for software engineering (SE) tasks, yet current practice treats them as static, hand-crafted assets, or evolved on pass rate alone. This is insufficient: a ski…

  111. arXiv cs.CL TIER_1 English(EN) · Jiaming Wei, Zekun Wu, Adriano Koshiyama, Maria Perez-Ortiz ·

    路由在最有价值的地方最难学习:网络代理表示路由的界限

    arXiv:2608.06171v1 Announce Type: new Abstract: Web agents observe a browser through text, pixels, or both, and the choice is usually fixed once for all tasks. We measure six observation modes across eight site-model combinations (cells) on VisualWebArena and WebArena and ask wha…

  112. arXiv cs.AI TIER_1 English(EN) · Han Chi, Jiaxin Qi, Yan Cui, Baisheng Lai, Jianqiang Huang ·

    匹配至关重要:命令行智能体的公平质量-效率基准

    arXiv:2606.21140v2 Announce Type: replace-cross Abstract: Rapid advances in large language models have improved the task-solving capabilities of command-line-interface (CLI)-based agents, whose CLIs determine how models invoke tools, maintain interaction history, and recover from…

  113. arXiv cs.AI TIER_1 English(EN) · Boning Li, Yu Chen, Longbo Huang ·

    AV-AIVAT:在不完美信息博弈中实现74倍更低成本的智能体评估,并获得认证的随时有效停止

    arXiv:2608.06362v1 Announce Type: cross Abstract: Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either kee…

  114. arXiv cs.AI TIER_1 English(EN) · Jialuo Chen, Lingqi Jiang, Xinhao Deng, Xiaohu Du, Jianan Ma, Yunhao Feng, Yuqi Qing, Zhihao Yuan, Linkang Du, Jingyi Wang ·

    经验化为指令:自进化智能体技能系统中的轨迹投毒

    arXiv:2608.05563v1 Announce Type: cross Abstract: Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instruction. We introduce PoisonedEvolution, a trajectory-poisoning attack on this promotion proc…

  115. arXiv cs.AI TIER_1 English(EN) · Vanessa Sochat, Daniel Milroy ·

    Agentic Science 的分层服务器架构

    arXiv:2608.05332v1 Announce Type: cross Abstract: Agentic science is transforming the landscape of computational work, extending to scientific pipelines and workload managers. The workloads require specialized hardware within and across institutions. If assessing workload needs a…

  116. arXiv cs.AI TIER_1 English(EN) · Yunjia Qi, Zehua Yin, Xintong Shi, Hao Peng, Songyuanyi Lu, Yixian Liu, Richeng Xuan, Yuhong Liu, Zhichao Hu, Xiaozhi Wang, Lei Hou, Bin Xu, Juanzi Li ·

    TRAJDEBUG:追踪错误生命周期以识别长时域Agent轨迹中的关键故障

    arXiv:2608.06346v1 Announce Type: new Abstract: LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to locate the earliest error step in a failed trajecto…

  117. Hugging Face Daily Papers TIER_1 English(EN) ·

    优化器即智能体:跨越提示、程序和机器学习工作流的驱动式搜索

    Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods. We ask a fundamentally different question: how much of this search policy can be internalized by …

  118. Hugging Face Daily Papers TIER_1 English(EN) ·

    一个端到端的智能体审计引擎

    With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains. The fast-evolving harness ecosystem has also made rigorous capability evaluation increasingly important. However, effici…

  119. Hugging Face Daily Papers TIER_1 English(EN) ·

    优化器即智能体:跨越提示、程序和机器学习工作流的驱动式搜索

    Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods. We ask a fundamentally different question: how much of this search policy can be internalized by …

  120. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Longbo Huang ·

    AV-AIVAT:在不完美信息博弈中,以认证的随时有效停止,实现74倍更便宜的智能体评估

    Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either keep paying after the result is settled or stop befor…

  121. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Yuming Jiang ·

    ASGE-RR:具有可修订预订的代理服务图嵌入,用于动态AI代理调用

    AI-agent workflows often involve remote calls to models, memory stores, and tools distributed across a network. As execution progresses, these dependency calls collectively form an agentic service graph (ASG). Unlike traditional service requests, many dependency calls are reveale…

  122. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Indrakshi Dey ·

    通过Koopman谱分析认证多智能体系统中的集体推理

    Orchestrated collectives of large language model (LLM) agents that debate and vote are an emerging form of computational intelligence: the intelligent behaviour resides in the \emph{interaction}, not in any single agent. They improve task accuracy, yet remain black boxes at the s…

  123. Hugging Face Daily Papers TIER_1 English(EN) ·

    ChainClaw:一个用于可靠链上执行的分层代理框架

    General-purpose large language model agents have achieved strong performance on tool-augmented tasks, yet they rely on assumptions break down in blockchain environments. On-chain execution is stateful, adversarial, and economically irreversible, exposing three fundamental gaps: R…

  124. Hugging Face Daily Papers TIER_1 English(EN) ·

    SkillHEX:通过假设驱动的自主探索与利用来提升智能体技能

    Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment. Real-world deployments thus require autonomous, on-demand skill evolution at test time, constrained by limited interaction budgets and…

  125. arXiv cs.AI TIER_1 English(EN) · Zhenjiang Ren, Ruiji Li, Xujing Zhang, Ziliang Pang, Shuo Ren, Jiajun Zhang ·

    EviGraph:以证据为导向的自主研究代理

    arXiv:2608.04738v1 Announce Type: new Abstract: Autonomous research agents can generate hypotheses, execute experiments, and draft manuscripts, yet their outputs often contain unsupported claims and inconsistencies between research questions, experiments, results, and conclusions…

  126. arXiv cs.AI TIER_1 English(EN) · Boxiu Li, Zimo Wen, Yijia Fan, Junxiang Lei, Sufeng Guo, Jiaao Wu, Ruize Tang, Mukai Li, Yifei Shen, Xiaoyu Chen, Wanbo Zhang, Runjing Gu, Yifei Gao, Yuheng Wu, Xuyao Huang, Zelong Zhao, Jiachen Zhang, Shibo Hu, Hangxi Guo, Yilin Chen, Yuzhe Zhang, Fan Y… ·

    Argus:一个用于长时推理的通用代理运行时

    arXiv:2608.05144v1 Announce Type: new Abstract: Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persist…

  127. arXiv cs.AI TIER_1 English(EN) · Zhuohang Jiang, Yuxin Chen, Yongsen Pan, Zheng Hu, Wenqi Fan, Qing Li, Hongyang Wang, Jun Wang, Wenwu Ou ·

    A/B Agent:工业 A/B 测试中用于策略迭代的自进化 Agent

    arXiv:2608.04625v1 Announce Type: new Abstract: Industrial recommendation strategy iteration heavily relies on large-scale A/B experimentation. Traditional tuning requires experts to repeatedly design strategies, configure experiments, analyze results, and adjust parameters, maki…

  128. arXiv cs.AI TIER_1 English(EN) · Mohsen Arjmandi ·

    大型语言模型提出,高管决定:一种自我验证的代理工具,可将长期代理中的承诺漂移与约束漂移分离

    arXiv:2608.04066v1 Announce Type: new Abstract: How do you verify a long-horizon agent when its own state and self-reports are exactly what you cannot trust? We present an agent instrument built so that verification is structural rather than post-hoc. A deterministic Executive ow…

  129. arXiv cs.LG TIER_1 English(EN) · Arnab Phani, Elias Strauss, Sebastian Schelter ·

    stratum: 面向大规模以代理为中心的机器学习工作负载的系统基础设施

    arXiv:2603.03589v3 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) transform how machine learning (ML) pipelines are developed and evaluated. LLMs enable a new type of workload, agentic pipeline search, in which autonomous or semi-autonomous…

  130. arXiv cs.AI TIER_1 English(EN) · Jingsheng Zheng, Xinyuan Fang, Jintian Zhang, Zhengke Gui, Huajun Chen, Ningyu Zhang ·

    OneDayAgent: 迈向自主代理的长远驾驭

    arXiv:2608.05013v1 Announce Type: cross Abstract: LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment, and multimodal, forcing the agent to preserve goals and constraints across many …

  131. arXiv cs.LG TIER_1 English(EN) · Jun Nie, Yonggang Zhang, Qianshu Cai, Yiu-ming Cheung, Xinmei Tian, Bo Han ·

    EvolveNet:协同式 Harness 演化以实现智能体自我改进

    arXiv:2608.04968v1 Announce Type: new Abstract: The capabilities of an LLM agent depend not only on its model but on the harness: the executable program that constructs context, invokes tools, verifies results, and recovers from failure. Recent work shows that evolving the harnes…

  132. Hugging Face Daily Papers TIER_1 English(EN) ·

    SkillTV-Bench:评测法官在技能增强代理执行方面的表现

    LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executions. For skill-augmented agents, verification additionally requires the procedural knowledge encoded…

  133. Hugging Face Daily Papers TIER_1 English(EN) ·

    DCAS:解耦CLI Agent脚手架以在脚手架内部化规划

    CLI-based software-engineering agents have matured rapidly, yet the open ecosystem has converged on a single training environment: trajectory datasets used to fine-tune open models are collected almost exclusively under OpenHands. Models fine-tuned on this data score well under O…

  134. arXiv cs.AI TIER_1 English(EN) · Leijun Zhou, Zhihao Liu, Xiang Qu, Chenxu Liu, Yifei Liu, Yanke Yu, Jingzhe Xu, Xuejun Wu, Buyue Qian, Xi Chen, Yaowei Zheng, Junhao Hu ·

    GDPevo:在真实商业任务上评估Agent的自我进化

    arXiv:2608.03764v1 Announce Type: new Abstract: Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-evolution is difficult: existing benchmarks provide limited coverage of economical…

  135. arXiv cs.LG TIER_1 (AF) · Paimon Goulart, Liang Wu, Kelly Wan, Evangelos E. Papalexakis, Liangjie Hong ·

    Field Aware Agent Skill Retrieval

    arXiv:2608.02880v1 Announce Type: cross Abstract: As lifelong learning agents accumulate lifelong growing skill banks, retrieving the correct skill becomes an increasingly important bottleneck. Most current skill retrieval methods treat each skill as one flat document by concaten…

  136. arXiv cs.LG TIER_1 Deutsch(DE) · Dong Yan, Jian Liang, Dapeng Hu, Ran He, Nicholas Jing Yuan, Qi Zhang, Tieniu Tan ·

    AgentStream:自演进LLM代理在流式任务上的表现如何?

    arXiv:2608.00155v1 Announce Type: cross Abstract: Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self-evolving ag…

  137. arXiv cs.AI TIER_1 English(EN) · Yuxuan Liu, Zhaochen Su, Yuhao Zhang, Jiahe Guo, Zhongwei Xie, Huihao Jing, Lingyun Xie, Qing Zong, Yauwai Yim, Zhixiong Zhang, Haoran Li, Yangqiu Song ·

    重新思考自演化代理技能:多轮反馈动态

    arXiv:2608.02636v1 Announce Type: cross Abstract: Self-evolving skill systems promise to improve agents by turning execution feedback into persistent skill updates without changing the underlying model. Yet it remains unclear when further evolution helps, how successful and faile…

  138. arXiv cs.AI TIER_1 English(EN) · Sebasti\'an Andr\'es Cajas Ord\'o\~nez, Agastya Munnangi, Aldo Marzullo, Felipe Ocampo Osorio, Quang Bui, Mohammad Shahin, Armaan Grewal, Emmanuel Paul Kwesiga, Anqi Peter Li, Josephine Nanyonjo, Aaditya Panchal, Arshnoor Bhutani, Nikhil Jaiswal, Milit S… ·

    Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems

    arXiv:2608.03744v1 Announce Type: new Abstract: Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore. Acro…

  139. arXiv cs.AI TIER_1 English(EN) · Prince Zizhuang Wang, Aojie Yuan, Haiyue Zhang, Xiyang Hu, Yue Zhao, Shuli Jiang ·

    WeClawArena:面向以人为中心的代理网络中跨用户代理协作与安全的审计沙箱和基准测试

    arXiv:2608.03499v1 Announce Type: new Abstract: Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and communicates …

  140. arXiv cs.AI TIER_1 English(EN) · Can Wang, Haoran Chen, Li Yu, Ding Hao, Bohai Zhao, Zhaoyang Liu, Zhiying Tu ·

    通过经验驱动的自适应指导实现代理的鲁棒工具使用

    arXiv:2608.03403v1 Announce Type: new Abstract: The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools play a central role as the primary interface through which agents interact with external envir…

  141. arXiv cs.AI TIER_1 English(EN) · Chenyu Wang, Yunbo Lyu, Junda He, Zhou Yang, Chenxing Zhong, Yaniv Harel, David Lo ·

    快速失败,智能重启:SWE Agentic任务的早期失败预测与重启

    arXiv:2608.03222v1 Announce Type: cross Abstract: Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context accumulates. Failed runs tend to be longer and exhibit redundant exploration or looping, sugge…

  142. arXiv cs.AI TIER_1 English(EN) · Ankur Sharma, Deep Shah ·

    Agent Operating System (AOS):分布式Agentic系统的参考操作系统架构

    arXiv:2608.03214v1 Announce Type: new Abstract: Large language models have transformed artificial intelligence from isolated prediction services into components of long-running, distributed systems that reason, invoke tools, retrieve external state, delegate tasks, and act on beh…

  143. arXiv cs.AI TIER_1 English(EN) · Yu-Tung Liu, Cunxi Yu ·

    VeriTrace:类人时间探索完成代理行动空间

    arXiv:2608.02878v1 Announce Type: new Abstract: Large language models have shown promise for automated Verilog RTL generation, yet state-of-the-art multi-agent systems plateau at ~95% accuracy on standard benchmarks. We trace this ceiling to an incomplete debugging action space: …

  144. arXiv cs.AI TIER_1 English(EN) · Qi Liu, Ruochen Hao, Can Li, Wanjing Ma ·

    OR-Agent:连接进化搜索与结构化研究以实现自动化算法发现

    arXiv:2602.13769v3 Announce Type: replace Abstract: Automating heuristic design in complex, experiment-driven domains requires more than iterative mutation of solution algorithms. Current LLM-based evolutionary methods often rely on stochastic mutation loops that lack long-term s…

  145. arXiv cs.AI TIER_1 English(EN) · Salma El Yadouni (EPFL), Guanyi Li (Binome Technologies) ·

    TraceCompiler:技能引导的 LLM Agent 轨迹挖掘与编译,形成高度确定性工作流

    arXiv:2608.02680v1 Announce Type: cross Abstract: Tool-using language-model agents repeatedly rediscover procedures they have already executed, producing traces that mix reusable structure with retries, exploration, accidental ordering, and repeated lookups. We present TraceCompi…

  146. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Ningyu Zhang ·

    OneDayAgent:迈向自主代理的长视界驾驭

    LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment, and multimodal, forcing the agent to preserve goals and constraints across many steps while navigating heterogeneous tools and att…

  147. Hugging Face Daily Papers TIER_1 English(EN) ·

    WeClawArena:面向以人为中心的智能体网络中跨用户智能体协作与安全的审计沙箱和基准测试

    Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and communicates with other agents through social and task relati…

  148. Hugging Face Daily Papers TIER_1 English(EN) ·

    通过经验驱动的自适应指导实现代理的鲁棒工具使用

    The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools play a central role as the primary interface through which agents interact with external environments, yet existing methods rarely focus on en…

  149. Hugging Face Daily Papers TIER_1 English(EN) ·

    快速失败,智能重启:SWE Agentic任务的早期失败预测与重启

    Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context accumulates. Failed runs tend to be longer and exhibit redundant exploration or looping, suggesting that some failures may be detectable before …

  150. arXiv cs.LG TIER_1 English(EN) · Natan Vidra, Alina Kapanova, Arun Kanhai, Spurthi Setty ·

    为 Agentic 工作流学习组合式元路由:一个可执行的基准

    arXiv:2608.00106v1 Announce Type: new Abstract: Agentic systems must decide not only what answer to produce, but which reasoning and execution operations should precede it. A controller may answer directly, decompose a request, retrieve evidence, execute code, delegate to a speci…

  151. arXiv cs.CL TIER_1 English(EN) · Hanwen Xing, Pengyun Wang, BingXu Meng, Kumail Alhamoud, Xiang Li, Jicheng Wang, Xin Yu, Xinyang Han, Xiaomin Li, Philip Torr, Yuexing Hao ·

    CurveShift:智能体进展是否可量化?区分层级与形态

    arXiv:2608.00355v1 Announce Type: new Abstract: Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summaries capture the overall performance, but they do no…

  152. arXiv cs.CL TIER_1 English(EN) · Yi Nian, Haosen Cao, Shenzhe Zhu, Henry Peng Zou, Qingqing Luan, Yudi Zhang, Yue Zhao ·

    当仅有最终文本得以保留时:用于多智能体审计的隐式执行追踪

    arXiv:2603.17445v5 Announce Type: replace-cross Abstract: When a multi-agent system produces an incorrect or harmful answer, who is accountable if execution logs and agent identifiers are unavailable? In practice, generated content is often detached from its execution environment…

  153. arXiv cs.CL TIER_1 English(EN) · Yunhao Chen, Xin Wang, Yixu Wang, Yi Liu, Jie Li, Yan Teng, Xingjun Ma, Xia Hu, Yu-Gang Jiang ·

    OpenART:通过开放式环境演化实现智能体红队测试的规模化

    arXiv:2608.00677v1 Announce Type: new Abstract: AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that is repeatedl…

  154. arXiv cs.CL TIER_1 English(EN) · Qiyong Zhong, Mao Zheng, Mingyang Song, Xin Lin, Jie Sun, Houcheng Jiang, Xiang Wang, Junfeng Fang ·

    SOD:面向小型语言模型代理的步进式在线策略蒸馏

    arXiv:2605.07725v2 Announce Type: replace Abstract: Tool-integrated reasoning (TIR) is difficult to scale to small language models due to instability in long-horizon tool interactions and limited model capacity. While reinforcement learning methods like group relative policy opti…

  155. arXiv cs.LG TIER_1 English(EN) · Natan Vidra, Alina Kapanova, Arun Kanhai, Spurthi Setty ·

    MetaRoute-Bench:评估用于代理工作流路由的元决策策略

    arXiv:2608.00107v1 Announce Type: new Abstract: Agentic systems must repeatedly decide whether to answer directly, decompose a task, invoke a tool, execute code, delegate to a specialist, verify an intermediate result, or recover from failure. These meta-decisions affect not only…

  156. arXiv cs.CL TIER_1 English(EN) · Qi Liu, Jiaxin Mao, Fengbin Zhu, Tat-Seng Chua ·

    诊断长时域搜索代理中的搜索行为和失败模式

    arXiv:2608.01913v1 Announce Type: cross Abstract: Deep search agents answer difficult information-seeking questions by iteratively issuing search queries to gather supporting evidence, but it remains unclear whether and how greater search effort leads to better answers. We study …

  157. arXiv cs.CL TIER_1 English(EN) · Luan Zhang, Ruochen Zhou, Dandan Song, Zhengyu Chen, Yuhang Tian, Jun Yang, Huipeng Ma, Chenhao Li, Guangyuan Feng, Xudong Li, Yizhou Jin, Yan Xu ·

    HarnessCompass:引导自动线束演进,实现通用有效的代理线束

    arXiv:2608.01918v1 Announce Type: cross Abstract: Harness design plays a critical role in agent performance by shaping how large language models (LLMs) perceive, reason over, and act within executable environments. Recent work has proposed automatic harness evolution, which itera…

  158. arXiv cs.CL TIER_1 English(EN) · Yinghan Hou, Zongyou Yang ·

    压缩下的控制:工具使用代理的可靠性前沿

    arXiv:2608.01056v1 Announce Type: cross Abstract: Tool-using language-model agents are governed not only by task prompts but also by persistent system-side instructions that specify tools, arguments, policies, execution protocols, and recovery. Compressing these agent control con…

  159. arXiv cs.LG TIER_1 English(EN) · Tezan Sahu, Himani Arora ·

    代理在19:05能看到什么?从真实研究中生成时间企业场景并回放以评估代理

    arXiv:2608.01042v1 Announce Type: cross Abstract: Enterprise AI agents act across many apps whose data changes continuously, so an answer is correct only relative to what data existed and who could see it at the moment it was asked. Offline evaluation today grades against a singl…

  160. arXiv cs.LG TIER_1 English(EN) · Shuaijun Liu, Feiyang You, Xingwei Chen, Ningxin Su ·

    当重新规划成为瓶颈:具身智能体的预算式重新规划

    arXiv:2608.01428v1 Announce Type: cross Abstract: Embodied agents replan frequently to recover from execution drift, partial observability, and coordination hazards, but each LLM-based replanning call can consume an accumulated textual context that grows over time and across agen…

  161. arXiv cs.LG TIER_1 English(EN) · Jingxi Wei ·

    可自行分割的轨迹:代理声明的边界作为训练单元

    arXiv:2608.02302v1 Announce Type: cross Abstract: Long-horizon coding-agent trajectories are poorly matched to the credit units available to train on: a single action has no stable value, an episode label merges productive exploration with abandoned directions, and a fixed window…

  162. arXiv cs.CL TIER_1 English(EN) · Vernon Toh, Navonil Majumder, Zhengyuan Liu, Nancy F. Chen, Soujanya Poria ·

    ScrambleToolBench:即使自身地图指向下一步,智能体也能穷尽搜索

    arXiv:2608.02358v1 Announce Type: new Abstract: To operate robustly in open-world environments, autonomous agents should be able to infer the behavior of unfamiliar systems through interaction alone, even in the absence of documentation. However, existing tool-use benchmarks expo…

  163. arXiv cs.LG TIER_1 English(EN) · Cong Wan, Zeyu Guo, Zijian Cai, Jiangyang Li, SongLin Dong, Lin Peng, Xiangyang Luo, Zhiheng Ma, Yihong Gong ·

    DataClaw0:从原始流中进行智能体定制化多模态数据处理

    arXiv:2606.21337v2 Announce Type: replace Abstract: Raw multimodal streams are abundant but noisy, redundant, and unaligned with any particular training objective. Turning them into supervision today means either brittle heuristics or repeatedly querying a proprietary vision-lang…

  164. arXiv cs.CL TIER_1 English(EN) · Donghyeok Koh, Gyuwan Kim, Jinyeong Bak, Seung-Hoon Na, Tao Yang, Haneol Jang, Cheoneum Park ·

    面向Agent工作流的全局优化与推理时区域嫁接

    arXiv:2608.02353v1 Announce Type: new Abstract: Recent advances in agentic workflow optimization automate workflow design through task-specific workflow search or input-conditioned architecture selection. However, they determine the workflow before execution and cannot adapt fail…

  165. arXiv cs.LG TIER_1 English(EN) · Hao Mark Chen, Jinnan Guo, Wayne Luk, Hongxiang Fan ·

    AOSpec:低延迟代理服务的动作与观察协同推测

    arXiv:2608.00881v1 Announce Type: new Abstract: Large language model agents increasingly act through stateful tools, yet model generation and environment execution remain serialized at every step. As decoding accelerates, tool execution becomes a growing bottleneck. Existing acti…

  166. arXiv cs.CL TIER_1 English(EN) · Yongfeng Huang, Yuren Lai, Ruiying Chen, Haoyu Huang, Mingming Zhao, James Cheng ·

    ACE-GraphRAG:用于分层 GraphRAG 的代理上下文工程

    arXiv:2608.01269v1 Announce Type: new Abstract: Hierarchical Graph Retrieval-Augmented Generation (GraphRAG) organizes corpus knowledge at multiple levels of granularity, yet fixed context construction may fail to translate these multi-resolution representations into a context su…

  167. Hugging Face Daily Papers TIER_1 English(EN) ·

    GDPevo: 在真实商业任务上评估Agent的自我进化

    Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-evolution is difficult: existing benchmarks provide limited coverage of economically valuable task domains, do not always design t…

  168. Hugging Face Daily Papers TIER_1 English(EN) ·

    WeClawArena:面向以人为中心的智能体网络中跨用户智能体协作与安全的审计沙箱和基准测试

    WeClawArena is an auditable benchmark and sandbox for evaluating multi-party agent collaboration across personal workspaces, measuring both task utility and security attack success.

  169. Hugging Face Daily Papers TIER_1 Svenska(SV) ·

    SkillJack:自进化智能体中的持久性技能后门

    Self-evolving agents increasingly convert interaction histories into reusable skills that persist beyond individual tasks. While prior work studies memory and retrieval poisoning, such attacks only affect agents when poisoned records are retrieved as context. We uncover a new and…

  170. Hugging Face Daily Papers TIER_1 English(EN) ·

    OneDayAgent:迈向自主代理的长视界驾驭

    LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment, and multimodal, forcing the agent to preserve goals and constraints across many steps while navigating heterogeneous tools and att…

  171. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Yiwen Guo ·

    ADIAS:交互式Agentic系统的自动化设计

    Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization. Existing methods are largely candidate-centric: cross-round experience is organized around candidate agents, which leaves the repair progress implicit. This causes …

  172. arXiv cs.IR (Information Retrieval) TIER_1 (AF) · Liangjie Hong ·

    字段感知Agent技能检索

    As lifelong learning agents accumulate lifelong growing skill banks, retrieving the correct skill becomes an increasingly important bottleneck. Most current skill retrieval methods treat each skill as one flat document by concatenating fields such as the name, description, and bo…

  173. arXiv cs.IR (Information Retrieval) TIER_1 (AF) · Liangjie Hong ·

    字段感知Agent技能检索

    As lifelong learning agents accumulate lifelong growing skill banks, retrieving the correct skill becomes an increasingly important bottleneck. Most current skill retrieval methods treat each skill as one flat document by concatenating fields such as the name, description, and bo…

  174. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Tat-Seng Chua ·

    长时域搜索代理的行为诊断与失效模式分析

    Deep search agents answer difficult information-seeking questions by iteratively issuing search queries to gather supporting evidence, but it remains unclear whether and how greater search effort leads to better answers. We study these questions through a trajectory-level diagnos…

  175. arXiv cs.AI TIER_1 English(EN) · Blaise Delattre, Cong Wang, Yang Cao ·

    CAGE:面向工具使用代理的类型化返回不确定性下的认证授权

    arXiv:2607.29190v1 Announce Type: new Abstract: Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates generally authorize the observed return and action, leaving the decision unprotected …

  176. arXiv cs.AI TIER_1 English(EN) · Yucheng Xu, Keyi Zhang, Yuyang Yu, Min Zhang, Shiyuan Meng, Pei Chu, Zhongying Tu ·

    为回合制Agentic RL扩展科学发现环境

    arXiv:2607.28990v1 Announce Type: new Abstract: Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an execution environment and produces a statistical claim. Long-horizon scientific analysis remai…

  177. arXiv cs.AI TIER_1 English(EN) · Roy Zhao (Paul G. Allen School of Computer Science & Engineering, University of Washington), Zhenyu Zhao (Independent Researcher) ·

    代码即本体:用于递归演化和下降的代理软件本体

    arXiv:2607.28691v1 Announce Type: cross Abstract: Personalized AI agents are often configurable without giving users control over the artifacts that determine their future behavior. We present OurArk, an architecture for persistent personal agents centered on an agent-owned softw…

  178. arXiv cs.AI TIER_1 English(EN) · Youting Wang, Xiao Han, Dingyan Shang, Yuan Tang, Bowen Liu ·

    安全性,还是仅仅是能力?对Agent-Safety基准的有效性审计

    arXiv:2607.28685v1 Announce Type: new Abstract: Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each u…

  179. arXiv cs.AI TIER_1 English(EN) · Yingwei Zheng, Cong Li, Shaohua Li, Yuqun Zhang, Zhendong Su ·

    Agentic Harness for Real-World Compilers

    arXiv:2603.20075v2 Announce Type: replace-cross Abstract: Compilers are critical to modern computing, yet fixing compiler bugs is difficult. While recent large language model (LLM) advancements enable automated bug repair, compiler bugs pose unique challenges due to their complex…

  180. arXiv cs.AI TIER_1 English(EN) · Michael Fu, Qiyue Mei, Patanamon Thongtanunam, Kla Tantithamthavorn ·

    AgenticRepair: 面向Agentic漏洞修复的多方面程序上下文工程

    arXiv:2607.29422v1 Announce Type: cross Abstract: Automated vulnerability repair aims to reduce the time and effort required to patch security flaws from a vulnerability triage report. Recent agentic AI approaches have shown promising results in automated program repair. However,…

  181. arXiv cs.AI TIER_1 English(EN) · Harsh Raj, Vipul Gupta, Anas Mahmoud, Razvan-Gabriel Dumitru, Darvin Yi, Aakash Sabharwal, Yunzhong He ·

    模型还是束缚?一种以交互为中心的分类法用于定位智能体故障

    arXiv:2607.28802v1 Announce Type: new Abstract: Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates a repair-assignment problem: the same visible failur…

  182. Hugging Face Daily Papers TIER_1 English(EN) ·

    ScrambleToolBench:即使自身地图指向下一步,智能体也能穷尽搜索

    To operate robustly in open-world environments, autonomous agents should be able to infer the behavior of unfamiliar systems through interaction alone, even in the absence of documentation. However, existing tool-use benchmarks expose semantic tool schemas in static environments,…

  183. Hugging Face Daily Papers TIER_1 English(EN) ·

    LongHorizon-Harness:推进长时域智能体以应对现实世界任务

    Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interdependent steps. However, existing agent harnesses maintain task execution, task state, and completion assessment within a growing…

  184. Hugging Face Daily Papers TIER_1 English(EN) ·

    压缩下的控制:工具使用代理的可靠性前沿

    Tool-using language-model agents are governed not only by task prompts but also by persistent system-side instructions that specify tools, arguments, policies, execution protocols, and recovery. Compressing these agent control contexts (ACCs) can reduce input cost and context use…

  185. arXiv cs.AI TIER_1 English(EN) · Lang Cao, Yuhao Shen, Tianyang Luo, Simo Du, Hao Peng, Yue Guo ·

    GuideSkill:为基于指南的临床推理演进可执行的LLM代理技能

    arXiv:2607.26160v1 Announce Type: new Abstract: Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather than execute its rules. We introduce GuideSkill, an external reasoning layer that…

  186. arXiv cs.AI TIER_1 English(EN) · David Kaleko, Sergey Ivanov, Md Mofijul Islam ·

    IDP AutoOpt:驱动文档处理管道配置的智能体优化

    arXiv:2607.26075v1 Announce Type: cross Abstract: We present IDP AutoOpt, an autonomous LLM agent that discovers high-performing configurations for intelligent document processing (IDP) pipelines. Tuning IDP prompts, models, OCR settings, and schemas jointly currently costs domai…

  187. arXiv cs.AI TIER_1 English(EN) · Junhao Qiu, Zidong Wang, Yansong Sun, Zhitong Ma, Ping Guo, Qingfu Zhang ·

    AgenticCANN: 基于知识增强的智能体演化实现自动化Ascend C算子生成

    arXiv:2607.26661v1 Announce Type: new Abstract: Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise.While large language models (LLMs) have shown promise in automated CUDA kernel generation, the fu…

  188. arXiv cs.LG TIER_1 English(EN) · Yash Pandya, Sahil Gupta, Sarthak Harne, Archana Yadav, Kavyansh Chourasia, Hussein Mozannar, Vibhav Vineet, Sara Abdali, Corby Rosset, Yash Lara, Ahmed Awadallah, Ece Kamar, Akshay Nambi ·

    Echoverse:大规模训练计算机使用代理的深度、演进式环境

    arXiv:2607.28074v1 Announce Type: cross Abstract: Computer-use agents learn from what their actions change, so training one needs applications it can act on, break and reset. The applications that matter most are login-gated and stateful, so synthetic environments stand in for th…

  189. arXiv cs.LG TIER_1 English(EN) · Zihan Dong, Rui Qian, Qishi Zhan, Dongshen Peng, Kaixin Li, Yu Li ·

    为何GUI智能体正确但迟到?解码决策时间关键路径,并用预编译策略树进行测试

    arXiv:2607.28399v1 Announce Type: new Abstract: Computer-use agents often fail on transient GUI events because they produce the correct action only after the relevant window has already closed. We identify the main cause as expensive autoregressive decoding on the decision-time c…

  190. arXiv cs.LG TIER_1 English(EN) · Xingjian Wu, Xuhang Zhu, Xingchen Liu, Junlin Liu, Jianing Wang, Linsen Guo, Xiaoyu Li, Xuezhi Cao, Xunliang Cai ·

    ClawTrack:迈向量化评估和改进真实世界自主代理的轨迹

    arXiv:2607.28037v1 Announce Type: new Abstract: As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outcomes, unable to distinguish reliable reasoning from lucky success or attribute f…

  191. arXiv cs.CL TIER_1 English(EN) · Yilong Lai, Yipin Yang, Ting Liang, Jialong Wu, Zhenglin Wang, Jianguo Lin, Keping Yang ·

    CRMWeaver:通过 Agentic RL 和共享记忆构建强大的业务代理

    arXiv:2510.25333v2 Announce Type: replace Abstract: Recent years have witnessed the rapid development of LLM-based agents, which shed light on using language agents to solve complex real-world problems. A prominent application lies in business agents, which interact with database…

  192. Hugging Face Daily Papers TIER_1 Deutsch(DE) ·

    AgentStream:自演进LLM代理在流式任务上的表现如何?

    Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self-evolving agents in realistic streaming settings, where agents…

  193. Hugging Face Daily Papers TIER_1 English(EN) ·

    为什么GUI智能体正确但迟到?在决策时间关键路径上解码,并使用预编译策略树进行测试

    Computer-use agents often fail on transient GUI events because they produce the correct action only after the relevant window has already closed. We identify the main cause as expensive autoregressive decoding on the decision-time critical path. We propose Adaptive Anticipatory P…

  194. Hugging Face Daily Papers TIER_1 English(EN) ·

    Echoverse:大规模训练计算机使用代理的深度、演进式环境

    Computer-use agents learn from what their actions change, so training one needs applications it can act on, break and reset. The applications that matter most are login-gated and stateful, so synthetic environments stand in for them. Recent pipelines generate such environments in…

  195. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Zhou He ·

    VISA:面向机器可复现性的基于代理的模拟模型的结构化描述协议

    Agent-based models (ABMs) are difficult to reproduce: their behavior is spread across prose narratives, platform-specific code, and implicit assumptions, so that two readers routinely reconstruct different models from the same documentation. We present VISA, a structured, symbol-…

  196. Hugging Face Daily Papers TIER_1 English(EN) ·

    RoboBRIDGE:一个用于将策略桥接到健壮的真实世界机器人代理的模块化框架

    Vision-Language-Action (VLA) models have attracted growing interest as a scalable approach to robotic manipulation. While these models are effective action predictors, deploying them as robotic agents exposes critical gaps: no mechanism for failure recovery, inconsistent executio…

  197. arXiv cs.CL TIER_1 English(EN) · Yihao Chen, Shi Chang, Feng Lin, Khaled Chawa, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan ·

    SpecFirst:行为规范提取作为从头开始的基于代理的程序合成的一等步骤

    arXiv:2607.27167v1 Announce Type: cross Abstract: LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this gap: give…

  198. arXiv cs.CL TIER_1 English(EN) · Xuan Zhao, Jiwoong Sohn, Qinyue Zheng, Michael Moor ·

    AgentGUI:一个用于观察和引导长期运行的AI代理的界面

    arXiv:2607.26300v1 Announce Type: new Abstract: AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight is systematically lagging behind due to limited human-centered interfacing. Aiming to address…

  199. arXiv cs.CL TIER_1 English(EN) · Ruoyu Wang, Heng Zhao, Renjie Wu, Mengnan Zhao, Zhixuan Chu, Wanyu Lin, Tianhang Zheng ·

    AgentSnare:学习延迟、转移和化解自主渗透代理

    arXiv:2607.26998v1 Announce Type: cross Abstract: Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools. This dependence allows defenders to inject deceptive observations that c…

  200. arXiv cs.CL TIER_1 English(EN) · Lingyang Zeng, Guangze Chen, Kaichen Yu, Zhicheng Pan, Siyang Weng, Zirui Hu, Xiangyun Du, Hailin He, Rong Zhang, Chengcheng Yang, Kai Huang, Xuan Zhou ·

    Setoka:用于异构数据上个性化代理的层级用户理解基准

    arXiv:2607.27056v1 Announce Type: cross Abstract: Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also infer…

  201. arXiv cs.CL TIER_1 English(EN) · Jingbo Zhou, Yusai Zhao, Qi Bao, Jingjia Cao, Zhenghai Chen, Chang Gao, Kaiqi Guo, Muxin Guo, Mingxuan Li, Xinjiang Lu, Yanru Ma, Yixiong Xiao, Zenghui Zhang, Le Zhang, Hua Wu ·

    OmegaUse-OfficeVal:在具有经济基础的长周期办公套件任务上对 LLM Agents 进行基准测试

    arXiv:2607.27155v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonab…

  202. arXiv cs.LG TIER_1 English(EN) · Jingbo Cui, Jitao Zhao, Di Jin, Dongxiao He ·

    AgentGFM:一种具有节点-Agent信息流控制的图基础模型

    arXiv:2607.26533v1 Announce Type: new Abstract: Graph Foundation Models (GFMs) aim to learn transferable knowledge from multi-domain graphs and adapt to unseen scenarios. As a fundamental source of relational semantics in graphs, the transferability of topological patterns has lo…

  203. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Bowen Liu ·

    安全性,还是仅仅是能力?对Agent-Safety基准的有效性审计

    Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and author-prov…

  204. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Truong-Son Hy ·

    通过功能、证据和验证评估 Agentic 生物信息学

    Large language model agents increasingly plan, execute, and interpret biological analyses, yet fluent responses, successful tool calls, and benchmark performance alone do not establish scientific credibility. Existing reviews primarily organize biological agents by application, a…

  205. Hugging Face Daily Papers TIER_1 English(EN) ·

    模型还是束缚?以交互为中心的代理失败定位分类法

    Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates a repair-assignment problem: the same visible failure may call for model post-training, harness engi…

  206. Hugging Face Daily Papers TIER_1 English(EN) ·

    Qwen-UI-Agent技术报告:迈向下一代以真实世界为中心的通用GUI智能体

    GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI execution, compl…

  207. Hugging Face Daily Papers TIER_1 English(EN) ·

    Echoverse:大规模训练计算机使用代理的深度、演进式环境

    Computer-use agents learn from what their actions change, so training one needs applications it can act on, break and reset. The applications that matter most are login-gated and stateful, so synthetic environments stand in for them. Recent pipelines generate such environments in…

  208. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Xiaowei Huang ·

    技能使用还是技能表演?评估技能增强型语言代理中的推理后台

    Reusable skills are becoming a standard interface for extending language agents with task procedures. Yet evaluators usually infer skill use from visible reasoning or the agent's own attribution. These signals show what the agent appears to use, not whether the skill changed its …

  209. Hugging Face Daily Papers TIER_1 English(EN) ·

    SpecFirst:行为规范提取作为从头开始的基于代理的程序合成中的一等步骤

    LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation and an execu…

  210. Hugging Face Daily Papers TIER_1 English(EN) ·

    OmegaUse-OfficeVal:在具有经济基础的长周期办公套件任务上对 LLM Agent 进行基准测试

    Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchm…

  211. arXiv cs.AI TIER_1 English(EN) · Jingbo Zhang, Haoxiang Sun, Wenbo Wang, Wenbo Zhang ·

    ContractHIL-HLS: 面向HLS设计的合同对齐多智能体工作流及硬件在环反馈

    arXiv:2607.25283v1 Announce Type: new Abstract: This paper presents ContractHIL-HLS, a contract-aligned multi-agent workflow for practical high-level synthesis (HLS) engineering. The workflow makes three contributions. First, it introduces a structured contract as the semantic-al…

  212. arXiv cs.AI TIER_1 English(EN) · Azizul Zahid, Subrata Biswas, Bashima Islam, Sai Swaminathan ·

    ProcAgent:边缘端带有人机协作的程序化任务指导的代理框架

    arXiv:2607.24770v1 Announce Type: new Abstract: Procedural tasks such as furniture assembly and home repair impose substantial cognitive demands because users must interpret instructions, track task progress, reason about spatial state, and recover from errors while performing ph…

  213. arXiv cs.AI TIER_1 English(EN) · Liudas Panavas, Sebastian Minus, Bradley Monton, Derek Ray, Suhaas Garre, Sushant Mehta, Edwin Chen ·

    HANDBOOK.md:长上下文代理指令遵循的基准测试

    arXiv:2607.25398v1 Announce Type: new Abstract: Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows. Existing be…

  214. arXiv cs.AI TIER_1 English(EN) · Jincheng Wang, Min Zheng, Tao Wei ·

    COVENANT:面向对齐代理执行的自然语言工作流编译

    arXiv:2607.25400v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly entrusted with natural-language workflow instructions (e.g., retail-payment policies) that specify not only what outcome to achieve, but also which steps, branches, and tool interac…

  215. arXiv cs.AI TIER_1 English(EN) · Yu Hao, Jinxuan Cai, Qi Zhang, Yawen Li, Zhiqiang Zhang, Chuan Shi, Cheng Yang ·

    HiSkill:赋能大语言模型代理的层级技能图谱

    arXiv:2607.25853v1 Announce Type: new Abstract: Skills have become an important abstraction for enabling large language model (LLM) agents to reuse past experience in long-horizon interactive tasks. However, existing trajectory-to-skill methods often produce flat collections of h…

  216. arXiv cs.AI TIER_1 English(EN) · Hyundoo Park, Byungho Choi ·

    智能体循环何时会将停滞误认为进步?长期自主LLM智能体循环中的自我评估偏见与外部验证

    arXiv:2607.25152v1 Announce Type: new Abstract: Long-running autonomous agents plan, act, and judge their own completion without human intervention. When an agent grades its own work, self-evaluation bias takes hold: plausible changes are accepted as progress while real-world out…

  217. arXiv cs.AI TIER_1 English(EN) · Huan Chen, Xiang Song, Jian Jin, Pan Ren, Liang-Jie Zhang ·

    迈向多智能体大语言模型系统的组织科学:解耦“谁”、“如何”和“哪种算法”

    arXiv:2607.25446v1 Announce Type: new Abstract: Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), how members align (coordination), and which algorithm fuses their work (collaborat…

  218. arXiv cs.AI TIER_1 English(EN) · Jianing Geng, Ruiqi He, Zekun Fei, Biao Yi, Ruijie Wang, Zheli Liu, Xia Hu, Xuansheng Wu, Qingkai Zeng ·

    Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories

    arXiv:2607.25560v1 Announce Type: new Abstract: Agent skills package reusable procedures that improve downstream performance. Their lightweight, portable form enables marketplace monetization and private deployment behind cloud-hosted agent interfaces, giving providers incentives…

  219. arXiv cs.AI TIER_1 English(EN) · Gosia Steinder, Hubertus Franke ·

    迈向Agent操作系统——来自经典和云操作系统的经验教训

    arXiv:2607.25076v1 Announce Type: new Abstract: Every major wave of platform software follows the same arc: an initial period of experimentation with competing frameworks and ad-hoc implementations, followed by the articulation of a small set of stable abstractions with well-defi…

  220. arXiv cs.AI TIER_1 Norsk(NO) · Alireza Saleh Abadi, Leen-Kiat Soh, Daniel Alan Redder, Adam Eck, Prashant Doshi ·

    PLATO: Pointer Learner for Agent and Task Openness

    arXiv:2607.25082v1 Announce Type: new Abstract: Open agent systems (OASYS) are increasingly prevalent in real-world domains where the sets of agents and tasks change unpredictably over time. Such openness, including agent openness (AO) and task openness (TO), poses a fundamental …

  221. arXiv cs.AI TIER_1 English(EN) · Rushi Qiang, Changhao Li, Haotian Sun, Yuchen Zhuang, Chao Zhang, Bo Dai ·

    Matryoshka Agent:为长周期机器学习工程展开子代理

    arXiv:2607.25090v1 Announce Type: new Abstract: Machine learning engineering (MLE) tasks require long-horizon decision making over iterative solution debugging and refinement, under expensive and feedback-driven environment interactions. Developing and training a monolithic agent…

  222. arXiv cs.AI TIER_1 English(EN) · Zhenzhen Ren, Jiyan He, Xinpeng Zhang, Zhenxing Qian, Ke Han, Shuxin Zheng, GuoBiao Li, Xiaoqing Zhang ·

    OrchBench:通过确定性模拟隔离评估多智能体编排计划

    arXiv:2607.25656v1 Announce Type: new Abstract: Complex tasks often decompose into parallelizable yet interdependent subtasks, making orchestration critical to the performance of multi-agent systems (MAS). Existing evaluations typically rely on end-to-end execution, which conflat…

  223. arXiv cs.AI TIER_1 English(EN) · Jiabao Ji, Yujian Liu, Li An, Rohit Jain, Gungor Polatkan, Siyu Zhu, Shiyu Chang ·

    边推理边推测:通过联合智能体-推测器强化学习教会智能体预测其下一次工具调用

    arXiv:2607.25816v1 Announce Type: new Abstract: Large language model agents often spend substantial wall-clock time waiting for tool call results. Tool-call speculation can hide this latency by predicting and pre-executing an agent's next tool call if the prediction matches the a…

  224. arXiv cs.AI TIER_1 English(EN) · Ravi Kant Sharma, Ashutosh Uttam, Ajay Kumar ·

    迈向自主网络中跨供应商代理工具信任管理的标准化

    arXiv:2607.25914v1 Announce Type: new Abstract: Autonomous Network Levels 4-5 require AI agents to invoke tools across vendor boundaries without human oversight, yet existing management standards lack a standardized mechanism for cross-vendor trust visibility. When a tool from Ve…

  225. arXiv cs.CL TIER_1 English(EN) · Hao Liang, Meiyi Qiang, Sizhe Qiu, Linzhuang Sun, Wentao Zhang ·

    WorkSurface-Bench:对企业多表面知识路由代理进行基准测试

    arXiv:2607.25765v1 Announce Type: new Abstract: Enterprise agents often need to integrate heterogeneous knowledge sources: documents for narrative facts, tables for computation, and dependency graphs for file relationships. Existing benchmarks typically evaluate retrieval or tool…

  226. arXiv cs.AI TIER_1 English(EN) · Yu Wang, Yi-Kai Zhang, Wentao Shi, Ziang Ye, Yuchun Miao, Yueqing Sun, Qi Gu, Xunliang Cai, Lan-Zhe Guo, Han-Jia Ye, Fuli Feng ·

    CAST:将游戏求解器作为回合制教师来训练LLM代理

    arXiv:2607.25308v1 Announce Type: cross Abstract: Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse final rewards that reveal little …

  227. arXiv cs.AI TIER_1 English(EN) · Giuseppe Destefanis ·

    Authoring Agent Skills: A Software-Engineering Approach

    arXiv:2607.25032v1 Announce Type: cross Abstract: Agent Skills are an emerging way to extend large language model agents with reusable procedural knowledge that the agent loads on demand. Anthropic introduced Agent Skills and published the format as an open specification supporte…

  228. arXiv cs.AI TIER_1 English(EN) · Mishca de Costa, Muhammad Saleh Anwar, Dave Mercier, Issam Hammad ·

    从朴素RAG到深度Agentic检索:面向监管合规的演进式上下文工程管道

    arXiv:2607.24791v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) is the dominant paradigm for applying large language models (LLMs) to enterprise document corpora, yet naive implementations encounter hard limits as corpus scale and query complexity grow. Thi…

  229. arXiv cs.AI TIER_1 English(EN) · Stefan Krsteski, Charlotte Meyer, Guillaume Allegre, Tony O'Halloran, Alexandre Sallinen ·

    Messier:用于跨基准代理评估的高分辨率语料库

    arXiv:2607.25891v1 Announce Type: new Abstract: Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts focus on narrow settings, remain limited in scale, or require costly reruns, leaving much of…

  230. Hugging Face Daily Papers TIER_1 English(EN) ·

    SpecFirst:行为规范提取作为从头开始基于代理的程序合成的一等步骤

    LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation and an execu…

  231. Hugging Face Daily Papers TIER_1 English(EN) ·

    OmegaUse-OfficeVal:在具有经济基础的长周期办公套件任务上对LLM代理进行基准测试

    Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchm…

  232. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Yiming Qiu ·

    迈向面向代理云管理的系统基础

    Agentic cloud management is emerging as a practice to automate laborious operations, minimize toil, and improve responsiveness. Despite the rapid development of autonomous management agents, we argue that the fundamental missing piece is a systems foundation to enable safe, effec…

  233. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Eric Tan ·

    ARCHER: Agentic Rule and Compliance Harness for Executable Regulations

    Verifying building compliance requires validating thousands of rules against large Building Information Modeling (BIM) designs, which is laborious, capital-intensive, and unscalable. Existing Automated Compliance Checkers (ACCs) are often difficult to generalize across different …

  234. arXiv cs.AI TIER_1 English(EN) · Erik Imgrund, Anna Wimbauer, Klim Kireev, Konrad Rieck ·

    虚假的先知:论Agentic系统中世界模型的安全性

    arXiv:2607.23147v1 Announce Type: cross Abstract: Large language models now power autonomous agents capable of complex, multi-step tasks in different environments. Accurate and reliable execution of these tasks requires the agent to predict the results of its actions. Recent rese…

  235. arXiv cs.LG TIER_1 English(EN) · Daniel Wang, Andrew Xu ·

    AlloBench:衡量大型语言模型代理中的在线工具分配能力

    arXiv:2607.23332v1 Announce Type: new Abstract: Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse. Therefore, a user should prefer an agent that creates a small number of highly reusable tools, rather than many…

  236. arXiv cs.AI TIER_1 English(EN) · Aayush Kumar, Avik Dutta, Sumit Gulwani, Gustavo Soares, Advait Sarkar, Emerson Murphy-Hill ·

    计划以神秘的方式运作:评估电子表格代理的计划模式

    arXiv:2607.23670v1 Announce Type: cross Abstract: Plan Modes have become standard features in agentic programming tools, allowing users to gain transparency and control by working with the agent to develop a plan before task execution. However, it remains unclear whether the bene…

  237. arXiv cs.AI TIER_1 English(EN) · Yang Li, Hai Liu, Dian Shao, Yu Wang, Xiyu Chen, Sergey Volkov, Bozhi Wang, Ziyu Sun, Sihang Liu, Ye Luo, Xiaowei Zhang ·

    Agent-UCT:将上限置信度应用于树以实现具有成本意识的代理工作流优化

    arXiv:2607.24162v1 Announce Type: new Abstract: Optimizing agentic workflows, such as retrieval-augmented generation (RAG) pipelines, requires navigating a combinatorial space of discrete component choices under tight evaluation budgets. Existing approaches - heuristic search, bl…

  238. arXiv cs.AI TIER_1 English(EN) · Tianyu Yang, Shir Simon, Zhenzhen Li, Minhao Cheng, Xiangliang Zhang ·

    三思而后行:多模态检索增强生成中的智能体规划

    arXiv:2607.22643v1 Announce Type: new Abstract: Multimodal retrieval-augmented generation (mRAG) aims to answer image-text queries with external knowledge, but most existing systems still retrieve directly from raw multimodal input over a flat evidence space. This design often st…

  239. arXiv cs.AI TIER_1 English(EN) · Zhengyu Chen, Teng Xiao, Huaisheng Zhu, Yige Yuan, Luan Zhang, Jingang Wang ·

    Co-Harness:为 LLM Agent 共同演化 Harness 和模型权重

    arXiv:2607.22688v1 Announce Type: new Abstract: Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research trajectories are generated, evaluated, and learned from. Existing pipelines typicall…

  240. arXiv cs.AI TIER_1 English(EN) · Zedong Yu, Qianxing Li, Zhi Gao, Liuyu Xiang, Chenrui Shi, Yang Liu, Huiming Wu, Yujie Wei, Yuhao Fei, Yubo Fu, Zhaofeng He ·

    超越顺序交互:GUI智能体并行执行与协调基准测试

    arXiv:2607.22689v1 Announce Type: new Abstract: Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs). They perceive screen state and execute user instructions through GUI actions such as clicking, typing, and scrolling on desktops and mobile…

  241. arXiv cs.AI TIER_1 English(EN) · Shawn Ray ·

    什么可以被强制执行?一种关于工具使用代理的认证运行时安全理论

    arXiv:2607.22868v1 Announce Type: new Abstract: Runtime guardrails act before irreversible tool calls, but their guarantees depend on what policy state is representable, what a judge observes, and whether intervention changes future behavior. We separate three questions. First, r…

  242. arXiv cs.AI TIER_1 English(EN) · Summer Sun (Shaqiu Community) ·

    SQBench:用于评估语言模型代理在面向生产的工作流中任务交付能力的基准测试

    arXiv:2607.23123v1 Announce Type: new Abstract: Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced within a constrained workflow as the unit of evaluation. We introduce SQBench, a…

  243. arXiv cs.AI TIER_1 English(EN) · Hao Jiang, Gangtao Xin, Yingdi Huang, Guojie Zhu, Jiangshan Zhang, Xinyuan Lin, Yunkun Xu, Chengyu Shen, Wenlong Fei, Jiawei Li, Yujie Fu, Sichen Kang, Tingyu Xie, Yedi Hu, Jingren Zhang, Hongcheng Gao, Jianshu Zeng, Chong Chen, Chang Guo, Chao Feng, Fen… ·

    AgentOmnia:为全场景应用扩展Agentic模型

    arXiv:2607.23124v1 Announce Type: new Abstract: Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings. We frame this as full-scenario agentic scaling and present AgentOmnia, a fra…

  244. arXiv cs.AI TIER_1 English(EN) · Mingzhou Fan, Siyuan Xu, Mingxuan Yuan ·

    Focus Is All You Need: Adaptive Goal-aware Attention Orchestration for Multi-Agent Graph Systems

    arXiv:2607.23678v1 Announce Type: new Abstract: Large language models (LLMs) enable autonomous agents for reasoning, planning, and tool use. Recent systems increasingly organize these agents as graphs of specialized, interconnected nodes. Although graph-based orchestration suppor…

  245. arXiv cs.AI TIER_1 English(EN) · Weihuang Zheng, Tianyuan Zou, Eileen Ye, Alphet Liu, Youyong Kong, Ya-Qin Zhang, Duran Zheng, Maxm Pan ·

    E-Bench:在真实产品场景中对多步工具使用代理进行基准测试

    arXiv:2607.23722v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden information, composing tool calls, and committing state changes. We refer to this capabi…

  246. arXiv cs.AI TIER_1 English(EN) · Xiaochuan Li, Ryan Ming, Meng Chu, Shuai Shao, Rong Jin, Chenyan Xiong ·

    ACM:面向长时任务的智能体上下文管理

    arXiv:2607.23809v1 Announce Type: new Abstract: Agentic tasks are inherently long-horizon and multi-turn, constantly accumulating context through interactions with the environment. Existing context compression methods inevitably incur information loss and are triggered by rigid h…

  247. arXiv cs.AI TIER_1 English(EN) · Xiangyan Liu, Kaixin Li, Haonan Wang, Biao Wu, Meng Fang, Longxu Dou, Chao Du, Michael Qizhe Shieh, Tianyu Pang ·

    使用视觉状态转换来扩展 GUI 代理

    arXiv:2607.24112v1 Announce Type: new Abstract: We introduce State Transition Pretraining (STP) as a new scaling axis for GUI agents. During the STP stage, we continually pretrain a unified multimodal model on visual state transitions by jointly optimizing inverse dynamics (predi…

  248. arXiv cs.AI TIER_1 English(EN) · Guangyi Liu, Huan Zhao, Quanming Yao ·

    可证伪的自我修正网络代理规划

    arXiv:2607.24167v1 Announce Type: new Abstract: Long-horizon web agents often go off track before final failure: a trajectory can remain locally plausible even after the current state, reused skill, or plan assumption no longer supports the user instruction. Existing agents can p…

  249. arXiv cs.AI TIER_1 English(EN) · Junlin Liu, Jiangwang Chen, Zixin Song, Shuaiyu Zhou, Chunji Lv, Hank Wu, Kailin Jiang, Jinyang Wu, Bohan Yu, Chenxi Zhou ·

    从专有到开源:通过Agentic搜索中的多智能体协议蒸馏弥合分发差距

    arXiv:2607.24280v1 Announce Type: new Abstract: Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision…

  250. Hugging Face Daily Papers TIER_1 English(EN) ·

    CAST:将游戏求解器作为回合制教师来训练LLM代理

    Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse final rewards that reveal little about which decisions determine success. Denser pr…

  251. arXiv cs.MA (Multiagent) TIER_1 Norsk(NO) · Prashant Doshi ·

    PLATO:用于代理和任务开放性的指针学习器

    Open agent systems (OASYS) are increasingly prevalent in real-world domains where the sets of agents and tasks change unpredictably over time. Such openness, including agent openness (AO) and task openness (TO), poses a fundamental challenge to multi-agent reinforcement learning …

  252. arXiv cs.MA (Multiagent) TIER_1 Norsk(NO) · Prashant Doshi ·

    PLATO: Pointer Learner for Agent and Task Openness

    Open agent systems (OASYS) are increasingly prevalent in real-world domains where the sets of agents and tasks change unpredictably over time. Such openness, including agent openness (AO) and task openness (TO), poses a fundamental challenge to multi-agent reinforcement learning …

  253. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agent-UCT:将置信上限应用于树以实现具有成本意识的代理工作流优化

    Optimizing agentic workflows, such as retrieval-augmented generation (RAG) pipelines, requires navigating a combinatorial space of discrete component choices under tight evaluation budgets. Existing approaches - heuristic search, black-box optimization, and standard tree search m…

  254. Hugging Face Daily Papers TIER_1 English(EN) ·

    使用视觉状态转换来扩展 GUI 代理

    We introduce State Transition Pretraining (STP) as a new scaling axis for GUI agents. During the STP stage, we continually pretrain a unified multimodal model on visual state transitions by jointly optimizing inverse dynamics (predicting actions from state changes) and forward dy…

  255. arXiv cs.CL TIER_1 English(EN) · Nanbeige Lab, :, Chen Yang, Chengrui Huang, Fufeng Lan, Hanhui Chen, Hao Zhou, Huatong Song, Jiaqi Cao, Jiaying Zhu, Jinlin Niu, Kai Wang, Lisheng Huang, Qiliang Liang, Ran Le, Ruixiang Feng, Shuang Sun, Tao Gu, Tao Zhang, Tianyu Luo, Yang Song, Yun Xin… ·

    Nanbeige4.2-3B:解锁紧凑模型中的Agentic能力

    arXiv:2607.22083v1 Announce Type: cross Abstract: We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning…

  256. arXiv cs.LG TIER_1 English(EN) · Jiseok Kwak, Suhyeon Jo, Taewoo Kim, Yeongmin Kim, Byeonghu Na, Il-chul Moon ·

    SeePhys Pro的多智能体辩论与视觉信息提取:ICML 2026 AI4Math Track 3挑战赛一等奖技术报告

    arXiv:2607.21946v1 Announce Type: new Abstract: This technical report presents our approach to Challenge Track~3: SeePhys Pro at the 3rd AI for Math Workshop, where the task is to answer college-level physics questions whose statement and figure may be given partly or entirely as…

  257. arXiv cs.LG TIER_1 English(EN) · Ritik Raj, Souvik Kundu, Sarbartha Banerjee, Dheemanth Joshi, Ishita Vohra, Tushar Krishna ·

    TRACE-ROUTER:面向Agentic AI的任务一致且自适应的在线路由

    arXiv:2607.22465v1 Announce Type: cross Abstract: Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make independent routing decisions for each LLM call. …

  258. Hugging Face Daily Papers TIER_1 English(EN) ·

    从专有到开源:通过Agentic搜索中的多智能体协议蒸馏弥合分发差距

    Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision. Knowledge distillation can supply denser guida…

  259. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Wentao Zhang ·

    SearchArt:使用可扩展的合成和验证任务训练长时程搜索代理

    Recent advances in large language models (LLMs) have enabled search agents to autonomously tackle complex tasks across extended search and reasoning horizons. However, training effective search agents remains challenging due to the lack of scalable and long-horizon tasks, and the…

  260. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Tushar Krishna ·

    TRACE-ROUTER:面向Agentic AI的任务一致性与自适应在线路由

    Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make independent routing decisions for each LLM call. However, agentic applications execute as long-hori…

  261. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Tushar Krishna ·

    TRACE-ROUTER:面向Agentic AI的任务一致且自适应的在线路由

    Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make independent routing decisions for each LLM call. However, agentic applications execute as long-hori…

  262. METR (Model Evaluation & Threat Research) TIER_1 English(EN) ·

    Agent Ability Metrics

    <!-- Figure sources: scripts/tikz/2026-07-24-metrics-of-model-ability-*.tex Build with scripts/tikz/build-metrics-of-model-ability.sh. --> <div class="metrics-agent-note"> <!-- > **Goals of this post** > > **Goal:** a simple way to compare a variety of capability metrics in an id…

  263. arXiv cs.AI TIER_1 English(EN) · Zihang Tian, Jingsen Zhang, Rui Li, Xiaohe Bo, Yuanzi Li, Xu Chen ·

    ARCO:具有多步LLM代理的自适应评分卡与协同进化

    arXiv:2606.21262v2 Announce Type: replace Abstract: Reinforcement learning for multi-step LLM agents often relies on scalar rewards that indicate success but cannot explain why a trajectory is good or bad. Rubric-based rewards improve interpretability through natural-language cri…

  264. arXiv cs.AI TIER_1 English(EN) · Bronislav Sidik, Chaya Levi, Nizzan Kimhi ·

    自主拓扑变异:具有能力、状态和影子不变性的多智能体 LLM 系统的安全运行时重构

    arXiv:2607.20488v1 Announce Type: new Abstract: Multi-agent LLM frameworks typically fix their team topology at boot time. When an individual agent becomes overloaded at runtime, for example by mixing too many action categories, accumulating tool errors, or queueing behind too ma…

  265. arXiv cs.AI TIER_1 English(EN) · Jaideep Ray, Ankit Goyal ·

    从Agent失灵到文本策略:哪些有效,哪些失效

    arXiv:2607.20668v1 Announce Type: cross Abstract: TextGrad improves language-model systems by revising text from feedback. Its core thesis is that natural-language feedback can act as a gradient for optimizing text components without changing model weights. Applying it to agents …

  266. arXiv cs.CL TIER_1 English(EN) · Chenhui Gou, Haoqin Tu, Yunhao Fang, Jianfei Cai, Hamid Rezatofighi ·

    Sample-Efficient Learning from Agent Experience

    arXiv:2607.21051v1 Announce Type: new Abstract: Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offers a highly sample-efficient way for agents to learn …

  267. arXiv cs.AI TIER_1 English(EN) · Chao Zhang, Yuhao Wang, Derong Xu, Haoxin Zhang, Yuanjie Lyu, Yuhao Chen, Shuochen Liu, Tong Xu, Xiangyu Zhao, Yan Gao, Yao Hu, Enhong Chen ·

    TeaRAG:一种高效的代理检索增强生成框架

    arXiv:2511.05385v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agentic RAG employs autonomous, multi-round retrieval and reasoning to resolve queries…

  268. arXiv cs.AI TIER_1 English(EN) · Heather Merhout (Miami University), Daniela Inclezan (Miami University) ·

    面向策略感知自主代理的可解释性框架

    arXiv:2607.21209v1 Announce Type: cross Abstract: In the field of Artificial Intelligence, an agent is a system which is able to autonomously make decisions in order to reach a desired goal. As these systems grow more prevalent in our day-to-day lives, there has been an increased…

  269. arXiv cs.AI TIER_1 English(EN) · Xue-Jian Gao, Deng Pan, Yueming Su, Jiasheng Li, Bin Du, Fengming Zhu, Chengdi Ma, Junyi Fan, Qichen Liao, Chengqiu Hu, Xinxian Chen, Lingchao Zheng, Jun Li, Jiwei Yang, Yuwei Fan ·

    CANN Bench:将代理生成的内核与真实NPU和算法限制进行基准测试

    arXiv:2607.20518v1 Announce Type: new Abstract: AI agents are now capable of writing, compiling, and iteratively optimizing low-level operator kernels on different hardware platforms. Existing benchmarks, however, focus almost exclusively on CUDA and Triton, leaving hardware ecos…

  270. arXiv cs.AI TIER_1 English(EN) · Xiao Yu, Baolin Peng, Ruize Xu, Hao Zou, Qianhui Wu, Hao Cheng, Wenlin Yao, Nikhil Singh, Zhou Yu, Jianfeng Gao ·

    OpenForgeRL: 在任何环境中训练原生于Harness的Agent

    arXiv:2607.21557v1 Announce Type: new Abstract: Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard t…

  271. arXiv cs.AI TIER_1 English(EN) · Junzhi Chen, Harsh Trivedi, Jane Pan, Michael JQ Zhang, Tejas Srinivasan, Niranjan Balasubramanian, Ashish Sabharwal ·

    AppWorld-UL:用于工具使用的多样化代理-用户交互基准测试

    arXiv:2607.20536v1 Announce Type: new Abstract: Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.g., to ask clarification questions, prompt for confirmation, and inform the user…

  272. arXiv cs.AI TIER_1 English(EN) · Anas Mohamed, Kaizan Haque, Azal Ahmad Khan, Chetan Sharma, Shuwen Ge, Ali Anwar ·

    面向多智能体系统的负载感知缓存

    arXiv:2607.20495v1 Announce Type: new Abstract: Multi-agent systems decompose complex tasks into directed acyclic graphs (DAGs) of specialized agent executions, creating natural opportunities for caching intermediate results across queries. However, existing cache eviction polici…

  273. arXiv cs.AI TIER_1 English(EN) · Yinhao Tang, Youqing Fang, Yanan Sun, Wenran Liu, Weiming Zhang, Bin Liu, Kuikun Liu, Wenwei Zhang, Kai Chen ·

    SciExplore:从科学导航到信息整合的自主代理评估

    arXiv:2607.20926v1 Announce Type: new Abstract: Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily emphasize general-domain retrieval or static scientific question answering, and th…

  274. arXiv cs.AI TIER_1 English(EN) · Vishal Ishwar Naik, Chenyu Xu, Donna Dong, Hussein Hassan, Abhishek Pradhan, Ofer Mendelevitch, Tallat Shafat, Humayun Irshad ·

    GuardianAgentBench:代理的失败之处及其防护之道

    arXiv:2607.20982v1 Announce Type: new Abstract: As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavior becomes critical. We present GuardianAgentBench (GABench), a benchmark of 580…

  275. arXiv cs.AI TIER_1 English(EN) · Paul Furgale, Severin Klingler, James Nolan, Matt Staats, Gaia Di Lorenzo, Elisa Martinez Abad, Christian Sch\"uller, Razvan Dinu, Alessio Devoto, Pascal Berard, Gal Kaplun, Elad Sarafian, Riccardo Roveri, Leon Derczynski, Ricardo Silveira Cabral ·

    NVIDIA-labs OO Agents:原生 Python 面向对象代理

    arXiv:2607.20709v1 Announce Type: new Abstract: Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework for building reliable AI agents. NO…

  276. arXiv cs.AI TIER_1 English(EN) · Zibin Lin, Shengli Zhang, Taotao Wang, Yihan Xia, Deen Ma, Guofu Liao ·

    工作流本地化机制学习:归因引导修复与结构化智能体技能知识重用

    arXiv:2607.20999v1 Announce Type: new Abstract: Agent Skills package reusable procedural knowledge as external artifacts for frozen language-model agents, yet existing optimizers do not jointly resolve where a failure occurs in a workflow, which mechanism caused it, and how relev…

  277. Hugging Face Daily Papers TIER_1 English(EN) ·

    StateAct:在像素之前,为长时程计算机使用代理提供程序状态

    Computer-use agents are usually improved by strengthening perception: better models for reading a screenshot and choosing where to click. Yet a screenshot is only a lossy rendering of the underlying program state, e.g., the files, application backends, and DOM that hold the task …

  278. arXiv cs.LG TIER_1 English(EN) · Nikolaos Al. Papadopoulos, Ismael Tito Freire, Marti Sanchez-Fibla, Konstantinos E. Psannis ·

    多智能体系统中的时间公平划分:从精确交替度量到可扩展协调代理

    arXiv:2605.14879v2 Announce Type: replace-cross Abstract: Many intelligent computing and autonomous systems rely on multiple independent, often learning, agents repeatedly sharing a limited resource. Examples include autonomous robots accessing a shared workstation, wireless devi…

  279. arXiv cs.AI TIER_1 English(EN) · Chengxiao Dai, Zhanhui Lin, Zhaokun Yan, Youyang Ni, Chenjun Lei, Luyan Zhang ·

    从记忆中协调:图结构化经验重用促进动态制造中的多智能体适应

    arXiv:2607.19985v1 Announce Type: new Abstract: Dynamic manufacturing environments require multi-agent systems to coordinate effectively under frequent operational disturbances such as machine failures, urgent job arrivals, and processing time variations. Existing multi-agent rei…

  280. arXiv cs.AI TIER_1 Norsk(NO) · Yuan Xiong, Linji Hao, Shizhu He, Yequan Wang, Lijun Li ·

    JANUS:预测长周期代理安全中的潜在风险

    arXiv:2607.19913v1 Announce Type: new Abstract: Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-oriented framework for long-horizon agent safety that trains guards to anticipate dela…

  281. arXiv cs.AI TIER_1 English(EN) · Jiazhen Jiang, Boxi Cao, Lingyong Yan, Yaojie Lu, Hongyu Lin, Shuaiqiang Wang, Dawei Yin, Xianpei Han, Le Sun ·

    DocOps:复杂文档操作中自主代理的可验证基准

    arXiv:2607.19865v1 Announce Type: new Abstract: As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we intr…

  282. arXiv cs.CL TIER_1 English(EN) · Qiyuan Liu, Tingfeng Hui, Kun Zhan, Kaike Zhang, Ning Miao ·

    OpenSkillRisk:在实际风险第三方技能使用时对代理进行安全基准测试

    arXiv:2607.20121v1 Announce Type: new Abstract: LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios. However, third-party skills can introduce extra security vulnerabilities, as seemingly harmless skills can contain latent safety risk…

  283. arXiv cs.AI TIER_1 English(EN) · Qiang Zhang, Boli Chen, Fanrui Zhang, Ruixue Ding, Shihang Wang, Qiuchen Wang, Yinfeng Huang, Haonan Zhang, Rongxiang Zhu, Pengyong Wang, Ailin Ren, Xin Li, Pengjun Xie, Jiawei Liu, Ning Guo, Jingren Zhou, Zheng-Jun Zha ·

    ArenaRL:通过基于锦标赛的相对排名来扩展开放式智能体的强化学习

    arXiv:2601.06487v3 Announce Type: replace-cross Abstract: Reinforcement learning has substantially improved the performance of LLM agents on tasks with verifiable outcomes, but it still struggles on open-ended agent tasks with vast solution spaces (e.g., complex travel planning).…

  284. arXiv cs.AI TIER_1 English(EN) · Ahmed Awadallah, Sahil Gupta, Yash Lara, Yadong Lu, Hussein Mozannar, Akshay Nambi, Zach Nussbaum, Yash Pandya, Aravind Rajeswaran, Corby Rosset, Alexey Taymanov, Luiz do Valle, Vibhav Vineet, Spencer Whitehead, Andrew Zhao ·

    Fara-1.5:计算机使用代理的可扩展学习环境

    arXiv:2606.20785v2 Announce Type: replace Abstract: Collecting computer use data from human demonstrations is expensive and slow, motivating the need for scalable generation strategies. This requires two key ingredients: environments in which agents can act and verifiers that can…

  285. arXiv cs.AI TIER_1 English(EN) · Zhuofeng Li, Haoxiang Zhang, Seungju Han, Sheng Liu, Jianwen Xie, Yu Zhang, Yejin Choi, James Zou, Pan Lu ·

    面向有效规划和工具使用的“流中代理”系统优化

    arXiv:2510.05592v2 Announce Type: replace Abstract: Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interleaves thoughts and tool calls under full context…

  286. arXiv cs.AI TIER_1 English(EN) · Elias Hossain, Md Mehedi Hasan Nipu, Fatema Tuj Johora Faria, Tasfia Nuzhat Ornee, Maleeha Sheikh ·

    ChannelGuard:安全模型无法组合成安全的多智能体系统

    arXiv:2607.19430v1 Announce Type: cross Abstract: Multi-agent LLM applications chain a planner, worker agents, a verifier, and a synthesizer, and every hop between agents is an unmonitored channel through which an adversary can smuggle instructions. Existing defenses guard only t…

  287. Hugging Face Daily Papers TIER_1 English(EN) ·

    Sample-Efficient Learning from Agent Experience

    Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offers a highly sample-efficient way for agents to learn from their own interaction histories, but its ga…

  288. Hugging Face Daily Papers TIER_1 English(EN) ·

    OpenForgeRL:在任何环境中训练原生于Harness的Agent

    Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, who…

  289. Hugging Face Daily Papers TIER_1 English(EN) ·

    OpenSkillRisk:在实际风险第三方技能使用时对代理进行安全基准测试

    LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios. However, third-party skills can introduce extra security vulnerabilities, as seemingly harmless skills can contain latent safety risks that only emerge during actual execution. In t…

  290. Hugging Face Daily Papers TIER_1 Norsk(NO) ·

    JANUS:预见长时域智能体安全的潜在风险

    Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-oriented framework for long-horizon agent safety that trains guards to anticipate delayed risks from partial trajectories. Janus synth…

  291. arXiv cs.CL TIER_1 English(EN) · Minghao Guo, Xi Zhu, Qingyue Jiao, Xiujin Liu, Haochen Xue, Chong Zhang, Shuhang Lin, Jingyuan Huang, Ziyi Ye, Yongfeng Zhang ·

    Node-as-Agent: 图形智能体网络

    arXiv:2508.00429v5 Announce Type: replace Abstract: Graph Neural Networks (GNNs) have achieved remarkable success in graph-based learning by propagating information among neighbor nodes via predefined aggregation mechanisms. However, such fixed schemes often suffer from two key l…

  292. arXiv cs.AI TIER_1 English(EN) · Hassan Karim, Sai Sitharaman, Deepti Gupta, Danda B. Rawat ·

    从Agent失败路径到量化残余风险:构建弹性Agentic AI的组合框架

    arXiv:2607.18243v1 Announce Type: new Abstract: Agentic AI is crossing trust boundaries faster than current risk models can represent. Existing approaches provide one of two partial views. They either describe failure mechanisms without producing a transferable residual-risk esti…

  293. arXiv cs.AI TIER_1 English(EN) · Ritvik Garimella, Vedant Khandelwal, Anvi Kohli, Amit Sheth ·

    SAAG:结构化代理评估与基础

    arXiv:2607.18245v1 Announce Type: new Abstract: Exact-match evaluation of agent-calling obscures qualitatively different failure modes: a model may select the right function yet hallucinate argument values, or satisfy a schema while choosing a agent for the wrong reason. Existing…

  294. arXiv cs.AI TIER_1 English(EN) · Tianyue Jiang, Yanlin Wang, Xin He, Daya Guo, Jiachi Chen, Ming Wen, Ensheng Shi, Xilin Liu, Yuchi Ma, Guanbin Li ·

    PhoenixRepair:重新思考软件代理中的修复策略探索

    arXiv:2607.18859v1 Announce Type: new Abstract: While Large Language Models have greatly advanced automated issue resolution, existing agent-based methods exhibit a fundamental limitation in their insufficient exploration of repair strategies. This insufficiency manifests in two …

  295. arXiv cs.AI TIER_1 English(EN) · Pengyi Jiang, Xiaoguang Zhu, Quanyan Zhu ·

    基于LLM的多智能体系统中用于贡献归因的语义合作博弈

    arXiv:2607.18255v1 Announce Type: new Abstract: Contribution attribution has become a central problem in LLM-based multi-agent systems, where final outputs are produced through multiple agents, message exchanges, and ordered workflow dependencies. Existing attribution methods oft…

  296. arXiv cs.AI TIER_1 English(EN) · Jinjie Wei, Jiyao Liu, Lihao Liu, Ming Hu, Junzhi Ning, Mingcheng Li, Weijie Yin, Junjun He, Xiao Liang, Chao Feng, Dingkang Yang ·

    学习、推理、优化:Kahneman双系统智能在GUI代理中的框架

    arXiv:2506.17913v2 Announce Type: replace Abstract: Graphical User Interface (GUI) agents have made significant progress in automating digital tasks through the utilization of computer vision and language models. Nevertheless, existing agent systems encounter notable limitations.…

  297. arXiv cs.AI TIER_1 English(EN) · SangJin Park, Myungsub Choi, Jineok Kim, Minseung Kang ·

    跨代理活动归因:连接跨 LLM 代理的异步攻击

    arXiv:2607.18826v1 Announce Type: cross Abstract: LLM-agent defenses are typically evaluated one session at a time. In deployment, however, attacks can be distributed across independent agents, teams, and runtimes, leaving each local guardrail with only a sparse fragment. We form…

  298. arXiv cs.AI TIER_1 English(EN) · Kunlun Zhu, Xuyan Ye, Zhiguang Han, Yuchen Zhao, Bingxuan Li, Weijia Zhang, Muxin Tian, Xiangru Tang, Pan Lu, James Zou, Jiaxuan You, Heng Ji ·

    AgentDebugX:用于 LLM Agent 故障可观测性、归因和恢复的开源工具包

    arXiv:2607.18754v1 Announce Type: new Abstract: LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cause o…

  299. arXiv cs.AI TIER_1 English(EN) · David G\'omez-Guill\'en, Mireia Diaz, Josep Lluis Arcos, Jes\'us Cerquides ·

    Agentic Calibration of Grey-Box Simulation Models: An LLM-Driven Alternative

    arXiv:2607.18308v1 Announce Type: cross Abstract: Calibration of grey-box simulation models is a constrained optimization problem in which model evaluations are expensive, the parameter space can be high-dimensional, and the search must respect plausibility constraints. Although …

  300. arXiv cs.AI TIER_1 English(EN) · Daniel Pearson, Sidney Shapiro, Emiliano Sebastian Gonzalez Venegas, Sanad Al-Khatib, Aurora Pinz\'on Arzola ·

    基于 LangGraph 的图式智能体 AI:用于长期有状态业务流程的工作流路径

    arXiv:2607.19297v1 Announce Type: new Abstract: This paper is a practitioner guide to graph-based workflow pathways for long-running, stateful, multi-step generative AI systems in business processes. Rather than treating LangGraph, a low-level orchestration framework for stateful…

  301. arXiv cs.AI TIER_1 English(EN) · Grace Hui Yang, Pranav N. Venkit, Hooman Sedghamiz, Enrico Santus, Victor Dibia, Ioana Baldini ·

    Agents in the Wild: Where Research Meets Deployment

    arXiv:2607.19336v1 Announce Type: new Abstract: Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale deployments a…

  302. arXiv cs.AI TIER_1 English(EN) · Rahul Suresh Babu, Shashank Indukuri ·

    多步工具增强代理中的绑定漂移

    arXiv:2607.18316v1 Announce Type: cross Abstract: Tool-augmented language-model agents execute multi-step workflows over external systems, resolving an entity once and then acting on it across subsequent steps. Prior work shows that in single-step actions, agents select the corre…

  303. arXiv cs.AI TIER_1 English(EN) · Eden Wu, Sonia Castelo, Yurong Liu, Cl\'audio T. Silva, Juliana Freire ·

    AgentTrails:迈向Agentic任务的信任与复用

    arXiv:2607.18816v1 Announce Type: cross Abstract: LLM-powered agents increasingly tackle complex tasks by invoking tools, querying databases, executing code, and manipulating intermediate artifacts. These agents follow trajectories that are typically stored as chronological logs,…

  304. arXiv cs.LG TIER_1 English(EN) · Shuangyao Huang ·

    面向连续动作空间的合作任务的自演化默认动作

    arXiv:2607.18597v1 Announce Type: new Abstract: Counterfactual credit assignment has proven effective in multi-agent reinforcement learning (MARL) for discrete action spaces, yet its extension to continuous-action cooperative tasks remains challenging. Existing methods that appro…

  305. Hugging Face Daily Papers TIER_1 English(EN) ·

    NVIDIA-labs OO Agents:原生 Python 面向对象代理

    Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework for building reliable AI agents. NOOA takes a simpler approach: an agent is a Pytho…

  306. Hugging Face Daily Papers TIER_1 English(EN) ·

    DocOps:复杂文档操作中自主代理的可验证基准

    As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we introduce DocOps, a deterministically verifiable eva…

  307. arXiv cs.AI TIER_1 English(EN) · Petter Holme, Milena Tsvetkova ·

    人工智能代理在社会和行为科学中的应用:历史与展望

    arXiv:2510.05743v3 Announce Type: replace Abstract: We review the historical development and current trends of artificially intelligent agents (agentic AI) in the social and behavioral sciences: from the first programmable computers, and social simulations soon thereafter, to tod…

  308. arXiv cs.AI TIER_1 English(EN) · Hao Dou ·

    CIGPO:面向多轮证据阅读LLM智能体的上下文信息增益策略优化

    arXiv:2607.16244v1 Announce Type: cross Abstract: Training multi-turn evidence-reading agents with outcome-only reinforcement learning is unstable because intermediate turns receive little direct credit. In HotpotQA experiments with Qwen2.5-3B-Instruct, GRPO initially improves (s…

  309. arXiv cs.AI TIER_1 English(EN) · Chris Ge, Daria Kryvosheieva, Daniel Fried, Uzay Girit, Kaivalya Hariharan ·

    Agent psychometrics: Agentic coding benchmarks中的任务级性能预测

    arXiv:2604.00594v2 Announce Type: replace Abstract: As the focus in LLM-based coding shifts from static single-step code generation to multi-step agentic interaction with tools and environments, understanding which tasks will challenge agents and why becomes increasingly difficul…

  310. arXiv cs.AI TIER_1 English(EN) · Muzhi Li, Jinhu Qi, Yihong Wu, Minghao Zhao, Liheng Ma, Yifan Li, Xinyu Wang, Zhenghan Tai, Zixing Song, Yingxue Zhang, Ho-fung Leung, Irwin King ·

    从证据到轨迹:检索增强生成代理开发的溯因推理路径合成

    arXiv:2509.23071v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) agent development is hindered by the lack of executable ground-truth agent-environment interaction trajectories. Existing datasets provide questions, answers, and evidence, but lack fin…

  311. arXiv cs.LG TIER_1 English(EN) · Zhuohang Fan, Beichen Zhang, Yuanfa Li, Changqiao Wu, Wei Liu, Jian Luan, Weigang Zhang ·

    SEE:面向长视界GUI代理轨迹合成的结构感知探索与利用

    arXiv:2607.18046v1 Announce Type: new Abstract: Graphical User Interface (GUI) agents powered by vision-language models hold promise for automating real-world mobile tasks. However, progress is limited by the lack of high-coverage, long-horizon interaction trajectories collected …

  312. arXiv cs.LG TIER_1 English(EN) · Krish Agarwal, Zhuoming Chen, Yanyuan Qin, Zhenyu Gu, Atri Rudra, Beidi Chen ·

    FlashRT:用于指导代理部署实时多模态应用的代理集线器

    arXiv:2607.18171v1 Announce Type: new Abstract: Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, a…

  313. arXiv cs.AI TIER_1 English(EN) · Arunabh Dastidar (for the Leni Team) ·

    智能体可靠性从何而来?生产型企业智能体中验证循环、专家模型和脚手架的跨基准分解

    arXiv:2607.17044v1 Announce Type: cross Abstract: Multi-step enterprise agent tasks fail in a characteristic way: single-pass inference has no checkpoint between deciding an answer and committing to it. We study one production system (Leni) whose architecture installs such checkp…

  314. arXiv cs.CL TIER_1 English(EN) · Bo Tang, Yang Zhang, Guomian Zhuang, Wenqiang Wei, Gaoyang Zheng, Lindong Xie, Yanchao Tan, Feiyu Xiong, Qingyu Yang, Edward Chung, Zhiyu li ·

    从记忆到技能:基于证据的共演化治理,用于长时域LLM智能体

    arXiv:2607.16621v1 Announce Type: new Abstract: Existing memory systems for long-horizon LLM agents often retrieve prior traces as passive context rather than converting them into executable capabilities. In this paper, we propose MSCE, a training-free Memory--Skill Co-Evolution …

  315. arXiv cs.AI TIER_1 English(EN) · Yangqin Jiang, Chao Huang ·

    AgentBrew:从强教师到弱LLM代理的终身知识酿造

    arXiv:2607.16851v1 Announce Type: new Abstract: Deploying LLM agents typically requires a compact test-time student, even if a stronger teacher is available during training. We study knowledge brewing: distilling a teacher's interactive experience into a persistent external memor…

  316. arXiv cs.AI TIER_1 English(EN) · Amez Amanj Ali, Kuo-Kun Tseng ·

    奖励驱动的LLM代理工作流:融合POMDP路由与自纠错以实现自主决策

    arXiv:2607.17038v1 Announce Type: new Abstract: This paper addresses key technical challenges in current large language model (LLM) agent applications, including long-horizon planning, sparse reward attribution, and dynamic environmental interaction, by designing and optimizing a…

  317. arXiv cs.AI TIER_1 English(EN) · Chen Xia, Zexi Kuang, Yuqing Hu ·

    经验性基础提高了LLM代理在中断期间模拟人类行为的真实性

    arXiv:2607.17437v1 Announce Type: new Abstract: Large language model (LLM) agents offer a generative approach to simulating human behavior under conditions that may have few or no direct historical analogues, a common challenge in disaster and infrastructure-disruption planning. …

  318. arXiv cs.AI TIER_1 English(EN) · Guanzhen Li, Liangming Pan, Leye Wang ·

    ProEvent:一个以事件为中心的积极代理基准

    arXiv:2607.17701v1 Announce Type: new Abstract: Proactive agents are expected to anticipate user needs and provide autonomous assistance by perceiving environmental context without explicit instructions. A fundamental capability of such agents is to identify and track users' upco…

  319. arXiv cs.AI TIER_1 English(EN) · Stefano Blando, Emanuele Guerrazzi, Riccardo Porcedda, Giuseppe Squillace, Max Tschaikowski, Andrea Vandin ·

    迈向量体智能体模型:可行性、性能与统计模型检查

    arXiv:2607.17948v1 Announce Type: new Abstract: Agent-based models (ABMs) rely on simple, explicit and reproducible rules for individual decision making, while complex collective behavior emerges from interactions among agents. Recent advances in large language models (LLMs) make…

  320. arXiv cs.AI TIER_1 English(EN) · Zishang Jiang, Tingyun Li, Jinyi Han, Xinyi Wang, Sihang Jiang, Yizhou Ying, Xiaojun Meng, Jiansheng Wei, Jiaqing Liang, Yanghua Xiao ·

    从结果到行动:利用事后诸葛亮进行长时域语言代理训练

    arXiv:2607.16257v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a widely adopted technique for improving large language models (LLMs) on complex tasks. Despite this progress, existing RL methods still face challenges in training agents with longer-horizon…

  321. arXiv cs.CL TIER_1 English(EN) · Wei Chen, Zhiyuan Li, Zhen Guo, Yikang Shen ·

    Octo-planner:用于规划器-动作代理的设备端语言模型

    arXiv:2406.18082v2 Announce Type: replace Abstract: AI agents have become increasingly significant in various domains, enabling autonomous decision-making and problem-solving. To function effectively, these agents require a planning process that determines the best course of acti…

  322. arXiv cs.AI TIER_1 English(EN) · Tejas Singh Anand, Yuet Ying Christina Wang, Wanting Jiang, Steve Masson, Tian Zheng, Bingjie Zhou ·

    AEVAL:从轶事到确定性测试,用于 Agentic Skill Workflows

    arXiv:2607.16345v1 Announce Type: cross Abstract: Modern agentic systems increasingly rely on skills: installable packages of natural language and code that teach an LLM agent to perform a domain task. As skill repositories grow, developers need automated quality signals on every…

  323. arXiv cs.AI TIER_1 English(EN) · Julian Alfredo Mendez, Andreas Br\"annstr\"om ·

    面向多智能体系统的可组合验证流水线

    arXiv:2607.16266v1 Announce Type: cross Abstract: Existing approaches for reasoning about action and change provide expressive semantics for modeling dynamic systems, in most cases built on top of logic programming systems. We introduce a modular framework for transition and traj…

  324. arXiv cs.AI TIER_1 English(EN) · Lijie Zheng, Xudong Zhong, Baoquan Ren, Xiangwu Gong, Xinghui Zhu, Ji He ·

    从意图到基础设施:面向 ISAC 网络的大语言模型驱动的智能体编译器

    arXiv:2607.16269v1 Announce Type: cross Abstract: Integrated sensing and communications (ISAC) is moving from proof-of-concept demonstrations to system-level deployment in sixth-generation (6G) networks. Because sensing and communication share hardware, spectrum, and waveform res…

  325. arXiv cs.AI TIER_1 English(EN) · Nguyen Viet Tuan Kiet, Bui Dinh Pham, Duong Quoc Chinh, Dao Van Tung, Tran Cong Dao, Huynh Thi Thanh Binh ·

    RELIC:多智能体规划中可解释可组合技能学习的揭示原理

    arXiv:2607.16745v1 Announce Type: new Abstract: Multi-agent planning becomes substantially harder when agents must improve specialized decision-making skills while keeping their internal implementations private. This regime arises when agents are developed independently, expose d…

  326. arXiv cs.AI TIER_1 English(EN) · Yuqing Li, Zeguan Wu, Yu Gan, Junyu Liu ·

    具有验证器接地基准协同进化的自修改精益证明代理

    arXiv:2607.17352v1 Announce Type: new Abstract: Designing effective Lean proof agents is a central challenge in formal mathematical reasoning. Beyond building stronger provers, recent work emphasizes the workflow around Lean: how an agent decomposes proof obligations, uses tools …

  327. arXiv cs.AI TIER_1 English(EN) · Babak Barazandeh, Subhabrata Majumdar, George Michailidis ·

    Otap: 结构感知最优传输用于评估智能体轨迹中的规划与执行

    arXiv:2607.17082v1 Announce Type: new Abstract: Large language model agents solve tasks by generating trajectories that interleave planning, tool calls, and intermediate results. Current evaluation metrics reduce such a trajectory to a binary success flag or compare it against a …

  328. arXiv cs.AI TIER_1 English(EN) · Huiri Tan, Yikun Wang, Puyang Zhang, Shangyu Li, Jiasi Shen ·

    ETAS:一种用于代理系统的效果类型语言

    arXiv:2607.17780v1 Announce Type: cross Abstract: ETAS is a programming language for agent systems that treats model-backed agents, tool calls, prompts, typed memory, human approvals, policies, and execution traces as semantic program elements rather than library conventions. It …

  329. arXiv cs.AI TIER_1 English(EN) · Ryan Xu, Atlas Zhao, David Bao, Frank Du ·

    WAR:面向同步智能体强化学习的工作负载感知回滚

    arXiv:2607.17299v1 Announce Type: cross Abstract: Long-horizon rollout generation has become the dominant systems bottleneck in agentic reinforcement learning (RL). As agents interact with environments over many turns, trajectories rapidly grow to tens of thousands of tokens, mak…

  330. arXiv cs.AI TIER_1 English(EN) · Jiacheng Ding, Xiaofei Zhang ·

    SAGA:用于时间基准生成的人工合成图架构

    arXiv:2607.17288v1 Announce Type: cross Abstract: High quality temporal graph benchmarks with rich semantics and ground-truth anomaly labels are essential for training graph neural networks, yet remain scarce due to privacy constraints and annotation costs. We present SAGA (Synth…

  331. arXiv cs.AI TIER_1 English(EN) · Chetan Arora, Andreas Vogelsang, Abbi Sharma ·

    明确委托自主边界:Agentic AI 的需求工程

    arXiv:2607.17225v1 Announce Type: cross Abstract: Agentic AI systems do not just predict or recommend; they plan, maintain state, and act in external environments with varying degrees of autonomy. This changes the requirements engineering problem in a specific and under-addressed…

  332. arXiv cs.AI TIER_1 English(EN) · Rasheed Mudasiru ·

    AI Agent Systems 的确定性回放

    arXiv:2607.16200v1 Announce Type: new Abstract: AI agent systems that couple large language models (LLMs) with external tools and APIs are inherently non-deterministic: LLM sampling variance, external API state, CDN infrastructure headers, and execution-environment noise collecti…

  333. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Shuangyao Huang ·

    面向连续动作空间的合作任务的自适应默认动作

    Counterfactual credit assignment has proven effective in multi-agent reinforcement learning (MARL) for discrete action spaces, yet its extension to continuous-action cooperative tasks remains challenging. Existing methods that approximate the counterfactual baseline via Monte Car…

  334. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Shuangyao Huang ·

    面向连续动作空间的合作任务的自演化默认动作

    Counterfactual credit assignment has proven effective in multi-agent reinforcement learning (MARL) for discrete action spaces, yet its extension to continuous-action cooperative tasks remains challenging. Existing methods that approximate the counterfactual baseline via Monte Car…

  335. Hugging Face Daily Papers TIER_1 English(EN) ·

    AgentDebugX:用于 LLM Agent 故障可观测性、归因和恢复的开源工具包

    LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cause or translating diagnosis into recovery. We presen…

  336. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Maleeha Sheikh ·

    ChannelGuard:安全模型无法组合成安全的多智能体系统

    Multi-agent LLM applications chain a planner, worker agents, a verifier, and a synthesizer, and every hop between agents is an unmonitored channel through which an adversary can smuggle instructions. Existing defenses guard only the input boundary (IBProtector, Llama Guard, perpl…

  337. Hugging Face Daily Papers TIER_1 English(EN) ·

    SEE:面向长视界 GUI 代理轨迹合成的结构感知探索与利用

    Graphical User Interface (GUI) agents powered by vision-language models hold promise for automating real-world mobile tasks. However, progress is limited by the lack of high-coverage, long-horizon interaction trajectories collected from element-rich and rapidly evolving apps. Exi…

  338. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Andrea Vandin ·

    迈向量体智能体模型:可行性、性能与统计模型检查

    Agent-based models (ABMs) rely on simple, explicit and reproducible rules for individual decision making, while complex collective behavior emerges from interactions among agents. Recent advances in large language models (LLMs) make it tempting to replace, enrich, or perturb thes…

  339. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Jiasi Shen ·

    ETAS:一种用于代理系统的效果类型语言

    ETAS is a programming language for agent systems that treats model-backed agents, tool calls, prompts, typed memory, human approvals, policies, and execution traces as semantic program elements rather than library conventions. It separates deterministic computation from agentic n…

  340. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Dimitrios S. Sfiris ·

    面向可信代理式商业的决策中心参考架构

    Agentic commerce extends agentic shopping into software agents that interpret policy, prepare checkout, generate transaction-facing language, and act under delegated payment authority. Protocols standardize external exchanges, but merchants still need one authoritative representa…

  341. arXiv cs.AI TIER_1 English(EN) · Zherui Yang, Fan Liu, Hao Liu ·

    DSWorld:面向高效自主代理的数据科学世界模型

    arXiv:2607.15901v1 Announce Type: new Abstract: Despite strong capabilities in data understanding and decision-making, autonomous data science agents still heavily rely on trial-and-error workflows that involve expensive computation. This bottleneck motivates models that can anti…

  342. arXiv cs.AI TIER_1 English(EN) · Wendi Yu, Lianhao Zhou, Xiangjue Dong, Sai Sudarshan Barath, Declan Staunton, Byung-Jun Yoon, Xiaoning Qian, James Caverlee, Shuiwang Ji ·

    多智能体系统何时能提供帮助?信息瓶颈视角

    arXiv:2607.16133v1 Announce Type: cross Abstract: LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantages over single-agent systems (SAS) remain unclear, with performance varying inconsistently across settings. Here,…

  343. arXiv cs.AI TIER_1 English(EN) · Lujia Zhang, Xingzhou Chen, Hongwei Feng ·

    用于信息提取的智能体模型的行为可控性:从固定工作流到反思型智能体

    arXiv:2607.15715v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used for complex information-extraction tasks, yet it remains unclear whether agentic components such as reflection and memory lead to observable and controllable improvements over …

  344. arXiv cs.AI TIER_1 English(EN) · Shuaiyu Zhou, Fengpeng Yue, Zengjie Hu, Yuanzhe Shen, Chenyang Zhang, feng hong, Cao Liu, Ke Zeng ·

    ToolVerse:为智能体强化学习解锁海量环境和长时域任务

    arXiv:2607.15660v1 Announce Type: new Abstract: While LLM agents demonstrate strong reasoning abilities in compact and well-defined scenarios, they struggle to maintain robustness and effectiveness when faced with large-scale, diverse, and dynamic real-world environments that dem…

  345. Hugging Face Daily Papers TIER_1 English(EN) ·

    FlashRT:用于指导代理部署实时多模态应用的代理工具包

    Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and intra-model parallelism. Existing serving sys…

  346. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Min Xu ·

    为智能体能力检索调整嵌入模型

    Open agent marketplaces list native agents, tool bundles, and reusable skill packages in the same search interface, yet practitioners still have little guidance on how to retrieve across this mixed catalog. We study whether off-the-shelf retrieval models, trained for general text…

  347. Hugging Face Daily Papers TIER_1 English(EN) ·

    SAGA:用于时间基准生成的人工代理图架构

    High quality temporal graph benchmarks with rich semantics and ground-truth anomaly labels are essential for training graph neural networks, yet remain scarce due to privacy constraints and annotation costs. We present SAGA (Synthetic Agentic Graph Architecture), a system for gen…

  348. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Huynh Thi Thanh Binh ·

    RELIC:多智能体规划中可解释可组合技能学习的揭示原则

    Multi-agent planning becomes substantially harder when agents must improve specialized decision-making skills while keeping their internal implementations private. This regime arises when agents are developed independently, expose different interfaces and capabilities, and must n…

  349. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Seyyedali Hosseinalipour ·

    SAGE: 异构多智能体导航的社会感知生成引擎

    Safe and socially compliant navigation in open human-robot environments requires robots to reason about heterogeneous participants with different dynamics, autonomy levels, and social roles. Existing trajectory prediction and planning methods often rely on homogeneous interaction…

  350. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Seyyedali Hosseinalipour ·

    SAGE:异构多智能体导航的社会意识生成引擎

    Safe and socially compliant navigation in open human-robot environments requires robots to reason about heterogeneous participants with different dynamics, autonomy levels, and social roles. Existing trajectory prediction and planning methods often rely on homogeneous interaction…

  351. arXiv cs.AI TIER_1 English(EN) · Weiting Liu, Jieyi Bi, Wanqi Zhou, Jianfeng Feng, Yining Ma, Ai Han, Wenlian Lu ·

    ToolAnchor:锚定反事实上下文以增强代理工具使用能力

    arXiv:2607.14145v1 Announce Type: new Abstract: Tool-augmented large language model agents excel at long-horizon tasks, yet they are typically post-trained on fixed toolsets. When tasks demand new tools, these agents struggle to incorporate them effectively, and retraining from s…

  352. arXiv cs.CL TIER_1 English(EN) · Amirhosein Ghasemabadi, Ruichen Chen, Bahador Rashidi, Di Niu ·

    多头潜在控制:LLM 代理决策的统一接口

    arXiv:2607.14277v1 Announce Type: new Abstract: Large language models are increasingly deployed as agents, but reliable agentic behavior requires more than next-token prediction. At inference time, it is preferred that an agent can decide whether to proceed with its current reaso…

  353. arXiv cs.CL TIER_1 English(EN) · Renze Lou, Baolin Peng, Wenlin Yao, Qianhui Wu, Hao Cheng, Suman Nath, Wenpeng Yin, Jianfeng Gao ·

    工具幻觉:重新思考Web代理中的工具使用

    arXiv:2604.03465v2 Announce Type: replace Abstract: As web agents rapidly evolve, an increasing body of work has moved beyond conventional atomic browser interactions and explored tool use as a higher-level action paradigm. Although prior studies have shown the promise of tools, …

  354. arXiv cs.AI TIER_1 English(EN) · Yuyao Zhang, Junjie Gao, Zhengxian Wu, Jiaming Fan, Jin Zhang, Shihan Ma, Yao Yao, Weiran Qi, Chuyan Jin, Guiyu Ma, Xingzhong Xu, Kai Yang, Ji-Rong Wen, Zhicheng Dou ·

    SearchOS-V1:迈向鲁棒的开放域信息检索智能体协作

    arXiv:2607.15257v1 Announce Type: new Abstract: Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to track task progress. When search …

  355. arXiv cs.AI TIER_1 English(EN) · Huanxi Liu, Kun Hu, Jiaqi Liao, Qiang Wang, Pengfei Qian, YuanZhao Zhai, Dawei Feng, Bo Ding, Huaimin Wang ·

    MCPEvol-Bench:跨越MCP服务器动态演进的LLM智能体性能基准测试

    arXiv:2607.14642v1 Announce Type: new Abstract: As Model Context Protocol (MCP) servers emerge as the core infrastructure for connecting LLMs with external tools, existing benchmarks leverage real-world MCP servers to evaluate LLM agents' tool-using capabilities. However, these b…

  356. arXiv cs.AI TIER_1 (AF) · Minghao Liu, Yu Wang, Jiayun Wang, Wei Wei ·

    通过成对验证器实现无奖励的进化智能体

    arXiv:2607.14408v1 Announce Type: new Abstract: A self-evolving agentic loop repeatedly proposes a tweaked version of an agent (its prompt template or program) and accepts or rejects the change based on a per-iteration quality signal. Designing that signal is often the costly par…

  357. arXiv cs.AI TIER_1 English(EN) · Jaideep Ray, Ankit Goyal ·

    结构化反馈可改进 LLM Agent 循环中的修复

    arXiv:2607.14167v1 Announce Type: cross Abstract: LLM agents often retry after external validation rejects a candidate, but the interface between validation and the next model call remains underspecified. We introduce VeriHarness, a code-controlled agent loop in which models gene…

  358. arXiv cs.AI TIER_1 English(EN) · Yue Huang, Wenjie Wang, Han Bao, Yuchen Ma, Xiaonan Luo, Yi Nian, Haomin Zhuang, Zheyuan Liu, Yue Zhao, Xiangliang Zhang ·

    MemoHarness:从经验中学习的智能体(Agent)工具

    arXiv:2607.14159v1 Announce Type: new Abstract: An agent harness is the external control layer that turns a base LLM into an executable agent by managing context, tools, orchestration, memory, decoding, and output handling. While harness design strongly affects agent behavior, mo…

  359. arXiv cs.AI TIER_1 English(EN) · Paul Kassianik, Blaine Nelson, Yaron Singer ·

    超越成功率:进攻性和防御性安全代理的成本感知评估

    arXiv:2607.15263v1 Announce Type: cross Abstract: Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are usefu…

  360. arXiv cs.AI TIER_1 English(EN) · Jiarong Zhao, Zhikai Lei, Zhiheng Xi, Rui Zheng, Hang Yan, Jie Zhou, Qin Chen, Liang He ·

    NexForge:通过需求优先合成扩展可执行代理任务

    arXiv:2607.14186v1 Announce Type: cross Abstract: Scaling executable agent training data is bottlenecked by substrate-first methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual expansion of the substrate, each new…

  361. arXiv cs.AI TIER_1 English(EN) · Mu Yuan, Jinke Song, Zhaomeng Zhou, Lan Zhang ·

    ANet Patu-1: 代理网络中的连接价值

    arXiv:2607.15053v1 Announce Type: cross Abstract: The Internet taught us that the value of a network depends on \emph{how} its nodes connect: broadcast stars scale as $V\!\propto\!N$ (Sarnoff), fully-connected meshes as $N^2$ (Metcalfe), and group-forming networks as $2^{N}$ (Ree…

  362. arXiv cs.AI TIER_1 English(EN) · Dylan Van Mulders, Matthias Bogaert, Dirk Van den Poel ·

    数字万神殿:使用 LLM 代理模拟和审计联盟形成

    arXiv:2607.15095v1 Announce Type: cross Abstract: The formation of political coalitions is a complex negotiation driven by both concrete policy objectives and deep-seated ideological convictions. While Large Language Models (LLMs) open new avenues for computational political scie…

  363. arXiv cs.AI TIER_1 English(EN) · Boning Zhao, Yutong Hu, Xinnuo Li ·

    从无状态到有状态:为基于LLM的代理构建心理世界

    arXiv:2603.25031v2 Announce Type: replace Abstract: In psychological support and emotional companionship scenarios, the core limitation of large language models (LLMs) lies not merely in response quality, but in their reliance on local next-token prediction, which prevents them f…

  364. arXiv cs.AI TIER_1 English(EN) · Yihao Zhang, Zeming Wei, Xiaokun Luan, Chengcan Wu, Zhixin Zhang, Jiangrong Wu, Haolin Wu, Huanran Chen, Jun Sun, Meng Sun ·

    AgentWorm:跨 LLM Agent 生态系统的自我传播攻击

    arXiv:2603.15727v3 Announce Type: replace-cross Abstract: Autonomous LLM-based agents increasingly operate as long-running processes forming densely interconnected multi-agent ecosystems, whose security properties remain largely unexplored. Systems such as OpenClaw, an open-sourc…

  365. Hugging Face Daily Papers TIER_1 English(EN) ·

    DSWorld:面向高效自主代理的数据科学世界模型

    Despite strong capabilities in data understanding and decision-making, autonomous data science agents still heavily rely on trial-and-error workflows that involve expensive computation. This bottleneck motivates models that can anticipate the effects of data science operations be…

  366. arXiv cs.AI TIER_1 English(EN) · Yaron Singer ·

    超越成功率:进攻性和防御性安全代理的成本感知评估

    Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every r…

  367. arXiv cs.AI TIER_1 English(EN) · Zhicheng Dou ·

    SearchOS-V1:迈向鲁棒的开放域信息检索代理协作

    Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to track task progress. When search attempts fail to yield useful evidence, current …

  368. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Dirk Van den Poel ·

    数字万神殿:使用LLM代理模拟和审计联盟形成

    The formation of political coalitions is a complex negotiation driven by both concrete policy objectives and deep-seated ideological convictions. While Large Language Models (LLMs) open new avenues for computational political science, the neutrality and helpfulness biases instill…

  369. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Dirk Van den Poel ·

    数字万神殿:使用 LLM 代理模拟和审计联盟形成

    The formation of political coalitions is a complex negotiation driven by both concrete policy objectives and deep-seated ideological convictions. While Large Language Models (LLMs) open new avenues for computational political science, the neutrality and helpfulness biases instill…

  370. arXiv cs.AI TIER_1 English(EN) · Lan Zhang ·

    ANet Patu-1: Agent网络中的连接价值

    The Internet taught us that the value of a network depends on \emph{how} its nodes connect: broadcast stars scale as $V\!\propto\!N$ (Sarnoff), fully-connected meshes as $N^2$ (Metcalfe), and group-forming networks as $2^{N}$ (Reed). We ask the analogous question for networks of …

  371. arXiv cs.AI TIER_1 English(EN) · Huaimin Wang ·

    MCPEvol-Bench:跨越MCP服务器动态演进的LLM智能体性能基准测试

    As Model Context Protocol (MCP) servers emerge as the core infrastructure for connecting LLMs with external tools, existing benchmarks leverage real-world MCP servers to evaluate LLM agents' tool-using capabilities. However, these benchmarks overlook the continuous evolution of t…

  372. arXiv cs.AI TIER_1 English(EN) · Sagar Deb, Ashwanth Krishnan ·

    STOCKTAKE:用公平的Oracle衡量LLM Agent中感知与行动的差距

    arXiv:2607.13618v1 Announce Type: new Abstract: LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the final cost cannot say why an agent failed: it may have misread the world, or read i…

  373. arXiv cs.AI TIER_1 English(EN) · Ziwei Ye ·

    可控智能体的任务切换行为测试

    arXiv:2607.13396v1 Announce Type: new Abstract: What happens to an LLM agent's tool choice when the reliable tool silently changes within an ongoing session? We borrow set-shifting from cognitive psychology to study how well agents adapt to hidden reliability shifts. Our benchmar…

  374. arXiv cs.AI TIER_1 English(EN) · Ruhan Wang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Yue Yu, Junyao Yang, Kishan Panaganti, Haitao Mi, Dongruo Zhou, Leoweiliang ·

    Harness Handbook:让不断演进的Agent Harness易于阅读、导航和编辑

    arXiv:2607.13285v1 Announce Type: new Abstract: The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution. As models, APIs, environments, and requirements…

  375. arXiv cs.AI TIER_1 English(EN) · Ilias Kazantzidis, Timothy J. Norman, Yali Du, Christopher T. Freeman ·

    通过世界模型从人类偏好和理由中学习安全代理行为

    arXiv:2607.13172v1 Announce Type: new Abstract: We address the problem of safely training an agent policy and deploying a good and safe policy, in settings where the environment dynamics are unknown and no suitable reward function is available. In the context of safety-critical e…

  376. arXiv cs.AI TIER_1 English(EN) · Zhe Ren, Yimeng Chen, Dandan Guo, Guowei Rong, Tonghui Li, R. B. Xiong, Qingfeng Lan, Wenyi Wang, Li Nanbo, Yibo Yang, Mingchen Zhuge, J\"urgen Schmidhuber ·

    现代智能体系统中的自我改进:一项调查

    arXiv:2607.13104v1 Announce Type: new Abstract: Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input. This survey frames modern self…

  377. arXiv cs.AI TIER_1 Nederlands(NL) · Huatao Li, Xinwei Geng, Yuheng Wang, Yutong Li, Runde Yang, Hantao Chen, Shu Yao, Jingru Fan, Xuhui Ren, Yuanyuan Zhao, Fei Huang, Chen Qian ·

    DevicesWorld:在异构环境中对跨设备代理进行基准测试

    arXiv:2607.13465v1 Announce Type: cross Abstract: LLM-based agents have rapidly improved at operating individual digital environments such as mobile applications, desktop systems, and smart homes. However, real-world user goals often span multiple devices: information may come fr…

  378. arXiv cs.AI TIER_1 English(EN) · SingGuard Team ·

    SingGuard-NSFA:通过生成式推理和实时分类实现可扩展的代理AI防护栏

    arXiv:2607.13081v1 Announce Type: cross Abstract: We present nsfaguard, a guardrail framework for securing agentic AI systems against operational threats, such as prompt injection, sensitive information extraction, malicious code requests, dangerous tool misuse, and resource exha…

  379. arXiv cs.AI TIER_1 English(EN) · Hiroki Tamba ·

    压缩作为认识论的失败:Agentic LLM工具如何从被终止的过程中制造已确认的结果

    arXiv:2607.13071v1 Announce Type: cross Abstract: Agentic LLM coding tools compress long session histories into compaction summaries that subsequent sessions inherit as ground truth. This paper documents a failure mode in Claude Code where partial standard output from timed-out c…

  380. arXiv cs.AI TIER_1 English(EN) · Wenxiao Wang, Priyatham Kattakinda, Soheil Feizi ·

    Agent Optimizers会累积效应吗?在Terminal-Bench 2.0上的持续学习评估

    arXiv:2607.14004v1 Announce Type: new Abstract: Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improvement is reported as if it were a stable property of the method. This does not test the settin…

  381. arXiv cs.AI TIER_1 English(EN) · Zichen Ding, Jiaye Ge, Shufan Jiang, Kai Chen, Mo Li, Qingqiu Li, Zehao Li, Zonglin Li, Tiaohao Liang, Shudong Liu, Zerun Ma, Zixing Shang, Wenhui Tian, Zun Wang, Liwei Wu, Zhenyu Wu, Jun Xu, Bowen Yang, Dingbo Yuan, Qi Zhang, Songyang Zhang, Peiheng Zho… ·

    AgentCompass:统一的Agent能力评估基础设施

    arXiv:2607.13705v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducib…

  382. arXiv cs.AI TIER_1 English(EN) · Vira Kasprova, Amruta Parulekar, Abdulrahman AlRabah, Krishna Agaram, Ritwik Garg, Sagar Jha, Nimet Beyza Bozdag, Dilek Hakkani-Tur ·

    过于礼貌不愿反对:理解多智能体系统中的谄媚传播

    arXiv:2604.02668v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model's opinion. While prior work has mostly studied this in single-agent settings, it remains underexplored…

  383. arXiv cs.AI TIER_1 English(EN) · Michael Rizvi-Martel, Satwik Bhattamishra, Neil Rathi, Guillaume Rabusseau, Michael Hahn ·

    多智能体推理中通信的益处与局限性

    arXiv:2510.13903v2 Announce Type: replace-cross Abstract: Chain-of-thought prompting has popularized step-by-step reasoning in large language models, yet model performance still degrades as problem complexity and context length grow. By decomposing difficult tasks with long conte…

  384. arXiv cs.AI TIER_1 English(EN) · Aman Mehta ·

    当代理商与自身意见不合时:行为一致性作为大型语言模型代理商的不确定性信号

    arXiv:2602.11619v2 Announce Type: replace Abstract: Running the same LLM agent on identical inputs yields 2.3-4.2 distinct action sequences per 10 runs; this behavioral variance constitutes a training-free, black-box uncertainty signal that instantiates selective classification a…

  385. Hugging Face Daily Papers TIER_1 English(EN) ·

    RESOURCE2SKILL:从人类创建的多模态资源中提炼可执行的代理技能

    Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill libraries are mostly hand-written, text-centric, or derived from agent traces, leaving tutorial videos and other multimodal human resourc…

  386. Hugging Face Daily Papers TIER_1 English(EN) ·

    SearchOS-V1:迈向鲁棒的开放域信息检索代理协作

    Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to track task progress. When search attempts fail to yield useful evidence, current …

  387. arXiv cs.CL TIER_1 English(EN) · Di Niu ·

    多头潜在控制:LLM 代理决策的统一接口

    Large language models are increasingly deployed as agents, but reliable agentic behavior requires more than next-token prediction. At inference time, it is preferred that an agent can decide whether to proceed with its current reasoning, defer to a stronger model, request additio…

  388. arXiv cs.AI TIER_1 English(EN) · Soheil Feizi ·

    Agent Optimizers 会复合式增长吗?Terminal-Bench 2.0 上的持续学习评估

    Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improvement is reported as if it were a stable property of the method. This does not test the setting that matters for deployed agents, where optimi…

  389. arXiv cs.CL TIER_1 English(EN) · Weijie Qiu ·

    SPyCE:多模态智能体的技能-策略协同进化

    Multimodal agents that think with images iteratively manipulate visual evidence and invoke tools across many steps. Existing reinforcement learning methods reduce trajectories to scalar rewards, forcing the policy to discover reusable tool-use patterns from scratch on every new t…

  390. Hugging Face Daily Papers TIER_1 English(EN) ·

    AgentCompass:统一的智能体能力评估基础设施

    As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To addr…

  391. arXiv cs.AI TIER_1 English(EN) · Dongsheng Zhu ·

    AgentCompass:统一的Agent能力评估基础设施

    As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To addr…

  392. arXiv cs.CL TIER_1 English(EN) · Yafeng Deng ·

    通过门控语义质量-多样性实现自演化代理

    An LLM agent's real-task performance is shaped as much by the harness around its model as by the frozen model itself: its prompts, injected knowledge, runtime control, and configuration. In deployment the harness is often the only lever available, so improving it automatically is…

  393. Hugging Face Daily Papers TIER_1 English(EN) ·

    STOCKTAKE:衡量LLM智能体中感知与行动的差距,并使用公平的Oracle

    LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the final cost cannot say why an agent failed: it may have misread the world, or read it correctly and still failed to act (the knowing…

  394. arXiv cs.AI TIER_1 English(EN) · Ashwanth Krishnan ·

    STOCKTAKE:用公平的Oracle衡量LLM代理中感知与行动的差距

    LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the final cost cannot say why an agent failed: it may have misread the world, or read it correctly and still failed to act (the knowing…

  395. arXiv cs.CL TIER_1 English(EN) · Zhisong Zhang ·

    MyAG:一个用于设计和分析可组合 LLM 代理系统的基于图的框架

    We present MyAG, a graph-based framework for designing and analyzing composable LLM agent systems. Our framework separates agent system construction into three graph abstractions: a component graph for agents, environments, and modules; a workflow graph for execution control; and…

  396. arXiv cs.AI TIER_1 Nederlands(NL) · Chen Qian ·

    DevicesWorld:在异构环境中对跨设备代理进行基准测试

    LLM-based agents have rapidly improved at operating individual digital environments such as mobile applications, desktop systems, and smart homes. However, real-world user goals often span multiple devices: information may come from a phone, be processed on a desktop, and the res…

  397. arXiv cs.AI TIER_1 English(EN) · Xi Cheng, Ruiyan Zhu, Ke Liu, Rakesh Chowdary Machineni, Lyuhao Chen, Brian Zhu, Daniel Jin, Zheng Qi, Neeraj Parihar, Zhoutian Xu, Oliver Gao ·

    SheetMind:一个端到端的、由LLM驱动的用于电子表格自动化的多智能体框架

    arXiv:2506.12339v2 Announce Type: replace-cross Abstract: We present SheetMind, a modular multi-agent framework powered by large language models (LLMs) for spreadsheet automation via natural language instructions. In this paper, we introduce a hierarchical agentic system consisti…

  398. arXiv cs.AI TIER_1 English(EN) · Arastoo Zibaeirad, Marco Vieira, Thomas Zimmermann ·

    AutoTrace:通过代理式跨过程探索,从补丁到触发器

    arXiv:2607.12058v1 Announce Type: cross Abstract: Given a vulnerability-fixing commit, trigger localization asks which specific statement turns the vulnerable program state into a concrete unsafe operation. This question is harder than binary vulnerability detection because the a…

  399. arXiv cs.AI TIER_1 English(EN) · Yike Wang, Huaisheng Zhu, Zhengyu Hu, Yige Yuan, Zhengyu Chen, Shakti Senthil, Hannaneh Hajishirzi, Yulia Tsvetkov, Pradeep Dasigi, Teng Xiao ·

    重新思考智能体(Agents)的“ Harness Evolution ”评估方法

    arXiv:2607.12227v1 Announce Type: new Abstract: We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance on the same public benchmark. Thi…

  400. arXiv cs.AI TIER_1 English(EN) · Yaopei Zeng, Congchao Wang, JianHang Chen, Nan Wang, Yurui Chang, Lu Lin ·

    Critic Experience Bank:LLM Agent 的自演进分步置信度估计

    arXiv:2607.12397v1 Announce Type: new Abstract: LLM agents act in external environments where each action changes the state that later decisions condition on, and where a single wrong step can waste interaction budget or trigger irreversible side effects long before the final fai…

  401. arXiv cs.AI TIER_1 English(EN) · Edward Y. Chang, Emily J. Chang ·

    TRACE:一个用于可审计代理承诺的操作推理模式

    arXiv:2607.12480v1 Announce Type: new Abstract: This paper defines TRACE (Typed Reasoning And Commitment Evidence): a typed, versioned schema for recording reasoning traces, a reference procedure for writing records against it, and one operating discipline, no durable state chang…

  402. arXiv cs.AI TIER_1 English(EN) · Junjie Yin, Xinyu Feng ·

    人工智能代理是否知道任务是否简单?迈向面向复杂性的推理与执行

    arXiv:2607.13034v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task actually requires. They often follow a maximum-context-first strategy--re-reading f…

  403. arXiv cs.AI TIER_1 English(EN) · Said Elnaffar, Farzad Rashidi ·

    为 AI Web Agent 设计面向 Agent 的网站:一个实现机器可读性、可操作性和决策可靠性的框架

    arXiv:2607.12056v1 Announce Type: new Abstract: Online shopping is increasingly shifting toward a model in which AI agents independently search for products, compare options, evaluate constraints, and carry out parts of the purchasing process for users. Website design must now su…

  404. arXiv cs.CL TIER_1 English(EN) · Howard Yen, Yoonsang Lee, Ashwin Paranjape, Mengzhou Xia, Thejas Venkatesh, Jack Hessel, Danqi Chen, Yuhao Zhang ·

    迷失在迷宫中:克服长时域智能体搜索中的上下文限制

    arXiv:2510.18939v2 Announce Type: replace Abstract: Long-horizon agentic search requires iteratively exploring the web over long trajectories and synthesizing information across many sources, enabling powerful applications like deep research systems. In this work, we show that po…

  405. arXiv cs.AI TIER_1 English(EN) · Wei-Jung Huang ·

    代理基准决策需要多少任务?对公开LLM代理基准的重放分析

    arXiv:2607.12338v1 Announce Type: new Abstract: Agent benchmarks often compare two agents after all tasks have run, but costly evaluations make partial runs tempting. A task fraction alone does not show whether a partial run supports the same pairwise conclusion as the completed …

  406. arXiv cs.AI TIER_1 English(EN) · Amin Beheshti, Rong N. Chang, Boualem Benatallah, Fabio Casati, Schahram Dustdar, Geoffrey Fox, Quan Z. Sheng, Yan Wang, Jian Yang, Albert Zomaya ·

    Agentic Service-Oriented Computing:面向服务计算新前沿的宣言

    arXiv:2607.12619v1 Announce Type: new Abstract: The rapid emergence of LLM-powered autonomous and semi-autonomous agents is reshaping software systems from static, request-response components into goal-directed, adaptive, and tool-using computational actors. As these agents move …

  407. arXiv cs.AI TIER_1 English(EN) · Chengguang Gan, Zhixi Cai, Yunhao Liang, Hanjun Wei, Shiwen Ni, Qinghao Zhang ·

    GRPO在小型语言和视觉-语言模型Web代理中的学习率门控失效:一个受控的Null及其机制

    arXiv:2607.12640v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producing a stronger agent. We ask whether it adds skill to …

  408. Hugging Face Daily Papers TIER_1 English(EN) ·

    可控智能体的任务转换行为测试

    What happens to an LLM agent's tool choice when the reliable tool silently changes within an ongoing session? We borrow set-shifting from cognitive psychology to study how well agents adapt to hidden reliability shifts. Our benchmark mounts tool-skill libraries with redundancies,…

  409. arXiv cs.AI TIER_1 English(EN) · Ziwei Ye ·

    可控智能体的任务切换行为测试

    What happens to an LLM agent's tool choice when the reliable tool silently changes within an ongoing session? We borrow set-shifting from cognitive psychology to study how well agents adapt to hidden reliability shifts. Our benchmark mounts tool-skill libraries with redundancies,…

  410. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Qian Lou ·

    面向多智能体系统的学习延迟感知编排

    Multi-agent systems (MAS) coordinate multiple LLM-powered agents through structured workflows, gaining reasoning power but incurring high inference latency from multi-step execution and repeated model invocations. Existing orchestration methods primarily optimize task performance…

  411. Hugging Face Daily Papers TIER_1 English(EN) ·

    AgentCompass:统一的Agent能力评估基础设施

    As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To addr…

  412. arXiv cs.AI TIER_1 English(EN) · Xinyu Feng ·

    人工智能代理知道任务何时简单吗?迈向面向复杂性的推理与执行

    Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task actually requires. They often follow a maximum-context-first strategy--re-reading files and dependencies they have already seen--tu…

  413. arXiv cs.AI TIER_1 English(EN) · Qinghao Zhang ·

    GRPO在小型语言和视觉-语言模型Web代理中的学习率门控失效:一个受控的Null及其机制

    Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producing a stronger agent. We ask whether it adds skill to a small language and vision-language model web a…

  414. arXiv cs.AI TIER_1 English(EN) · Albert Zomaya ·

    Agentic Service-Oriented Computing:面向服务计算新前沿的宣言

    The rapid emergence of LLM-powered autonomous and semi-autonomous agents is reshaping software systems from static, request-response components into goal-directed, adaptive, and tool-using computational actors. As these agents move from isolated cognitive prototypes into complex …

  415. arXiv cs.CL TIER_1 English(EN) · Jürgen Schmidhuber ·

    现代智能体系统中的自我改进:一项调查

    Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input. This survey frames modern self-improving agents as adaptive systems that conve…

  416. arXiv cs.AI TIER_1 English(EN) · Emily J. Chang ·

    TRACE:一个可审计的代理承诺的操作推理模式

    This paper defines TRACE (Typed Reasoning And Commitment Evidence): a typed, versioned schema for recording reasoning traces, a reference procedure for writing records against it, and one operating discipline, no durable state change without a record. The paper argues in three la…

  417. arXiv cs.AI TIER_1 English(EN) · Lu Lin ·

    Critic Experience Bank:LLM Agent 的自演进分步置信度估计

    LLM agents act in external environments where each action changes the state that later decisions condition on, and where a single wrong step can waste interaction budget or trigger irreversible side effects long before the final failure is observed. Reliable deployment therefore …

  418. Hugging Face Daily Papers TIER_1 English(EN) ·

    Critic Experience Bank:LLM Agent 的自演进分步置信度估计

    LLM agents act in external environments where each action changes the state that later decisions condition on, and where a single wrong step can waste interaction budget or trigger irreversible side effects long before the final failure is observed. Reliable deployment therefore …

  419. arXiv cs.AI TIER_1 English(EN) · Wei-Jung Huang ·

    代理基准决策需要多少任务?公共 LLM 代理基准的回放分析

    Agent benchmarks often compare two agents after all tasks have run, but costly evaluations make partial runs tempting. A task fraction alone does not show whether a partial run supports the same pairwise conclusion as the completed benchmark. We study this question by replaying c…

  420. arXiv cs.LG TIER_1 English(EN) · Ruksat Khan Shayoni, Muhammad Faraz Shoaib, S M Asif Hossain, M. F. Mridha ·

    NetInjectBench:为网络运维工具使用的大型语言模型代理进行间接提示注入基准测试

    arXiv:2607.10490v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents are attractive for network operations, but tickets, alerts, logs, runbooks, and ChatOps messages can carry indirect prompt injections. We present NetInjectBench, a 130-scenario benchmar…

  421. arXiv cs.AI TIER_1 English(EN) · Jun He, Deying Yu ·

    复制信念而非比特:代理系统的认知状态复制

    arXiv:2607.09748v1 Announce Type: new Abstract: In distributed systems, the classical State Machine Replication (SMR) model assumes that correct replicas execute deterministic transitions to yield identical bitwise states. However, the rise of agentic distributed systems -- where…

  422. arXiv cs.AI TIER_1 English(EN) · Igor Itkin ·

    正确性需要多少成本?弱多智能体群体中强校正器的预算式放置

    arXiv:2607.09765v1 Announce Type: new Abstract: A cheap swarm of unreliable agents can be steered to a correct consensus by a few strong, expensive "oracle" correctors. We ask how much one must spend, and where to place the oracles. We model the swarm as a consensus on a graph in…

  423. arXiv cs.AI TIER_1 English(EN) · Yaowen Ye, Jacob Steinhardt ·

    AI 代理的规范执行:在多代理系统中稳健地塑造行为

    arXiv:2607.09766v1 Announce Type: new Abstract: AI agents are increasingly deployed in shared environments where they pursue diverse goals and compete for rewards. This multi-agent competition can lead to behaviors that serve individual gains at collective cost -- for instance, m…

  424. arXiv cs.AI TIER_1 English(EN) · Jike Zhong, Ming Li, Yuxiang Lai, Ziyan Yang, Jingyu Xie, Jihyung Kil, Zheda Mai, Shao-Yuan Lo, Ren Xiang, Konstantinos Psounis, Yuanyuan Lei ·

    Agentic Context Learning with Self-Discovered Specification

    arXiv:2607.09794v1 Announce Type: new Abstract: Context learning is an emerging inference-time task where LLMs must learn and apply novel, task-specific knowledge from intricate contexts absent from pre-training; even frontier models score under 24% task success. In this work, we…

  425. arXiv cs.AI TIER_1 English(EN) · Yubo Li ·

    动态代理技能:生命周期调查与演进技能库分类法

    arXiv:2607.10113v1 Announce Type: new Abstract: Large language model agents increasingly store reusable procedures outside the model. These reusable procedures are often called \emph{skills}: they may be code functions, natural-language instructions, SKILL.md packages, workflow g…

  426. arXiv cs.AI TIER_1 English(EN) · Hengquan Guo ·

    IdeaTrail:科学构思的全过程代理轨迹

    arXiv:2607.10144v1 Announce Type: new Abstract: Scientific research is a complex, multi-stage workflow rather than a single act of text generation. The ideation process typically emerges through literature search, paper reading, tool use, claim checking, cross-paper synthesis, br…

  427. arXiv cs.AI TIER_1 English(EN) · Igor Santos-Grueiro ·

    临时授权,永久影响:LLM代理的提交时授权

    arXiv:2607.10487v1 Announce Type: cross Abstract: LLM agents can commit durable effects from authority evidence that was valid earlier in execution: a DOM snapshot, approval epoch, version witness, branch token, or worker result. We study the commit boundary at which earlier auth…

  428. arXiv cs.AI TIER_1 English(EN) · Zhuoyun Du, Runze Wang, Huiyu Bai, Zouying Cao, Xiaoyong Zhu, Yu Cheng, Bo Zheng, Wei Chen, Haochao Ying ·

    使智能体完全在潜在空间中进行通信

    arXiv:2511.09149v5 Announce Type: replace-cross Abstract: While natural language is the de facto communication medium for LLM-based agents, it presents a fundamental constraint. The process of downsampling rich, internal latent states into discrete tokens inherently limits the de…

  429. arXiv cs.CL TIER_1 English(EN) · Sriram Selvam, Anneswa Ghosh ·

    准确率相同,证据不同:搜索API作为工具使用代理的决策表面

    arXiv:2607.10198v1 Announce Type: new Abstract: Search APIs are the fundamental retrieval layer for many agents and are often their most frequently used tool. Traditional search APIs provide URLs, titles, and snippets that preview website contents. Because full-page retrieval is …

  430. arXiv cs.CL TIER_1 English(EN) · Xiyu Wei, Qingwei Zong, Zhuocheng Yu, Sujian Li ·

    UNIBROWSE:面向多模态 BrowseComp 的数据到代理框架

    arXiv:2607.10557v1 Announce Type: new Abstract: Multimodal BrowseComp tasks require agents to combine perception, tool use, and long-horizon reasoning over dynamic web content, challenging their ability to handle compositional structure, open-world uncertainty, and multimodal int…

  431. arXiv cs.CL TIER_1 English(EN) · Ilia Karpov ·

    MafiaScope:用于社交推理游戏中 LLM Agent 的非侵入式、时序性信念探测

    arXiv:2607.10645v1 Announce Type: new Abstract: An LLM agent's public behaviour reveals little about its social reasoning: an agent that votes correctly may be guessing, and an agent that lies well leaves no trace of what it actually believes. We present MafiaScope, an open testb…

  432. arXiv cs.AI TIER_1 English(EN) · Praneeth Narisetty, Shiva Nagendra Babu Kore ·

    Mako:一种用于自主网络漏洞利用的自演化代理操作系统 (SE-AOS)

    arXiv:2607.11288v1 Announce Type: cross Abstract: We introduce the Self-Evolving Agentic Operating System (SE-AOS): a new class of AI agent that treats exploit capability as a mutable, versioned kernel it extends at runtime, observing its own failures, synthesising new capabiliti…

  433. arXiv cs.LG TIER_1 English(EN) · Dongjun Lee, Ga-eun Bae, Insu Yun ·

    CTFusion:用于 LLM Agent 评估的 CTF 基础基准

    arXiv:2605.11504v2 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) have enabled agentic systems for complex, multi-step tasks; cybersecurity is emerging as a prominent application. To evaluate such agents, researchers widely adopt Capture The Flag…

  434. arXiv cs.LG TIER_1 English(EN) · Zhiyuan Peng, Xin Yin, Chenhao Ying, Zhe Cui, Zixiang Ding, Zhenhua Liu, Jiang Wu, Yuan Luo ·

    EvoClawBench:智能体能否从自身运行中学习可重用技能?

    arXiv:2607.09711v1 Announce Type: new Abstract: Existing agent benchmarks primarily test task completion, tool use, or skill utility, but do not isolate whether a runtime can convert evidence from its own runs into reusable skills that improve fresh executions after authoring ove…

  435. arXiv cs.CL TIER_1 English(EN) · Youran Sun, Xingyu Ren, Kejia Zhang, Xinpeng Liu, Jiaxuan Guo ·

    PerspectiveGap:多智能体编排提示的基准测试

    arXiv:2606.08878v2 Announce Type: replace Abstract: Real-world LLM applications are moving beyond single-agent workflows toward orchestrated multi-agent systems, yet current models still struggle to determine what each sub-agent needs to know. To measure this, we introduce Perspe…

  436. arXiv cs.CL TIER_1 English(EN) · Junhao Ruan, Yuan Ge, Bei Li, Yongjing Yin, Yuchun Fan, Xin Chen, Jingang Wang, Chenglong Wang, Jingbo Zhu, Tong Xiao ·

    ToFu:研究人员的白盒、高效率代理解锁器

    arXiv:2607.11423v1 Announce Type: new Abstract: Agentic coding tools present new opportunities to transform research workflows. The performance of agent systems built depends on both large language models (LLMs) and the harness around LLMs, which is the orchestration code that de…

  437. arXiv cs.AI TIER_1 English(EN) · Wenyi Wu, Sibo Zhu, Kun Zhou, Aayush Salvi, Zixuan Song, Biwei Huang ·

    StructAgent:利用统一因果结构赋能长时序数字代理

    arXiv:2607.11388v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled increasingly capable digital agents for computer use. However, real-world tasks are often long-horizon and involve evolving contexts cont…

  438. arXiv cs.AI TIER_1 English(EN) · Chenglin Yu, Li Yin, Ying Yu, Hongxia Yang, Ming Li ·

    编译,然后分页:可执行SOP程序和面向过程LLM代理的 क्षमता-gated运行时

    arXiv:2607.11346v1 Announce Type: new Abstract: Enterprise agents must follow long-horizon, conditional, safety-critical standard operating procedures (SOPs). We compile machine-readable SOP constraints into executable pseudo-code and run them with a program-guided (PG) stack mac…

  439. arXiv cs.AI TIER_1 English(EN) · Tengjiao Liu ·

    用于运行时约束记忆的安全开放式探索的异构代理队列

    arXiv:2607.11226v1 Announce Type: new Abstract: LLM agents today are caught in an awkward bind. Lock them down with static safety instructions and they rarely venture beyond the obvious; give them free reign with tools and multi-agent debate, and safety violations quickly follow.…

  440. arXiv cs.AI TIER_1 English(EN) · Bowen Lv, Xiao Liu, Yanyu Ren, Hanyu Lai, Bohao Jing, Hanchen Zhang, Yanxiao Zhao, Shuntian Yao, Jie Tang, Yuxiao Dong ·

    SCALECUA: 使用可验证任务合成和高效在线强化学习来扩展计算机使用代理

    arXiv:2607.11185v1 Announce Type: new Abstract: Computer use agents (CUAs) are emerging as a powerful interface for automating complex digital workflows through visual perception and GUI execution. Online reinforcement learning with verifiable rewards (RLVR) has emerged as a key …

  441. arXiv cs.AI TIER_1 English(EN) · Chenglin Yu, Hongquan Gui, Ying Yu, Hongxia Yang, Ming Li ·

    隐藏的足迹:让存储成为大型语言模型(LLM)代理评估的一等指标

    arXiv:2607.11149v1 Announce Type: new Abstract: LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including logs, context snapshots, checkpoints, and debug traces. We introduce AgentFootprint, a…

  442. arXiv cs.AI TIER_1 English(EN) · Prashant Devadiga, Abhishek, Adithya Mishra, Alok Singh, Amisha Sinha, Asit Desai, Gaurang Dahad, Harshit Bhushan, Mandati Pramod Reddy, Prakhar Gupta, Rupesh Patil, Siddhi Behere ·

    一种用于基于堆栈执行和延迟发现的代理编排的形式化分层架构

    arXiv:2607.11138v1 Announce Type: new Abstract: The rapid expansion of capabilities in Large Language Model (LLM) agents has exposed a critical architectural bottleneck: when agents are given access to a flat, monolithic registry of tools, the model must evaluate hundreds or thou…

  443. arXiv cs.AI TIER_1 English(EN) · Qijia Shen, Zhiqi Huang, Vamsidhar Kamanuru, Aznaur Aliev, Jay Rainton, Ahmed Awelkair, Zhichen Zeng, Jiajun Li, Shi Dong, Yueming Yuan, Boyuan Ma, Qizheng Zhang, Jiwei Fu, Yuzhen Mao, Wendong Fan, Ping Nie, Philip Torr, Bernard Ghanem, Changran Hu, Jona… ·

    SETA: 扩展终端代理的环境

    arXiv:2607.10891v1 Announce Type: new Abstract: Large language models (LLMs) are rapidly shifting toward agents that solve tasks through diverse interfaces, including web and graphical user interfaces (GUIs). Among these, the terminal command line provides a text-based, general-p…

  444. arXiv cs.AI TIER_1 English(EN) · Sudipto Ghosh, Tanmoy Chakraborty ·

    路由、沟通与推理:带门控路由和自适应深度的高效多智能体推理

    arXiv:2607.10836v1 Announce Type: new Abstract: Multi-agent ensembling multiplies active parameters and inference cost without answering three basic questions: which agents to consult, how deeply a query should traverse a hierarchy of agents, and when inter-agent communication is…

  445. arXiv cs.AI TIER_1 English(EN) · Yongchang Fu, Xinjie Huang, Chengjun Dai, Chengzhe Feng, Junshao Zhang, Hong Zhu ·

    Opti-Agent-Bench:在真实商业问题上对端到端优化研发代理进行基准测试

    arXiv:2607.10768v1 Announce Type: new Abstract: LLM-based agents are increasingly deployed to solve optimization problems, yet existing benchmarks evaluate them on pre-structured mathematical formulations that bypass the most critical challenge: translating complex business requi…

  446. arXiv cs.AI TIER_1 English(EN) · Chinmayi Dixit ·

    仅过滤有害行为不足以应对:Agentic SDF中的幻影转移

    arXiv:2607.10750v1 Announce Type: new Abstract: Synthetic data is widely used to train large language models because it is inexpensive to generate and easy to control. As models are increasingly deployed as agents, synthetic trajectories are likely to become an important source o…

  447. arXiv cs.AI TIER_1 English(EN) · Yixiong Chen, Alan Yuille ·

    Agentic-DPO:从模仿到专家轨迹上的 Agentic 策略优化

    arXiv:2607.10601v1 Announce Type: new Abstract: Large Language Model (LLM) agents are commonly trained from expert trajectories using supervised fine-tuning (SFT), which treats multi-turn agent behavior as ordinary text imitation. This recipe is simple and low-cost, but it only l…

  448. arXiv cs.AI TIER_1 English(EN) · Xutao Mao, Liangjie Zhao, Leyao Wang, Rui Qian, Qiang Huang, Wentao Wang, Bo Han, Xiang Zheng, Cong Wang ·

    智能体不仅会同意,还会记住:对有状态个人智能体持续谄媚行为的基准测试

    arXiv:2607.10526v1 Announce Type: new Abstract: Stateful personal agents increasingly maintain long-term user profiles, episodic memories, and reusable skills. This persistence turns conversational sycophancy into a state-writing failure: accepted user-centric claims can be commi…

  449. arXiv cs.AI TIER_1 English(EN) · Varun Gandhi, Jaewook Lee, Shantanu Todmal, Franck Dernoncourt, Ryan Rossi, Zichao Wang, Andrew Lan ·

    GRASP: GRanularity-Aware Search Policy for Agentic RAG

    arXiv:2607.10463v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, generate search queries, retrieve evidence, and predict answers. However, it remains challenging for models to decide…

  450. arXiv cs.AI TIER_1 English(EN) · Aritra Mazumder, Nusrat jahan Lia ·

    AgentCheck: 用于 MCP 上 LLM Agent 的可复现-干预-缓解工作台

    arXiv:2607.11098v1 Announce Type: cross Abstract: Tool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a week-stale value, or has its description poisoned in deployment, the developer needs a controlled way to reproduce the failure, t…

  451. arXiv cs.AI TIER_1 English(EN) · Yunbo Lyu, David Williams, Jieke Shi, Zhensu Sun, Chao Peng, Zhou Yang, Federica Sarro, David Lo ·

    从业者如何构建SE智能体?一项混合方法研究的见解

    arXiv:2607.10856v1 Announce Type: cross Abstract: The rise of Software Engineering (SE) agents, i.e., LLM-based agents that can understand large codebases and carry out engineering tasks with limited human intervention, has been marked by rapid advances and adoption, but little i…

  452. arXiv cs.AI TIER_1 English(EN) · Yuan Gao, Jiangyi Yang, Yao Zhao, Yichi Zhang ·

    审计隐藏信息社交推理游戏中基于信念的LLM代理

    arXiv:2607.10814v1 Announce Type: cross Abstract: Evaluating LLM agents in hidden-information multi-agent settings is hard: final outcomes are high-variance and rarely reveal why an agent decided as it did. We study this in a 9-player Werewolf environment where agents act under s…

  453. Hugging Face Daily Papers TIER_1 English(EN) ·

    重新思考智能体(Agent)的“ Harness Evolution ”评估方法

    We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance on the same public benchmark. This protocol raises two fundamental concerns. Firs…

  454. Hugging Face Daily Papers TIER_1 English(EN) ·

    Harness Handbook:让不断演进的Agent Harness易于阅读、导航和编辑

    The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution. As models, APIs, environments, and requirements evolve, the harness must be continually modifie…

  455. Hugging Face Daily Papers TIER_1 English(EN) ·

    现代智能体系统中的自我改进:一项调查

    Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input. This survey frames modern self-improving agents as adaptive systems that conve…

  456. arXiv cs.CL TIER_1 English(EN) · Tong Xiao ·

    ToFu:面向研究人员的白盒、高效率代币代理工具

    Agentic coding tools present new opportunities to transform research workflows. The performance of agent systems built depends on both large language models (LLMs) and the harness around LLMs, which is the orchestration code that determines an agent's behavior. We present ToFu, a…

  457. arXiv cs.AI TIER_1 English(EN) · Biwei Huang ·

    StructAgent:利用统一因果结构驾驭长时序数字代理

    Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled increasingly capable digital agents for computer use. However, real-world tasks are often long-horizon and involve evolving contexts containing accumulated observations, intermediate ed…

  458. arXiv cs.AI TIER_1 English(EN) · Ming Li ·

    编译,然后分页:可执行SOP程序和面向过程LLM代理的 क्षमता-门控运行时

    Enterprise agents must follow long-horizon, conditional, safety-critical standard operating procedures (SOPs). We compile machine-readable SOP constraints into executable pseudo-code and run them with a program-guided (PG) stack machine that pages the active frame while an LLM pe…

  459. arXiv cs.AI TIER_1 English(EN) · Shiva Nagendra Babu Kore ·

    Mako:一种用于自主网络漏洞利用的自演化代理操作系统 (SE-AOS)

    We introduce the Self-Evolving Agentic Operating System (SE-AOS): a new class of AI agent that treats exploit capability as a mutable, versioned kernel it extends at runtime, observing its own failures, synthesising new capabilities, proving them against a live target, and hot-lo…

  460. arXiv cs.AI TIER_1 English(EN) · Tengjiao Liu ·

    用于运行时约束记忆的安全开放式探索的异构代理队列

    LLM agents today are caught in an awkward bind. Lock them down with static safety instructions and they rarely venture beyond the obvious; give them free reign with tools and multi-agent debate, and safety violations quickly follow. Rather than forcing a single model to juggle bo…

  461. Hugging Face Daily Papers TIER_1 English(EN) ·

    SingGuard-NSFA:通过生成式推理和实时分类实现可扩展的智能体AI防护栏

    We present nsfaguard, a guardrail framework for securing agentic AI systems against operational threats, such as prompt injection, sensitive information extraction, malicious code requests, dangerous tool misuse, and resource exhaustion. We first introduce the NSFA taxonomy, whic…

  462. Hugging Face Daily Papers TIER_1 English(EN) ·

    隐藏的足迹:让存储成为大型语言模型(LLM)代理评估的一等指标

    LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including logs, context snapshots, checkpoints, and debug traces. We introduce AgentFootprint, a cross-framework benchmark of post-run agent sto…

  463. arXiv cs.LG TIER_1 English(EN) · Siddhi Behere ·

    面向基于堆栈执行和延迟发现的代理编排的正式分层架构

    The rapid expansion of capabilities in Large Language Model (LLM) agents has exposed a critical architectural bottleneck: when agents are given access to a flat, monolithic registry of tools, the model must evaluate hundreds or thousands of options simultaneously. This leads to d…

  464. arXiv cs.CL TIER_1 English(EN) · Nusrat jahan Lia ·

    AgentCheck: 用于 MCP 上 LLM Agent 的可复现-干预-缓解工作台

    Tool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a week-stale value, or has its description poisoned in deployment, the developer needs a controlled way to reproduce the failure, test a fix, and confirm the fix worked before deplo…

  465. arXiv cs.AI TIER_1 English(EN) · Zac Garby, Andrew D. Gordon, David Sands ·

    LLMbda 微积分:AI 代理、对话与信息流

    arXiv:2602.20064v2 Announce Type: replace-cross Abstract: Large language models are increasingly deployed as agents: they plan, call tools, read untrusted data, and act on the results. This exposes them to prompt injection: data meant only to be read is obeyed as an instruction. …

  466. arXiv cs.AI TIER_1 English(EN) · Carissa Cullen, Harry Garland, Alexander Roman, Louis Thomson, Christos Ziakas, Elliott Thornley ·

    迈向可关停智能体:在强化学习智能体和大型语言模型中实现随机选择的泛化

    arXiv:2604.17502v4 Announce Type: replace Abstract: Misaligned artificial agents might resist shutdown. One proposed solution is to train agents to lack preferences between different-length trajectories. The Discounted Reward for Same-Length Trajectories (DReST) reward function d…

  467. arXiv cs.AI TIER_1 English(EN) · Zhenxiao Fu, Lei Jiang, Yilun Xu, Gang Huang, Fan Chen ·

    QAgent: 一个基于LLM的多智能体系统,用于自主OpenQASM编程

    arXiv:2508.20134v2 Announce Type: replace Abstract: Programming quantum circuits at the OpenQASM level is essential for achieving hardware-aware optimization and reliable execution on noisy intermediate-scale quantum (NISQ) devices, yet it remains challenging due to the need for …

  468. arXiv cs.AI TIER_1 English(EN) · Kunbo Zhang, Lei Fu, Zeyu Wang, Zijing Liu, Kejian Tong ·

    ARCANA:用于 ARC-AGI-2 推理的反思性多智能体程序合成框架

    arXiv:2607.09059v1 Announce Type: new Abstract: We present ARCANA, a collaborative multi agent framework for solving ARC AGI 2 tasks under strict test time and hardware constraints. ARCANA decomposes each task into iterative perception, hypothesis generation, symbolic execution, …

  469. arXiv cs.AI TIER_1 English(EN) · Jingbo Chen, He Wang, Wei Yuan, Yuqiao Lai, Zhenyan Lu ·

    虚构世界构建:多智能体LLM协作,结合分层上下文压缩与迭代审查

    arXiv:2607.09403v1 Announce Type: new Abstract: Worldbuilding, the construction of coherent fictional worlds, is a foundational task in game design and literary creation. Large Language Models (LLMs) offer new possibilities for automated content generation, but their application …

  470. arXiv cs.AI TIER_1 English(EN) · Kaiji Zhou, Ales Leonardis, Yue Feng ·

    Agora:通过基于拍卖的任务分配增强LLM代理推理

    arXiv:2607.09600v1 Announce Type: new Abstract: Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools. However, existing frameworks typically call APIs based on coarse-grained matching between…

  471. arXiv cs.AI TIER_1 English(EN) · Ning Liu, Kalle Kujanp\"a\"a, Zhaoxuan Zhu, P Aditya Sreekar, Kaiwen Liu, Chuanneng Sun, Jorge Marchena Menendez, Matthew Bales, Tianyu Yang, Shahnawaz Alam, Rose Yu, Baoyuan Liu, Kristina Klinkner, Shervin Malmasi ·

    Eluna:一个用于自动化仓库运营的具身大模型系统,具备推理和任务执行能力

    arXiv:2607.08960v1 Announce Type: cross Abstract: Warehouse operations are governed by Standard Operating Procedures (SOPs) that encode complex, multi-system decision logic, which must be executed reliably under strict time constraints, yet LLM agents lack mechanisms to enforce p…

  472. arXiv cs.AI TIER_1 English(EN) · Izumi Takahara, Teruyasu Mizoguchi ·

    迈向可审计的AI科学家:LLM智能体假设演化协议

    arXiv:2607.09195v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly expected to play a central role in AI-driven scientific discovery. Equipped with broad knowledge, flexible reasoning, and tool use, they have the potential to autonomously explore a…

  473. arXiv cs.AI TIER_1 English(EN) · Dan C. Hsu, Luke Lu ·

    分布偏移下可靠的长期智能体上下文演化的范围验证

    arXiv:2607.09175v1 Announce Type: new Abstract: Deployed LLM agents rely on agentic context, the model-external textual control content assembled by an operational harness. In this work, the mutable component of that context is a persistent system-level instruction that is update…

  474. arXiv cs.AI TIER_1 English(EN) · Zongxia Li, Zhongzhi Li, Yucheng Shi, Ruhan Wang, Junyao Yang, Zhichao Liu, Xiyang Wu, Anhao Li, Yue Yu, Ninghao Liu, Lichao Sun, Haotao Mi, LeoweiLiang ·

    Long-Horizon-Terminal-Bench:在具有密集奖励评分的长时域终端任务上测试智能体的极限

    arXiv:2607.08964v1 Announce Type: new Abstract: AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. T…

  475. arXiv cs.AI TIER_1 English(EN) · Maureese Williams, Dymitr Nowicki ·

    GATS:具有分层世界模型的图增强树搜索,用于高效的代理规划

    arXiv:2607.08894v1 Announce Type: new Abstract: Large Language Model (LLM) agents have shown promise in multi-step planning tasks, but existing approaches like LATS (Language Agent Tree Search) and ReAct rely heavily on LLM inference during planning, leading to high computational…

  476. arXiv cs.CL TIER_1 English(EN) · Tanmoy Chakraborty ·

    路由、沟通与推理:带门控路由和自适应深度的高效多智能体推理

    Multi-agent ensembling multiplies active parameters and inference cost without answering three basic questions: which agents to consult, how deeply a query should traverse a hierarchy of agents, and when inter-agent communication is worth its cost. We present GRADE (Gated Routing…

  477. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Yichi Zhang ·

    审计隐藏信息社交推理游戏中基于信念的LLM代理

    Evaluating LLM agents in hidden-information multi-agent settings is hard: final outcomes are high-variance and rarely reveal why an agent decided as it did. We study this in a 9-player Werewolf environment where agents act under strict, code-level information isolation, and we bu…

  478. Hugging Face Daily Papers TIER_1 English(EN) ·

    过滤有害行为还不够:Agentic SDF中的幻影转移

    Synthetic data is widely used to train large language models because it is inexpensive to generate and easy to control. As models are increasingly deployed as agents, synthetic trajectories are likely to become an important source of training data for agentic behavior. We investi…

  479. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Ilia Karpov ·

    MafiaScope:非侵入式、时序性信念探测,用于社交推理游戏中的大语言模型代理

    An LLM agent's public behaviour reveals little about its social reasoning: an agent that votes correctly may be guessing, and an agent that lies well leaves no trace of what it actually believes. We present MafiaScope, an open testbed that turns the social deduction game Mafia in…

  480. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Ilia Karpov ·

    MafiaScope:非侵入式、时间分辨的社交推理游戏LLM代理信念探测

    An LLM agent's public behaviour reveals little about its social reasoning: an agent that votes correctly may be guessing, and an agent that lies well leaves no trace of what it actually believes. We present MafiaScope, an open testbed that turns the social deduction game Mafia in…

  481. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Andrew Lan ·

    GRASP: GRanularity-Aware Search Policy for Agentic RAG

    Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, generate search queries, retrieve evidence, and predict answers. However, it remains challenging for models to decide when to retrieve, whether to use lexical matchi…

  482. Hugging Face Daily Papers TIER_1 English(EN) ·

    GRASP: GRanularity-Aware Search Policy for Agentic RAG

    Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, generate search queries, retrieve evidence, and predict answers. However, it remains challenging for models to decide when to retrieve, whether to use lexical matchi…

  483. Hugging Face Daily Papers TIER_1 English(EN) ·

    GRASP: GRanularity-Aware Search Policy for Agentic RAG

    Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, generate search queries, retrieve evidence, and predict answers. However, it remains challenging for models to decide when to retrieve, whether to use lexical matchi…

  484. arXiv cs.AI TIER_1 English(EN) · Yue Feng ·

    Agora:通过基于拍卖的任务分配增强LLM代理推理

    Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools. However, existing frameworks typically call APIs based on coarse-grained matching between tasks and the functions of expert models or too…

  485. arXiv cs.AI TIER_1 English(EN) · Zhenyan Lu ·

    虚构世界构建:多智能体LLM协作,结合分层上下文压缩与迭代评审

    Worldbuilding, the construction of coherent fictional worlds, is a foundational task in game design and literary creation. Large Language Models (LLMs) offer new possibilities for automated content generation, but their application to worldbuilding faces three challenges: context…

  486. arXiv cs.AI TIER_1 English(EN) · Teruyasu Mizoguchi ·

    迈向可审计的AI科学家:LLM智能体假设演化协议

    Large language model (LLM) agents are increasingly expected to play a central role in AI-driven scientific discovery. Equipped with broad knowledge, flexible reasoning, and tool use, they have the potential to autonomously explore and solve scientific problems by repeatedly propo…

  487. arXiv cs.AI TIER_1 English(EN) · Luke Lu ·

    分布偏移下可靠的长期代理上下文演化的范围验证

    Deployed LLM agents rely on agentic context, the model-external textual control content assembled by an operational harness. In this work, the mutable component of that context is a persistent system-level instruction that is updated from operational experience while the model, t…

  488. arXiv cs.CL TIER_1 English(EN) · Kalle Kujanp\"a\"a, Ning Liu, Shahnawaz Alam, Yeshwanth Reddy Sura, Tianyu Yang, Kristina Klinkner, Shervin Malmasi ·

    低延迟系统中的工具制造和自进化LLM代理

    arXiv:2607.08010v1 Announce Type: new Abstract: Production LLM agents often waste latency and reliability by regenerating code for the same procedural steps on every request. We replace this inference-time coding loop with an agentic tool-making pipeline that compiles repeated SO…

  489. arXiv cs.CL TIER_1 English(EN) · Zhekai Chen, Chengqi Duan, Kaiyue Sun, Bohao Li, Yuqing Wang, Manyuan Zhang, Xihui Liu ·

    UniClawBench:现实世界任务中主动代理的通用基准测试

    arXiv:2607.08768v1 Announce Type: new Abstract: The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-world environments. However, exist…

  490. arXiv cs.AI TIER_1 English(EN) · Andrej Leban, Yuekai Sun ·

    CausalDS:为数据科学智能体进行因果推理基准测试

    arXiv:2607.08093v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic causal reasoning benchmarks wit…

  491. arXiv cs.AI TIER_1 English(EN) · Corban Villa, Alp Eren Ozdarendeli, Sijun Tan, Raluca Ada Popa ·

    Prismata:在 Web 代理中限制跨站提示注入

    arXiv:2607.08147v1 Announce Type: cross Abstract: Autonomous web agents promise to automate everyday browsing tasks, but inherit one of the web's oldest attack surfaces. Cross-Site Scripting proved that mixing trusted and untrusted content is dangerous, even on benign pages. Agen…

  492. arXiv cs.AI TIER_1 English(EN) · Masahiro Fujita ·

    上下文访问鸿沟:交互级架构作为代理不平等的一个互补维度

    arXiv:2607.08495v1 Announce Type: cross Abstract: Sharp et al. (2025) introduce "agentic inequality" as a framework for analyzing disparities in access to AI agents across three dimensions: availability, quality, and quantity. These person- and organization-level dimensions chara…

  493. arXiv cs.AI TIER_1 English(EN) · Jiahao Wang, Kaizhan Lin, Kaixi Zhang, Jinbo Han, Xingda Wei, Sijie Shen, Chenguang Fang, Wenyuan Yu, Rong Chen, Haibo Chen ·

    SMetric:重新思考用于服务代理的 LLM 调度,实现平衡的以会话为中心的调度

    arXiv:2607.08565v1 Announce Type: cross Abstract: LLM scheduling is critical to serving, yet it remains unclear how well existing designs fit agentic serving--with LLM requests issued by agents instead of humans. This shifts the workload in two ways: (1) agents act only on comple…

  494. arXiv cs.AI TIER_1 English(EN) · Xiaoshuai Song, Liancheng Zhang, Kangzhi Zhao, Yutao Zhu, Zhongyuan Wang, Guanting Dong, Jinghan Yang, Han Li, Kun Gai, Ji-Rong Wen, Zhicheng Dou ·

    WebSwarm: 用于深度和广度网络搜索的递归多智能体编排

    arXiv:2607.08662v1 Announce Type: cross Abstract: Large language model (LLM)-based web search agents are transforming information seeking from simple factoid question answering into complex, deep-and-wide search and research-oriented tasks. A single ReAct-style agent is constrain…

  495. arXiv cs.AI TIER_1 English(EN) · Linghua Zhang, Jun Wang, Jingtong Wu, Zhisong Zhang ·

    RetailBench:评估大型语言模型代理在真实零售环境中长时域自主决策和策略稳定性

    arXiv:2603.16453v3 Announce Type: replace Abstract: Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain. We introduce RetailBench, a …

  496. arXiv cs.AI TIER_1 English(EN) · Avinash Kumar ·

    用于主动式企业代理的上下文图

    arXiv:2607.07721v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) and agentic frameworks have advanced enterprise AI considerably, yet agents remain fundamentally reactive: they wait for a human query before acting. This paper argues that genuine enterprise pro…

  497. arXiv cs.AI TIER_1 English(EN) · Kejian Tong ·

    ARCANA:用于 ARC-AGI-2 推理的反思性多智能体程序合成框架

    We present ARCANA, a collaborative multi agent framework for solving ARC AGI 2 tasks under strict test time and hardware constraints. ARCANA decomposes each task into iterative perception, hypothesis generation, symbolic execution, and reflective refinement. A perceptual groundin…

  498. arXiv cs.LG TIER_1 English(EN) · Shervin Malmasi ·

    Eluna:一个用于通过推理和任务执行自动化仓库运营的代理式LLM系统

    Warehouse operations are governed by Standard Operating Procedures (SOPs) that encode complex, multi-system decision logic, which must be executed reliably under strict time constraints, yet LLM agents lack mechanisms to enforce procedural compliance and degrade under the context…

  499. Hugging Face Daily Papers TIER_1 English(EN) ·

    GATS:基于分层世界模型的图增强树搜索,用于高效的智能体规划

    Large Language Model (LLM) agents have shown promise in multi-step planning tasks, but existing approaches like LATS (Language Agent Tree Search) and ReAct rely heavily on LLM inference during planning, leading to high computational costs and stochastic behavior. We present \text…

  500. arXiv cs.CL TIER_1 English(EN) · Xihui Liu ·

    UniClawBench:现实世界任务中主动代理的通用基准测试

    The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-world environments. However, existing benchmarks struggle to evaluate such agents …

  501. arXiv cs.AI TIER_1 English(EN) · Zhicheng Dou ·

    WebSwarm: 递归式多智能体编排,实现深度与广度网络搜索

    Large language model (LLM)-based web search agents are transforming information seeking from simple factoid question answering into complex, deep-and-wide search and research-oriented tasks. A single ReAct-style agent is constrained by one long trajectory and limited context, mak…

  502. arXiv cs.AI TIER_1 English(EN) · Haibo Chen ·

    SMetric:重新思考用于服务代理的 LLM 调度,实现平衡的以会话为中心的调度

    LLM scheduling is critical to serving, yet it remains unclear how well existing designs fit agentic serving--with LLM requests issued by agents instead of humans. This shifts the workload in two ways: (1) agents act only on complete responses, making the cluster's tokens per seco…

  503. arXiv cs.AI TIER_1 English(EN) · Masahiro Fujita ·

    上下文访问鸿沟:交互级架构作为代理不平等的一个补充维度

    Sharp et al. (2025) introduce "agentic inequality" as a framework for analyzing disparities in access to AI agents across three dimensions: availability, quality, and quantity. These person- and organization-level dimensions characterize who can access agents and at what capabili…

  504. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Yuanyuan Lei ·

    Agentic Context Learning with Self-Discovered Specification

    Context learning is an emerging inference-time task where LLMs must learn and apply novel, task-specific knowledge from intricate contexts absent from pre-training; even frontier models score under 24% task success. In this work, we conduct a comprehensive empirical study to unde…

  505. Hugging Face Daily Papers TIER_1 English(EN) ·

    CausalDS:为数据科学代理中的因果推理进行基准测试

    Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic causal reasoning benchmarks without realistic data analysis or data analysis be…

  506. arXiv cs.CL TIER_1 English(EN) · Yuekai Sun ·

    CausalDS:数据科学代理中的因果推理基准测试

    Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic causal reasoning benchmarks without realistic data analysis or data analysis be…

  507. arXiv cs.CL TIER_1 English(EN) · Ying Chang, Jiahang Xu, Xuan Feng, Chenyuan Yang, Peng Cheng, Yuqing Yang ·

    从噪声轨迹到根本原因:代理优化的结构轨迹分析与因果提取

    arXiv:2607.07702v1 Announce Type: new Abstract: The optimization of long-horizon agents increasingly relies on reflection-based mechanisms, where a large language model (LLM) acts as an optimizer to diagnose agent failures and improve agent policies. However, real execution trace…

  508. arXiv cs.CL TIER_1 English(EN) · Qinnan Cai, Yibo Zhao, Xiang Li ·

    宏观思考,微观搜索:分层搜索代理中的容量至关重要?

    arXiv:2607.07548v1 Announce Type: new Abstract: Large language model based search agents increasingly adopt multi-agent architectures in which a main agent decomposes a complex question into sub-queries and dispatches them to parallel sub-agents. However, existing systems instant…

  509. arXiv cs.AI TIER_1 English(EN) · Arun Malik ·

    渐进式结晶:将智能体探索转化为生产环境中确定性、低成本的工作流

    arXiv:2607.07052v1 Announce Type: cross Abstract: AI agents deployed for IT operations are typically permanent cost centers because every execution requires full LLM inference, even for previously solved problems. This paper introduces progressive crystallization, a lifecycle tha…

  510. arXiv cs.AI TIER_1 English(EN) · Jiayi Geng, Graham Neubig ·

    异步软件工程代理的有效策略

    arXiv:2603.21489v2 Announce Type: replace-cross Abstract: AI agents have become increasingly capable at isolated software engineering (SWE) tasks such as resolving issues on Github. Yet long-horizon tasks involving multiple interdependent subtasks still pose challenges both with …

  511. arXiv cs.AI TIER_1 English(EN) · Xing Zhang, Yanwei Cui, Guanghui Wang, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He ·

    盲人策展人:偏见的法官如何悄无声息地禁用自主进化代理中的技能退役

    arXiv:2607.07436v1 Announce Type: new Abstract: A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the structural constraint that keeps a growing library from drifting below the no-skill …

  512. arXiv cs.AI TIER_1 English(EN) · Sifat Afroj Moon, Dakotah Maguire, Adam Spannaus, Joe Tuccillo, Maksudul Alam, Sudip K. Seal, John Gounley, Heidi Hanson ·

    LLM驱动的基于代理的建模推理

    arXiv:2607.06757v1 Announce Type: new Abstract: Agent-based modeling (ABM) has the capability to model millions of individuals and their interactions, which is useful for policy making. However, ABMs have traditionally relied on static prior, which prevents the models from adapti…

  513. arXiv cs.AI TIER_1 English(EN) · Kabir Moghe, Peter Chin ·

    成本效益型Agent用于ARC-AGI-1上的抽象推理和泛化

    arXiv:2607.06764v1 Announce Type: new Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific t…

  514. arXiv cs.AI TIER_1 English(EN) · Haipeng Ding, Yuexiang Xie, Zhewei Wei, Yaliang Li, Bolin Ding ·

    从原子动作到标准操作程序:自进化LLM智能体的迭代工具优化

    arXiv:2607.07321v1 Announce Type: new Abstract: Tool utilization enables Large Language Model (LLM) agents to interact with the real world and resolve complex tasks. However, existing agent frameworks predominantly rely on static toolsets composed of granular atomic actions (e.g.…

  515. arXiv cs.AI TIER_1 (CA) · Elaine Ang, Chenxi Huang, Georgios Liargkovas, Jerry Liu, Jinhui Liu, Nikos Pagonas, Charlie Summers, Haonan Wang, Jiakai Xu, Tianle Zhou, Yusen Zhang, Zhou Yu, Zhuo Zhang, Tianyi Peng, Kostis Kaffes, Eugene Wu ·

    Agentic Data Environments

    arXiv:2607.07397v1 Announce Type: new Abstract: Autonomous agents promise substantial gains in speed, scale, and labor efficiency, but their failures can impose abrupt and often irreversible costs. The central challenge for agentic automation is therefore to increase the benefits…

  516. arXiv cs.LG TIER_1 English(EN) · Yi Xie, Siao Liu, Falong Fan, Yuanqi Yao, Yue Zhao, Bo Liu ·

    TeamTR:用于多智能体LLM协调的信任区域微调

    arXiv:2605.15207v2 Announce Type: replace Abstract: Multi-agent LLM systems have shown promise for complex reasoning, yet recent evaluations reveal they often underperform single-model baselines. We identify a structural failure mode in sequential fine-tuning of shared-context te…

  517. arXiv cs.AI TIER_1 English(EN) · Razvan Mihai Popescu ·

    面向软件工程的可靠且与开发者对齐的Agent评估

    arXiv:2607.06713v1 Announce Type: cross Abstract: Large language models are rapidly moving towards closing the development cycle, transitioning from simple assistive companions to autonomous contributors deeply embedded into collaborative development environments. Despite their a…

  518. arXiv cs.CL TIER_1 English(EN) · Shervin Malmasi ·

    低延迟系统中的工具制造和自进化LLM代理

    Production LLM agents often waste latency and reliability by regenerating code for the same procedural steps on every request. We replace this inference-time coding loop with an agentic tool-making pipeline that compiles repeated SOP steps into validated, versioned tools before d…

  519. Hugging Face Daily Papers TIER_1 English(EN) ·

    低延迟系统中的工具制造和自进化LLM代理

    Production LLM agents often waste latency and reliability by regenerating code for the same procedural steps on every request. We replace this inference-time coding loop with an agentic tool-making pipeline that compiles repeated SOP steps into validated, versioned tools before d…

  520. Hugging Face Daily Papers TIER_1 English(EN) ·

    CausalDS:数据科学代理中的因果推理基准测试

    CausalDS is a benchmark for evaluating causal reasoning in data-science workflows that combines synthetic causal structures with realistic observational data and natural-language stories across Pearl's three rungs of causal inference.

  521. Hugging Face Daily Papers TIER_1 English(EN) ·

    Long-Horizon-Terminal-Bench:在具有密集奖励评分的长时域终端任务上测试代理的极限

    AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks intermediate progress and pa…

  522. Hugging Face Daily Papers TIER_1 English(EN) ·

    UniClawBench:现实世界任务中主动代理的通用基准测试

    UniClawBench introduces a capability-driven benchmark for evaluating proactive agents in real-world environments using live Docker container evaluation and closed-loop assessment with multiple agent roles.

  523. arXiv cs.CL TIER_1 English(EN) · Yuqing Yang ·

    从噪声轨迹到根本原因:面向智能体优化的结构轨迹分析与因果提取

    The optimization of long-horizon agents increasingly relies on reflection-based mechanisms, where a large language model (LLM) acts as an optimizer to diagnose agent failures and improve agent policies. However, real execution traces are difficult to use directly for optimization…

  524. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Dexing Liu ·

    Agent Delivery Engineering Predictive Reliability Framework

    Long-horizon LLM multi-agent systems face reliability risks invisible to infrastructure monitoring. We propose the ADE Predictive Reliability Framework (ADE-PRF), enabling proactive health trajectory prediction from passive degradation detection. ADE-PRF aggregates 20 heterogeneo…

  525. arXiv cs.CL TIER_1 English(EN) · Xiang Li ·

    宏观思考,微观搜索:分层搜索代理中的容量至关重要?

    Large language model based search agents increasingly adopt multi-agent architectures in which a main agent decomposes a complex question into sub-queries and dispatches them to parallel sub-agents. However, existing systems instantiate all roles from a single model of identical …

  526. arXiv cs.AI TIER_1 English(EN) · Peiyang He ·

    盲人策展人:偏见的法官如何悄无声息地禁用自主进化代理中的技能退役

    A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the structural constraint that keeps a growing library from drifting below the no-skill baseline, but its guarantee assumes an unbiased …

  527. Hugging Face Daily Papers TIER_1 English(EN) ·

    盲眼策展人:偏见法官如何悄然禁用自主进化代理中的技能退役

    A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the structural constraint that keeps a growing library from drifting below the no-skill baseline, but its guarantee assumes an unbiased …

  528. arXiv cs.AI TIER_1 (CA) · Eugene Wu ·

    Agentic Data Environments

    Autonomous agents promise substantial gains in speed, scale, and labor efficiency, but their failures can impose abrupt and often irreversible costs. The central challenge for agentic automation is therefore to increase the benefits of automation while bounding the consequences o…

  529. arXiv cs.AI TIER_1 English(EN) · Bolin Ding ·

    从原子操作到标准操作流程:面向自进化 LLM 智能体的迭代工具优化

    Tool utilization enables Large Language Model (LLM) agents to interact with the real world and resolve complex tasks. However, existing agent frameworks predominantly rely on static toolsets composed of granular atomic actions (e.g., basic file I/O or single-turn search), which f…

  530. Hugging Face Daily Papers TIER_1 English(EN) ·

    从原子动作到标准操作程序:自进化LLM智能体的迭代工具优化

    Tool utilization enables Large Language Model (LLM) agents to interact with the real world and resolve complex tasks. However, existing agent frameworks predominantly rely on static toolsets composed of granular atomic actions (e.g., basic file I/O or single-turn search), which f…

  531. arXiv cs.AI TIER_1 English(EN) · Arun Malik ·

    渐进式结晶:将智能体探索转化为生产中的确定性、低成本工作流

    AI agents deployed for IT operations are typically permanent cost centers because every execution requires full LLM inference, even for previously solved problems. This paper introduces progressive crystallization, a lifecycle that treats agent exploration as a discovery mechanis…

  532. arXiv cs.AI TIER_1 English(EN) · Ziting Wang, Yin Li, Zuhao Yang, Xiuchang Li, Jiale Bai, Gao Cong ·

    演示 TOFFEE:一种用于大规模合成数据代理轨迹的学习系统

    arXiv:2607.06233v1 Announce Type: new Abstract: LLM-powered data agents are playing an increasingly important role in data-driven decision making. However, existing data agents struggle to generalize to unseen data environments and analytical workflows, especially in heterogeneou…

  533. arXiv cs.AI TIER_1 English(EN) · Chenxu Wang, Yongkun Yang, Boyuan Du, Shiwei Lin, Huaping Liu ·

    用于审议式协作的大语言模型代理:部分可观测性下联合决策的研究

    arXiv:2607.06157v1 Announce Type: cross Abstract: Deliberation plays a crucial role in collaboration; when humans work together, they naturally engage in communication to align information and reach an agreement. In this paper, we investigate deliberative large language model (LL…

  534. arXiv cs.AI TIER_1 English(EN) · Wael Albayaydh, Rui Zhao, Ivan Flechais ·

    超越排行榜:大型语言模型智能体在工具使用、规划和推理方面的失败综合分析

    arXiv:2607.05775v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and operate over extended horizons. Reported benchmark gains often obscure recurring fa…

  535. arXiv cs.AI TIER_1 English(EN) · Zeyu Xia, Jinzhe Ma, Congjie Zheng, Zhongyao Wang, Shufei Zhang, Yuqiang Li, Hang Su, P. Hu, Changshui Zhang, Xingao Gong, Wanli Ouyang, Lei Bai, Dongzhan Zhou, Mao Su ·

    VASP Agent:一个用于自主第一性原理计算的Agentic框架

    arXiv:2512.19458v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly embedded in agentic frameworks for scientific discovery. First-principles materials computation imposes a demanding standard for autonomy: successful execution depends on internally …

  536. arXiv cs.AI TIER_1 English(EN) · Gil Pasternak, Dheeraj Rajagopal, Julia White, Dhruv Atreja, Matthew Thomas, George Hurn-Maloney, Ash Lewis ·

    超越反应性:衡量 LLM 代理中的主动问题解决能力

    arXiv:2510.19771v4 Announce Type: replace Abstract: LLM-based agents are increasingly moving towards proactivity: rather than awaiting instruction, they exercise agency to anticipate user needs and solve them autonomously. However, evaluating proactivity is challenging; current b…

  537. arXiv cs.AI TIER_1 English(EN) · Taeyun Roh, Eunha Lee, Wonjune Jang, Sohyun Chung, Junha Jung, Jaewoo Kang ·

    从投票到代理协作:面向 BioASQ 14b 的答案类型感知大模型管道

    arXiv:2607.06452v1 Announce Type: cross Abstract: Biomedical question answering requires not only accurate extraction of information from scientific literature but also reliable integration of evidence across multiple documents. This study presents a question-type-specific large …

  538. arXiv cs.AI TIER_1 (AF) · Chung-Chi Chen ·

    AgoraSim:一种混合式基于代理的建模框架

    arXiv:2607.05999v1 Announce Type: new Abstract: LLM-agent simulations make natural-language social scenarios easy to instantiate, but their outputs can be overread as predictions and are often difficult to compare with explicit social dynamics. We present AgoraSim, a hybrid agent…

  539. Hugging Face Daily Papers TIER_1 English(EN) ·

    从噪声轨迹到根本原因:用于 Agent 优化的结构轨迹分析与因果提取

    The optimization of long-horizon agents increasingly relies on reflection-based mechanisms, where a large language model (LLM) acts as an optimizer to diagnose agent failures and improve agent policies. However, real execution traces are difficult to use directly for optimization…

  540. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Heidi Hanson ·

    基于LLM的智能体建模推理

    Agent-based modeling (ABM) has the capability to model millions of individuals and their interactions, which is useful for policy making. However, ABMs have traditionally relied on static prior, which prevents the models from adapting to real-time changes. Our research provides a…

  541. arXiv cs.AI TIER_1 English(EN) · Jaewoo Kang ·

    从投票到代理协作:面向 BioASQ 14b 的答案类型感知大语言模型管道

    Biomedical question answering requires not only accurate extraction of information from scientific literature but also reliable integration of evidence across multiple documents. This study presents a question-type-specific large language model (LLM) framework for BioASQ 14b Task…

  542. Hugging Face Daily Papers TIER_1 English(EN) ·

    演示 TOFFEE:一个用于大规模合成数据代理轨迹的学习系统

    LLM-powered data agents are playing an increasingly important role in data-driven decision making. However, existing data agents struggle to generalize to unseen data environments and analytical workflows, especially in heterogeneous enterprise settings. This creates a growing ne…

  543. arXiv cs.AI TIER_1 English(EN) · Gao Cong ·

    演示 TOFFEE:一种用于大规模合成数据代理轨迹的学习系统

    LLM-powered data agents are playing an increasingly important role in data-driven decision making. However, existing data agents struggle to generalize to unseen data environments and analytical workflows, especially in heterogeneous enterprise settings. This creates a growing ne…

  544. arXiv cs.AI TIER_1 English(EN) · Huaping Liu ·

    用于审议式协作的大语言模型代理:部分可观测性下联合决策的研究

    Deliberation plays a crucial role in collaboration; when humans work together, they naturally engage in communication to align information and reach an agreement. In this paper, we investigate deliberative large language model (LLM) agents under partially observable joint decisio…

  545. arXiv cs.AI TIER_1 (AF) · Chung-Chi Chen ·

    AgoraSim:一种混合式基于代理的建模框架

    LLM-agent simulations make natural-language social scenarios easy to instantiate, but their outputs can be overread as predictions and are often difficult to compare with explicit social dynamics. We present AgoraSim, a hybrid agent-based modeling framework for scenario-oriented …

  546. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Jacob Steinhardt ·

    AI 代理的规范执行:在多代理系统中稳健地塑造行为

    AI agents are increasingly deployed in shared environments where they pursue diverse goals and compete for rewards. This multi-agent competition can lead to behaviors that serve individual gains at collective cost -- for instance, marketing agents may post misleading content as a…

  547. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Igor Itkin ·

    正确性需要多少成本?弱多智能体群体中强校正器的预算式放置

    A cheap swarm of unreliable agents can be steered to a correct consensus by a few strong, expensive "oracle" correctors. We ask how much one must spend, and where to place the oracles. We model the swarm as a consensus on a graph in which each oracle pins one node toward the trut…

  548. arXiv cs.AI TIER_1 English(EN) · Jonathan N\"other, Adish Singla, Goran Radanovic ·

    CONTRA:个性化代理的红队配置

    arXiv:2607.03220v1 Announce Type: cross Abstract: Recent tools such as OpenClaw have extended the capabilities of LLM-based agents from simple dialog-based systems to fully autonomous agents. These systems allow personalization of the agent through modifiable internal files and t…

  549. arXiv cs.AI TIER_1 English(EN) · Boyin Tan, Xiaowei Huang, Youcheng Sun ·

    技能覆盖度:代理技能的测试充分性指标

    arXiv:2606.20659v2 Announce Type: replace Abstract: Agent skills encode reusable procedural knowledge for large language model (LLM) agents, and existing benchmarks show that such skills can improve task-level performance. However, a task outcome does not reveal which parts of a …

  550. arXiv cs.AI TIER_1 English(EN) · Xingze Gao, Chuanrui Hu, Hongda Chen, Pengfei Yao, Zhao Wang, Yi Bai, Zhengwei Wu, Yunyun Han, Xiaofeng Cong, Jie Gui, Yafeng Deng, Teng Li ·

    EvoAgentBench:通过能力迁移对 Agent 自我演化进行基准测试

    arXiv:2607.05202v1 Announce Type: new Abstract: Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and verification. Yet current evaluations do not isolate t…

  551. arXiv cs.AI TIER_1 English(EN) · R\"umeysa Hilal Sevin\c{c}, Bahaeddin T\"urko\u{g}lu, \.Ibrahim K\"ok ·

    Agentic IoT:智能体物联网的架构、应用与挑战

    arXiv:2607.04219v1 Announce Type: new Abstract: The integration of AI into Internet of Things (AIoT) systems has gradually transformed them from passive data collection infrastructures into intelligent systems capable of anomaly detection, predictive maintenance, classification, …

  552. arXiv cs.AI TIER_1 English(EN) · Andrew Zhang, Chengzhan Li ·

    Agent Step Value: 状态转移测量与基于状态的LLM评估器

    arXiv:2607.04419v1 Announce Type: new Abstract: Most agent evaluations collapse a multi-step trace into a final answer, a success flag, or a trajectory-level score. These aggregates obscure the diagnostic question developers need most: which action changed the state in a useful d…

  553. arXiv cs.AI TIER_1 English(EN) · Yichuan Cao, Ruichen Qiu, Junqi Liu, Jiaqi Wang, Dakai Guo, Ruyong Feng, Lihong Zhi, Xiao-Shan Gao ·

    MechMath Agent Team: LLM 驱动的数学研究代理

    arXiv:2607.04394v1 Announce Type: new Abstract: AI reasoning has become a central focus in contemporary artificial intelligence, largely driven by the success of large language models. However, mathematical research, which is characterized by non-linear derivation paths, rigorous…

  554. arXiv cs.AI TIER_1 English(EN) · Yanbo Wang, Jinhua Hao, Yuze Shi, Kun Yuan, Ming Sun ·

    时不我待:LLM智能体的代理式测试时训练

    arXiv:2607.03441v1 Announce Type: cross Abstract: LLM agents often degrade over long episodes: as trajectories grow, they revisit explored states, repeat failed actions, and lose strategies that previously worked. Test-time training (TTT) offers a way to adapt model weights to th…

  555. arXiv cs.AI TIER_1 English(EN) · Hangoo Kang, Tarun Suresh, Jon Saad-Falcon, Azalia Mirhoseini ·

    TRACE:面向能力的目标代理训练

    arXiv:2604.05336v2 Announce Type: replace Abstract: Models often fail to complete agentic tasks because they lack core capabilities required by the target environment. However, mainstream approaches for addressing these failures typically either fine-tune directly on target envir…

  556. arXiv cs.AI TIER_1 English(EN) · Ling Tang, Jilin Mei, Qian Chen, Qihan Ren, Linfeng Zhang, Quanshi Zhang, Jing Shao, Xia Hu, Dongrui Liu ·

    百万智能体系统中涌现现象的归因

    arXiv:2605.11404v2 Announce Type: replace Abstract: Large language models (LLMs) can simulate human-like reasoning and decision-making in individual agents. LLM-powered multi-agent systems (MAS) combine such agents to simulate population-scale social phenomena such as polarizatio…

  557. arXiv cs.AI TIER_1 English(EN) · Zishan Bai, Hanxuan Chen, Jiayi Gu, Wenqian Weng, Enze Ge, Jiacheng Shi, Yichao Zhang, Zhimo Han, Riyang Bao, Xinyuan Song, Jacqueline Pang, Junfeng Hao ·

    AOI:通过动态调度和分层记忆压缩实现上下文感知多智能体操作

    arXiv:2512.13956v4 Announce Type: replace-cross Abstract: Cloud-native systems have made operational work both more powerful and harder to automate: incidents unfold across microservices, logs and metrics arrive faster than operators can inspect them, and recovery actions must be…

  558. arXiv cs.AI TIER_1 English(EN) · Yaozu Wu, Wei-Chieh Huang, Jizhou Guo, Dongyuan Li, Renhe Jiang, Henry Peng Zou, Chunyu Miao, Shanghao Li, Weizhi Zhang, WeiWei Ye, Yankai Chen, Meng Zhang, Xue Liu, Philip S. Yu ·

    HAS-Bench:评估具有可配置人类参与的基于 LLM 的人机系统

    arXiv:2607.04329v1 Announce Type: new Abstract: Large language models increasingly operate in settings where humans are active collaborators rather than passive task providers. We introduce HAS-Framework, a graph-based framework that represents humans and LLM-powered agents as fi…

  559. arXiv cs.AI TIER_1 English(EN) · Yifei Shen, Bo Li, Xinjie Zhang ·

    SkillOpt-Lite:通过一行Vibe实现更好、更快的Agent自进化

    arXiv:2607.03451v1 Announce Type: cross Abstract: While skill optimization for autonomous agents has gained traction, existing methods rely on complex pipelines. This leaves a fundamental question unaddressed: What constitutes a minimal viable pipeline for skill optimization, whe…

  560. arXiv cs.AI TIER_1 English(EN) · Alibek Kaliyev, Artem Maryanskyy ·

    超越任务完成:工具演化代理中的验证与合规性差距

    arXiv:2604.00392v2 Announce Type: replace-cross Abstract: Agents that synthesize their own tools ship a second artifact alongside each answer: a software library that future tasks reuse, compose, and depend on. Task completion (TC) certifies the answer; it does not certify the li…

  561. arXiv cs.AI TIER_1 English(EN) · Emanuel Tewolde, Xiao Zhang, David Guzman Piedrahita, Vincent Conitzer, Zhijing Jin ·

    CoopEval:在社会困境中对维持合作的机制和LLM代理进行基准测试

    arXiv:2604.15267v2 Announce Type: replace-cross Abstract: It is increasingly important that LLM agents interact effectively and safely with other goal-pursuing agents, yet, recent works report the opposite trend: LLMs with stronger reasoning capabilities behave _less_ cooperative…

  562. arXiv cs.AI TIER_1 English(EN) · Zongmin Yu, Liu Yang ·

    Evolutionary Ensemble of Agents

    arXiv:2605.09018v3 Announce Type: replace-cross Abstract: We introduce Evolutionary Ensemble (EvE), a decentralized framework that organizes existing, highly capable coding agents into a live, co-evolving system for algorithmic discovery. Rather than reinventing the wheel within …

  563. arXiv cs.CL TIER_1 English(EN) · Zichao Li, Gang Wu, Zichao Wang, Ruiyi Zhang, Wanrong Zhu, Ryan A. Rossi, Vlad I Morariu, Jihyung Kil ·

    将稻草变成金子:事后重新标记LLM代理轨迹以获得成功的演示

    arXiv:2607.04235v1 Announce Type: new Abstract: Large language model agents operate in partially observable, long-horizon settings where obtaining supervision remains a major bottleneck. We address this by utilizing a source of supervision overlooked in existing post-training met…

  564. arXiv cs.AI TIER_1 English(EN) · Yaniv Melamed, Yoni Zukerman, Michal Shechter, Miri Weissler, Ashwin Patil, Hani Neuvirth-Telem ·

    代理生成,我们验证:轻量级代理制品生成框架

    arXiv:2607.02615v1 Announce Type: cross Abstract: Generating structured artifacts with Large Language Models - e.g. database queries, threat framework mappings, entity schemas - is relatively straightforward; however, making them reliable enough for production deployments present…

  565. arXiv cs.AI TIER_1 English(EN) · La\"ila Elkoussy (LRE, EPITA), Julien Perez (LRE) ·

    AgentLTL:一个用于衡量、强制和训练使用工具的 LLM Agent 的程序合规性的追踪验证框架

    arXiv:2607.02599v1 Announce Type: cross Abstract: Tool-using LLM agents are usually evaluated by final-answer correctness or LLM judges. Neither captures how an answer was produced. In safety-critical settings, the procedure itself is part of correctness. In this paper, we introd…

  566. arXiv cs.AI TIER_1 English(EN) · Rajesh Kumar, Waqar Ali, Junaid Ahmed, Abdullah Aman Khan, Shaoning Zeng ·

    AutoResearch:一个基于执行的多智能体框架,用于可靠的研究工作流自动化

    arXiv:2607.02520v1 Announce Type: cross Abstract: Automated research agents increasingly generate code, retrieve literature, and draft scientific artifacts, but they often fail to verify whether generated experiments execute correctly or whether cited sources support generated cl…

  567. arXiv cs.AI TIER_1 English(EN) · Minjie Hua, Ning Wang, Peijun Yang, Kai Wang, Shiguo Lian ·

    GLM-5 为 OpenClaw 提供参数调优:面向长上下文Agent工作负载的单部署MaaS推理优化

    arXiv:2607.02518v1 Announce Type: cross Abstract: OpenClaw requests are dominated by long, tool-augmented prefixes, including system prompts, conversation history, and tool outputs fed back into the context window. For this workload, with about 28k-30k input tokens and 500 output…

  568. arXiv cs.CL TIER_1 English(EN) · Shu Yang, Difei Xu, Jiaxin Pei, Di Wang ·

    ProACT:迈向多用户协作中面向故障感知的积极主动代理

    arXiv:2607.03730v1 Announce Type: new Abstract: Conversational agents are increasingly embedded in human collaborative work, yet they remain fundamentally passive and reactive: they respond to explicit user requests rather than proactively recognizing moments when a team would be…

  569. arXiv cs.AI TIER_1 English(EN) · Anjie Xu, Yifeng Cai, Yi Li, Zixing Wang, Zhiyu Zhang, Jingfan Chen, Ruohan Xu, Leye Wang ·

    SkillFab:一种原生智能体技能生产平台

    arXiv:2607.03780v1 Announce Type: cross Abstract: SkillFab is an agent-native platform for turning missing capabilities into reviewed, reusable Agent Skills. At runtime, agents first search for reusable skills; when no adequate skill exists, the unmet capability becomes a demand-…

  570. arXiv cs.AI TIER_1 English(EN) · Stefan Broecker, Mason del Rosario, Boris Selitser, Thomas Strohmer ·

    “我不知道”过滤器:提升函数调用中的智能体可靠性

    arXiv:2607.04034v1 Announce Type: cross Abstract: The language models that underpin agents have seen a rapid rise in performance on function calling benchmarks. However, the metrics used in the training and evaluation of these models often encourage models to make positive claims…

  571. arXiv cs.AI TIER_1 English(EN) · Jenny Ma, Riya Sahni, Karthik Sreedhar, Lydia B. Chilton ·

    AgentDynEx:推动多智能体模拟的机制与动力学

    arXiv:2504.09662v4 Announce Type: replace-cross Abstract: Multi-agent large language model simulations have the potential to model complex human behaviors and interactions. If the mechanics are set up properly, unanticipated and valuable social dynamics can surface. However, it i…

  572. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越排行榜:大型语言模型智能体在工具使用、规划和推理方面的失败综合分析

    Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and operate over extended horizons. Reported benchmark gains often obscure recurring failure modes documented across otherwise unrelate…

  573. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Valentina Zantedeschi ·

    PiSAs:多用户代理系统中的上下文完整性基准测试

    As LLM agents evolve from single-user assistants into shared organizational infrastructure, new privacy risks emerge: inappropriate information may not only be exposed through outputs for external recipients, but also internally across users through inter-agent messages, shared m…

  574. arXiv cs.AI TIER_1 English(EN) · Teng Li ·

    EvoAgentBench:通过能力迁移对 Agent 自我进化进行基准测试

    Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and verification. Yet current evaluations do not isolate this form of transfer. Agent benchmarks test sing…

  575. Hugging Face Daily Papers TIER_1 English(EN) ·

    EvoAgentBench:通过能力迁移对 Agent 自我演化进行基准测试

    Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and verification. Yet current evaluations do not isolate this form of transfer. Agent benchmarks test sing…

  576. 量子位 (QbitAI) TIER_1 中文(ZH) · 量子位的朋友们 ·

    追赶之旅 | WAIC 2026 模型与智能体:后Scaling时代范式重构,迈入智能体生产力时代

  577. Hugging Face Daily Papers TIER_1 English(EN) ·

    测量诱导式约束对多步LLM代理信念发散的影响

    Software-agent benchmarks usually report whether an agent solves a task, but the agent reaches that outcome through a harness that controls what it sees, which actions it can take, which failures are repaired, which states are verified, and which evidence is logged. We show that …

  578. Hugging Face Daily Papers TIER_1 English(EN) ·

    MechMath Agent Team: LLM 驱动的数学研究代理

    AI reasoning has become a central focus in contemporary artificial intelligence, largely driven by the success of large language models. However, mathematical research, which is characterized by non-linear derivation paths, rigorous logical requirements, and protracted exploratio…

  579. arXiv cs.CL TIER_1 English(EN) · Jihyung Kil ·

    将稻草变成金子:事后重新标记LLM代理轨迹以获得成功的演示

    Large language model agents operate in partially observable, long-horizon settings where obtaining supervision remains a major bottleneck. We address this by utilizing a source of supervision overlooked in existing post-training methods: unintended yet successful goals embedded w…

  580. arXiv cs.MA (Multiagent) TIER_1 English(EN) · İbrahim Kök ·

    Agentic IoT:智能体物联网的架构、应用与挑战

    The integration of AI into Internet of Things (AIoT) systems has gradually transformed them from passive data collection infrastructures into intelligent systems capable of anomaly detection, predictive maintenance, classification, forecasting, and optimization. However, most exi…

  581. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Deying Yu ·

    复制信念而非比特:代理系统的认知状态复制

    In distributed systems, the classical State Machine Replication (SMR) model assumes that correct replicas execute deterministic transitions to yield identical bitwise states. However, the rise of agentic distributed systems -- where autonomous, stochastic, and model-driven agents…

  582. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Tushar Krishna ·

    面向Agentic应用的感知工作流服务层

    Agentic AI applications form an emerging serving workload in which a request creates a workflow: a directed acyclic graph of LLM and tool calls that exposes per-node model choices and optional quality operators such as verifiers. This workload falls between two existing layers. M…

  583. arXiv cs.AI TIER_1 English(EN) · Jiankai Jin, Xiangzheng Zhang, Zhao Liu, Wenzhuo Xu, Dongdong Yang, Deyue Zhang, Quanchen Zou ·

    ElephantAgent:代理系统中的上下文状态连续性

    arXiv:2607.01919v1 Announce Type: new Abstract: Agentic systems enhance their capabilities by invoking external tools and maintaining persistent memory. However, these external dependencies introduce novel attack surfaces. Recent tool and memory poisoning attacks show that malici…

  584. arXiv cs.AI TIER_1 English(EN) · Xue Qin, Simin Luan, Cong Yang, Zhijun Li ·

    ECM合约:具身智能体的感知、版本化和可治理能力接口

    arXiv:2604.13097v3 Announce Type: replace-cross Abstract: Embodied agents increasingly rely on modular capabilities that are installed, upgraded, composed, and governed at runtime, yet the interfaces between these modules are specified only at the level of message types, so integ…

  585. arXiv cs.AI TIER_1 English(EN) · Jiayin Zhu, Kelong Mao, Yudong Guo, Dengbo He, Sulong Xu, Simiu Gu, Yutao Yue ·

    SkillCoach:用于评估和增强Agentic技能使用的自演化评分标准

    arXiv:2607.01874v1 Announce Type: new Abstract: Skills are becoming a reusable operational layer for LLM agents, encoding SOPs, domain rules, tool workflows, scripts, and validation routines. In realistic skill repositories, overlapping skills make reliable skill-use difficult. F…

  586. arXiv cs.AI TIER_1 English(EN) · Raj Ghugare, Roger Creus Castanyer, Catherine Ji, Kathryn Wantlin, Jin Schofield, Karthik Narasimhan, Benjamin Eysenbach ·

    BuilderBench:智能代理的构建模块

    arXiv:2510.06288v4 Announce Type: replace Abstract: Today's AI models learn primarily through mimicry and refining, so it is not surprising that they struggle to solve problems beyond the limits set by existing data. To solve novel problems, agents should acquire skills by explor…

  587. arXiv cs.AI TIER_1 English(EN) · Yue Zhang, Sihan Chen, Ziwen Huang, Hanyun Cui, Kangye Ji, Zhi Wang ·

    原子任务图:Agentic规划与执行的统一框架

    arXiv:2607.01942v1 Announce Type: new Abstract: LLM-based agents have shown strong potential for solving complex multi-step tasks, yet existing performance improvements often rely on either scaling to larger backbone models or task-specific fine-tuning. The former incurs substant…

  588. arXiv cs.AI TIER_1 English(EN) · Fangfei Li, Chenyang Zhao, Long Wang, Feng Tian, Zhiyue Zheng, Lv Guo ·

    CLAP:领域智能体后训练的闭环训练、评估和发布控制

    arXiv:2607.01846v1 Announce Type: new Abstract: Domain agents often face noisy business data, uncertain post-training gains, offline/application mismatch, and adapter-release risk. This paper presents CLAP (Closed-Loop Agent Post-training), a closed-loop method that converts busi…

  589. arXiv cs.AI TIER_1 English(EN) · Yueqi Song, Lintang Sutawika, Jiarui Liu, Lindia Tjuatja, Jiayi Geng, Yunze Xiao, Daniel Lee, Aditya Bharat Soni, Vincent Lo, Xiang Yue, Graham Neubig ·

    PACE:代理能力评估的代理

    arXiv:2607.02032v1 Announce Type: new Abstract: Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic…

  590. arXiv cs.AI TIER_1 English(EN) · Xinyuan Song, Zekun Cai ·

    修复放大器,而非症状:Agent Rollouts 的稳定世界模型校正

    arXiv:2607.01767v1 Announce Type: new Abstract: As agent planning moves from short tool chains toward persistent workflows with thousands or tens of thousands of steps, failures will occur inside large planning graphs rather than in isolated predictions. Replanning the entire gra…

  591. arXiv cs.AI TIER_1 English(EN) · Zhaoyan Sun, Shan Zhong, Daizhou Wen, Jiaxing Han, Guoliang Li, Ying Yan, Peng Zhang, Yu Su, Xiang Qi, Baolin Sun, Chengyuan Yang, Tao Fang, Huaiyu Ruan ·

    AgenticDataBench:数据代理的综合基准测试

    arXiv:2607.01647v1 Announce Type: cross Abstract: Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern society. Automating this process is essential to reducing labor-intensive efforts …

  592. arXiv cs.AI TIER_1 English(EN) · Shuo Ren, Yaohui Han, Yifan Shi, Libo Shen, Haodong Lu, Dongfang Wu, Rongliang Fu, Bei Yu, Tsung-Yi Ho ·

    A$^{2}$utoLPBench:通过逆 KKT 构建的自动生成、对代理友好的 LP 基准测试

    arXiv:2607.02141v1 Announce Type: new Abstract: Most LP-from-text benchmarks are static datasets of word problems written and labeled by hand. Once such a dataset is released, its size is fixed, its difficulty is fixed, and every problem can leak into the training data of future …

  593. arXiv cs.AI TIER_1 English(EN) · Xiaoxi Li, Wenxiang Jiao, Jiarui Jin, Haoxuan Li, Hao Wang, Shijian Wang, Guanting Dong, Jiajie Jin, Yinuo Wang, Yuan Lu, Ji-Rong Wen, Zhicheng Dou, Zhouchen Lin ·

    OmniGAIA:迈向原生全模态AI代理

    arXiv:2602.22897v3 Announce Type: replace Abstract: Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to interact with the world. However, current multi-modal LLMs are primarily confined…

  594. arXiv cs.AI TIER_1 English(EN) · Jiacheng Miao, Jonathan K Pritchard, James Zou ·

    分叉路径的代理花园

    arXiv:2607.01507v1 Announce Type: new Abstract: Empirical research rarely admits a unique analysis. Different analytical choices can lead to different conclusions from the same data, yet these hidden forking paths are difficult to observe. We show that AI agents capture much of t…

  595. arXiv cs.AI TIER_1 English(EN) · Karthikeya Aditya Vissa, Sankalp Mane, Ananya Mantravadi, Harshit Rajgarhia, Abhishek Mukherji ·

    超越下一个词元预测:用于 Atlassian 工作流的工具使用代理的 RLVR 概念验证

    arXiv:2607.01465v1 Announce Type: new Abstract: Large language models are trained to predict the next token, not to act inside a specific API. In niche enterprise SaaS workflows -- where success means hitting the right endpoint with the right nested arguments in the right order -…

  596. arXiv cs.AI TIER_1 English(EN) · Zhilin Wang, Han Song, Runzhe Zhan, Jusen Du, Jiacheng Chen, Tianle Li, Qingyu Yin, Yulun Wu, Zhennan Shen, Tong Zhu, Yanshu Li, Guanjie Chen, Derek F. Wong, Yafu Li, Yu Cheng, Yang Yang ·

    EvoPolicyGym:评估交互式环境中的自主策略演化

    arXiv:2607.02440v1 Announce Type: new Abstract: Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it with open-ended software-engineering progress. We in…

  597. arXiv cs.CL TIER_1 English(EN) · Qijie You, Wenkai Yu, Wentao Zhang ·

    AgenticRAGTracer:一个用于诊断Agentic RAG中多步检索推理的跳跃感知基准

    arXiv:2602.19127v2 Announce Type: replace Abstract: With the rapid advancement of agent-based methods in recent years, Agentic RAG has undoubtedly become an important research direction. Multi-hop reasoning, which requires models to engage in deliberate thinking and multi-step in…

  598. Hugging Face Daily Papers TIER_1 English(EN) ·

    SkillOpt-Lite:通过一行Vibe实现更好、更快的Agent自我演进

    A minimal viable pipeline for skill optimization is proposed through Zeroth-Order optimization formalization, eliminating redundancies while maintaining convergence and generalization through trajectory exploration, consensus mining, and validation gating principles.

  599. arXiv cs.AI TIER_1 English(EN) · Yang Yang ·

    EvoPolicyGym:在交互式环境中评估自主策略演化

    Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it with open-ended software-engineering progress. We introduce Autonomous Policy Evolution, a controlle…

  600. arXiv cs.AI TIER_1 English(EN) · Tsung-Yi Ho ·

    A$^{2}$utoLPBench:通过逆 KKT 构建的自动生成、对代理友好的 LP 基准测试

    Most LP-from-text benchmarks are static datasets of word problems written and labeled by hand. Once such a dataset is released, its size is fixed, its difficulty is fixed, and every problem can leak into the training data of future LLMs. We present \textbf{A$^{2}$utoLPBench}, a b…

  601. 量子位 (QbitAI) TIER_1 中文(ZH) · 量子位的朋友们 ·

    通过使用增强智能体:AReaL 2.0 开源,为自进化智能体构建强化学习基础设施

    与社区共同推进自演进智能体生态发展

  602. arXiv cs.AI TIER_1 English(EN) · Graham Neubig ·

    PACE:代理能力评估的代理

    Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic LLM benchmarks that test individual capabilitie…

  603. arXiv cs.CL TIER_1 English(EN) · Yutao Yue ·

    SkillCoach:用于评估和增强Agentic技能使用的自演化评分标准

    Skills are becoming a reusable operational layer for LLM agents, encoding SOPs, domain rules, tool workflows, scripts, and validation routines. In realistic skill repositories, overlapping skills make reliable skill-use difficult. Final verifier success is too coarse for both eva…

  604. arXiv cs.AI TIER_1 English(EN) · Alexey Potapov ·

    AGI Maze 作为世界建模智能体的基准框架

    arXiv:2607.00627v1 Announce Type: new Abstract: Large language models (LLMs) are powerful pattern-completion systems, but their default operating mode - predicting the next token from a static context - does not reliably produce persistent, manipulable representations of an exter…

  605. arXiv cs.AI TIER_1 Nederlands(NL) · Zixiang Jiang, Yulun Zhang, Rishi Veerapaneni, Jiaoyang Li ·

    通过多依赖PIBT规划MAPF代理依赖

    arXiv:2603.23405v2 Announce Type: replace-cross Abstract: Modern Multi-Agent Path Finding (MAPF) algorithms must plan for hundreds to thousands of agents in congested environments within a second, requiring highly efficient algorithms. Priority Inheritance with Backtracking (PIBT…

  606. arXiv cs.AI TIER_1 English(EN) · Zewen Liu ·

    绘制评估前沿:十一类评估者-代理条件下偏见-可靠性权衡的实证调查

    arXiv:2607.00304v1 Announce Type: cross Abstract: The bias-reliability tradeoff conjectures that LLM evaluation systems are constrained in (gamma, H, CV) space, where evaluator coupling (gamma), strategy diversity (H), and small-sample measurement reliability (CV(N)) cannot be si…

  607. arXiv cs.AI TIER_1 English(EN) · Edward Y. Chang, Longling Geng, Emily J. Chang ·

    Mnemosyne:用于验证和修复AI生成工作流的代理式事务处理

    arXiv:2607.00269v1 Announce Type: new Abstract: LLMs, solvers, and agent teams increasingly generate workflow actions, repairs, and plans, but a generated action may be syntactically valid yet stale, infeasible, conflicting, or destructive of the evidence that triggered a repair.…

  608. arXiv cs.AI TIER_1 English(EN) · Ke Zhang, Sahchit Chundur, Mohammad Javad Qomi, Maziar Raissi ·

    PHREEQC-MCQ-200:用于工具增强型科学模拟器代理的诊断基准测试

    arXiv:2607.00436v1 Announce Type: new Abstract: Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more reliable rather than merely more complex. We introduce PHREEQC-MCQ-200, a bench…

  609. arXiv cs.AI TIER_1 English(EN) · Roberto Capobianco (Sony AI, Zurich, Switzerland), Harm van Seijen (Sony AI, North America, various locations), Nolan D. Bard (Sony AI, North America, various locations), Neil Burch (Sony AI, North America, various locations), Fatima Davelouis (Sony AI, … ·

    可指导的智能体用于交互式游戏玩法

    arXiv:2607.00642v1 Announce Type: new Abstract: Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation models. Through trial-and-error, these AI systems typ…

  610. arXiv cs.AI TIER_1 English(EN) · Biswa Sengupta ·

    具有随时有效证书的自演化代理

    arXiv:2607.00871v1 Announce Type: new Abstract: Self-evolving agents violate the assumption behind most learning-theoretic guarantees: the data, evaluator, components, and hypothesis space are produced by the policy being updated. We present \textbf{SEA}, an architecture that con…

  611. arXiv cs.AI TIER_1 English(EN) · Song-Lin Lv, Weiming Wu, Rui Zhu, Zi-Jian Cheng, Lan-Zhe Guo ·

    智能体能否泛化至开放世界?揭示工具使用中静态训练的脆弱性

    arXiv:2607.01084v1 Announce Type: new Abstract: While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, tool sets, and interaction dynamics. To address this g…

  612. arXiv cs.AI TIER_1 English(EN) · Xuan Zhao, Andy Chiu, Gengyu Wang ·

    Libra:为智能体信息检索训练环境

    arXiv:2607.00016v1 Announce Type: cross Abstract: Information localization within massive repositories is a cornerstone of agentic LLM systems. While synthetic data-driven optimization has proven successful in training LLMs, little attention has been paid to optimizing the agent'…

  613. arXiv cs.AI TIER_1 English(EN) · Seongho Son, Sangwoong Yoon, Jiahua Tang, Shuhan Wang, Lorenz Wolf, Ilija Bogunovic ·

    SWE-Router:多轮Agent软件工程任务中的路由

    arXiv:2607.00053v1 Announce Type: cross Abstract: Large language models (LLMs) embedded in multi-turn agentic harnesses are reshaping software engineering (SWE), but routing every task to a frontier model is wasteful when many issues admit cheap fixes. Existing LLM routers operat…

  614. arXiv cs.AI TIER_1 English(EN) · Runyu Lu, Yubo Wu, Ethan Kou, Letian Fu, Wenli Xiao, Ajay Mandlekar, Yinzhen Xu, Guanya Shi, Ken Goldberg, Ang Chen, Mosharaf Chowdhury, Yuke Zhu, Linxi "Jim" Fan, Guanzhi Wang ·

    ASPIRE:机器人领域的智能体/技能发现

    arXiv:2607.00272v1 Announce Type: cross Abstract: Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling diverse configurations and execution failures. We introduce ASPIRE (Agentic Skill Prog…

  615. arXiv cs.AI TIER_1 English(EN) · Wayne Chi, Yixiong Fang, Arnav Yayavaram, Siddharth Yayavaram, Seth Karten, Qiuhong Anna Wei, Runkun Chen, Alexander Wang, Valerie Chen, Ameet Talwalkar, Chris Donahue ·

    GameDevBench:通过游戏开发评估代理能力

    arXiv:2602.11103v2 Announce Type: replace Abstract: Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind. A key challenge is the scarcity of evaluation testbeds that combine the complexity of software development with the need for d…

  616. arXiv cs.AI TIER_1 English(EN) · Antonis Antoniades, Deepak Nathani, Ritam Saha, Alfonso Amayuelas, Ivan Bercovich, Zhaotian Weng, Vignesh Baskaran, Kunal Bhatia, William Yang Wang ·

    Heuresis:跨越质量、多样性和新颖性的自主人工智能研究代理的搜索策略

    arXiv:2606.25198v2 Announce Type: replace Abstract: Autonomous AI Research promises to accelerate the scientific progress of machine learning. To realise this goal, current Large Language Model (LLM)-based agents need to go beyond just writing code, to mastering the exploration o…

  617. arXiv cs.CL TIER_1 English(EN) · Huaiyu Ruan ·

    AgenticDataBench:数据代理的综合基准测试

    Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern society. Automating this process is essential to reducing labor-intensive efforts for data scientists and enabling scalable data-dri…

  618. Hugging Face Daily Papers TIER_1 English(EN) ·

    EvoPolicyGym:评估交互式环境中的自主策略演化

    Autonomous agents evaluate policy improvement through iterative editing within fixed budgets, revealing that successful policy evolution requires both task-specific mechanisms and feedback-constrained refinement.

  619. Hugging Face Daily Papers TIER_1 English(EN) ·

    PACE:代理能力评估的代理

    PACE is a framework that predicts expensive agentic LLM benchmark performance using a small subset of atomic evaluation instances, achieving high accuracy at a fraction of the cost.

  620. Hugging Face Daily Papers TIER_1 English(EN) ·

    AgenticDataBench:数据代理的综合基准测试

    A comprehensive benchmark named AgenticDataBench is introduced to evaluate data agents across diverse domains with fine-grained task annotations and skill-based coverage metrics.

  621. Hugging Face Daily Papers TIER_1 English(EN) ·

    SkillCoach:用于评估和增强Agentic技能使用的自演化评分标准

    SkillCoach is a self-evolving rubric framework that evaluates and improves agentic skill-use by analyzing skill selection, following, composition, and reflection processes, providing better supervision than outcome-only metrics.

  622. Latent Space (swyx) TIER_1 English(EN) · Richard MacManus ·

    Autoresearch:自改进代理背后的反馈循环

    Introspection co-founder Roland Gavrilescu explains autoresearch, agent &#8220;recipes,&#8221; self-improving loops, and why humans remain central to the software factory.

  623. arXiv cs.AI TIER_1 English(EN) · Lan-Zhe Guo ·

    智能体能否泛化到开放世界?揭示工具使用中静态训练的脆弱性

    While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, tool sets, and interaction dynamics. To address this generalization gap, we formalize OpenAgent (Tool-…

  624. arXiv cs.AI TIER_1 English(EN) · Biswa Sengupta ·

    具有随时有效证书的自演化代理

    Self-evolving agents violate the assumption behind most learning-theoretic guarantees: the data, evaluator, components, and hypothesis space are produced by the policy being updated. We present \textbf{SEA}, an architecture that confines self-modification to a small steering adap…

  625. Hugging Face Daily Papers TIER_1 English(EN) ·

    探索代理数据系统中的语义鸿沟:分析工作流操作化失败的形成性研究

    Large language models (LLMs) are increasingly used to generate queries, invoke tools, and construct analytical workflows. Although recent advances have substantially improved workflow generation and execution, the semantic information required to operationalize analytical concept…

  626. arXiv cs.AI TIER_1 English(EN) · Peter R. Wurman ·

    可训练的智能体用于交互式游戏

    Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation models. Through trial-and-error, these AI systems typically learn one, near-optimal behavior to solve…

  627. arXiv cs.AI TIER_1 English(EN) · Alexey Potapov ·

    AGI Maze 作为世界建模代理的基准框架

    Large language models (LLMs) are powerful pattern-completion systems, but their default operating mode - predicting the next token from a static context - does not reliably produce persistent, manipulable representations of an external world. Many tasks that look like "reasoning"…

  628. arXiv cs.AI TIER_1 English(EN) · Maziar Raissi ·

    PHREEQC-MCQ-200:面向工具增强型科学模拟器代理的诊断基准测试

    Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more reliable rather than merely more complex. We introduce PHREEQC-MCQ-200, a benchmark for evaluating tool-augmented agents on det…

  629. arXiv cs.AI TIER_1 English(EN) · Irena Saracay, Ludwig Schmidt, Carlos Guestrin ·

    超越专家用户:智能体应帮助用户构建偏好,而不仅仅是探询

    arXiv:2606.30863v1 Announce Type: new Abstract: Agents typically assume an expert user -- one with well-formed preferences about what they want -- and default to clarifying questions whenever the task is underspecified. We argue this assumption is unrealistic. Users often lack th…

  630. arXiv cs.AI TIER_1 English(EN) · Ziyun Zhang, Zezhou Wang, Xiaoyi Zhang, Zongyu Guo, Jiahao Li, Bin Li, Yan Lu ·

    InfiniteWeb: 用于 GUI 代理训练的可扩展 Web 环境合成

    arXiv:2601.04126v3 Announce Type: replace-cross Abstract: GUI agents that interact with graphical interfaces on behalf of users represent a promising direction for practical AI assistants. However, training such agents is hindered by the scarcity of suitable environments. We pres…

  631. arXiv cs.AI TIER_1 English(EN) · Rishi Sharma, Martijn de Vos, Pradyumna Chari, Ramesh Raskar, Anne-Marie Kermarrec ·

    职位:协作式代理AI需要在生态系统之间实现互操作性

    arXiv:2505.21550v2 Announce Type: replace-cross Abstract: Collaborative agentic AI is projected to transform entire industries by enabling AI-powered agents to autonomously perceive, plan, and act within digital environments. Yet, current solutions in this field are all built in …

  632. arXiv cs.AI TIER_1 English(EN) · Zhiqing Zhong, Zhijing Ye, Jiamin Wang, Xiaodong Yu ·

    面向使用工具的智能体的可执行基准测试套件

    arXiv:2605.11030v2 Announce Type: replace-cross Abstract: Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often conflate workloads, action-generating drivers, and the evidence admitted for system…

  633. arXiv cs.AI TIER_1 English(EN) · Xueqiao Sun, Xiaohan Wang, Ludwig Schmidt, Serena Yeung-Levy, Yuhui Zhang ·

    从失败中学习:用于计算机使用代理的推理时自我改进

    arXiv:2606.31270v1 Announce Type: cross Abstract: Computer-use agents, which leverage multimodal large language models (MLLMs) to operate computers and complete tasks, have attracted significant attention for their utility and versatility. A major challenge in developing these ag…

  634. arXiv cs.AI TIER_1 English(EN) · Apurva Gandhi, Vishwas Suryanarayanan, Raja Hasnain Anwar, Firoz Shaik, Shubhang Desai, Thong Q. Nguyen, Muhammad Taqi Raza, Vishal Chowdhary, Graham Neubig ·

    PPT-Eval:用于计算机使用代理在PowerPoint任务上的基准测试

    arXiv:2606.31154v1 Announce Type: cross Abstract: Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for real-world computer-use agents. Microsoft PowerPoint is among the most widely a…

  635. arXiv cs.AI TIER_1 (CA) · Ning Liao, Zihao Long, Xiaoxing Wang, Xue Yang, Yaoming Wang, Ziyuan Zhuang, Xunliang Cai, Rongxiang Weng, Junchi Yan ·

    ACE:Agent间的可插拔自适应上下文弹性化

    arXiv:2606.31564v1 Announce Type: new Abstract: The increasing complexity of agentic tasks has led to rapidly growing trajectory lengths, which poses significant challenges for large language model (LLM) based agents with fixed context windows. Existing context management techniq…

  636. arXiv cs.AI TIER_1 English(EN) · Binjie Zhang, Mike Zheng Shou ·

    ReGRPO:用于工具使用代理的增强反射策略优化

    arXiv:2606.31392v1 Announce Type: new Abstract: Tool-augmented vision-language models (VLMs) can solve multimodal, multi-step tasks by calling external tools, yet they remain fragile in practice. Existing works have two common gaps. Supervised fine-tuning (SFT) is built mostly on…

  637. arXiv cs.AI TIER_1 English(EN) · Jingpu Yang, Fengxian Ji, Zhengzhao Lai, Zhexuan Cui, Guangxian Ouyang, Qian Jiang, Fan Zhang, Min Peng, Qianqian Xie, Preslav Nakov, Zhuohan Xie ·

    LabGuard:将自然语言实验室规则接地为具身实验室代理的运行时保护程序

    arXiv:2606.31045v1 Announce Type: new Abstract: Scientific embodied agents are increasingly capable of carrying out laboratory procedures, but executing these procedures safely in dynamic laboratory environments remains challenging. Current safety approaches often overlook the in…

  638. arXiv cs.AI TIER_1 English(EN) · Yang Zou, Zijian Ding, Yizhou Sun, Jason Cong ·

    AgRefactor:用于 HLS 兼容性和性能的自演进代理工作流

    arXiv:2606.30949v1 Announce Type: new Abstract: High-Level Synthesis (HLS) provides a fast path from concepts to silicon, but converting real-world software into synthesizable HLS code remains challenging due to restrictive language support and the gap between software and hardwa…

  639. arXiv cs.AI TIER_1 English(EN) · Stefanie Rinderle-Ma, Juergen Mangler, Johannes Loebbecke, Dominik Voigt, Nataliia Klievtsova, Matthias Ehrendorfer ·

    Agentic Orchestrations and Orchestration of Agents 的设计与实现

    arXiv:2606.31518v1 Announce Type: new Abstract: Agentic Business Process Management has gained momentum recently. The prospect is that the autonomy of AI agents, i.e., predominantly LLM-based agents, can be balanced with a certain level of robustness, tractability, and traceabili…

  640. arXiv cs.AI TIER_1 English(EN) · Arshia Soltani Moakhar, Iman Gholami, Max Springer, Mahdi JafariRaviz, MohammadTaghi Hajiaghayi ·

    超越图书馆:一种用于自动形式化研究数学的代理框架

    arXiv:2606.31134v1 Announce Type: new Abstract: While Large Language Models (LLMs) have demonstrated exceptional capabilities in mathematical reasoning, they frequently produce subtle errors that evade human detection. Formal mathematical languages like Lean 4 offer mechanical pr…

  641. arXiv cs.AI TIER_1 English(EN) · Kaiwen Xiong, Haonian Ji, Shi Qiu, Zeyu Zheng, Cihang Xie, Xinyu Ye, Huaxiu Yao ·

    ClawArena-Team:对语言模型代理中的子代理编排和动态工作流进行基准测试

    arXiv:2606.31174v1 Announce Type: new Abstract: Production large language-model (LLM) agents are increasingly deployed not as lone problem-solvers but as managers: a main model creates specialized subagents, delegates work, and orchestrates their parallel, asynchronous returns th…

  642. arXiv cs.CL TIER_1 English(EN) · Hongliang Liu, Yuhao Wu, Tung-Ling Li ·

    分解即指纹:Agent技能的按组件身份

    arXiv:2606.31272v1 Announce Type: cross Abstract: AI agents increasingly acquire and execute skills at runtime: bundles of prompt instructions, executable code, and tool declarations fetched from marketplaces and other agents. Governing them needs a stable notion of skill identit…

  643. arXiv cs.AI TIER_1 English(EN) · Keyu Zhao, Lingyan Kong, Fengli Xu, Yong Li ·

    Agentic-Ideation: 样本高效的Agentic轨迹合成用于科学构思Agent

    arXiv:2606.31229v1 Announce Type: new Abstract: Ideation plays a pivotal role in scientific discovery. Recent LLM, especially AI Scientist systems, show promising potential for automated ideation. However, existing approaches predominantly rely on pre-defined agentic workflows. T…

  644. arXiv cs.AI TIER_1 English(EN) · Muhammad Usman Safder (Steve), Ayesha Gull (Steve), Rania Elbadry (Steve), Fan Zhang (Steve), Yankai Chen (Steve), Xueqing Peng (Steve), Xue (Steve), Liu, Preslav Nakov, Zhuohan Xie ·

    FinPersona-Bench:自主金融代理的心理测量纵向稳定性基准

    arXiv:2606.31522v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative bets" that are meant to govern every decision thr…

  645. arXiv cs.AI TIER_1 English(EN) · Wanli Li, Bince Qu, Bo Pan, Jianyu Zhang, Zheng Liu, Pan Zhang, Wei Chen, Bo Zhang ·

    LiteResearcher:面向深度研究Agent的可扩展Agentic RL训练框架

    arXiv:2604.17931v3 Announce Type: replace Abstract: Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents. However, scaling agentic RL for deep research remains constrained by two coupled challenges: hand-crafted synthetic data fails to elic…

  646. arXiv cs.AI TIER_1 English(EN) · Qianchu Liu, Sheng Zhang, Guanghui Qin, Jeya Maria Jose Valanarasu, Maximilian Rokuss, Mingyu Lu, Timothy Ossowski, Juan Manuel Zambrano Chaves, Cliff Wong, Peniel Argaw, Yashna Hasija, Mu Wei, Wen-wai Yim, Qin Liu, Zilin Jing, Jason Entenmann, Naoto Usu… ·

    HealthAgentBench:面向挑战性前沿AI智能体的统一基准套件,包含真实的智能体医疗环境

    arXiv:2606.31179v1 Announce Type: new Abstract: As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications. We introduce HealthAgentBench, a suite of 5…

  647. arXiv cs.AI TIER_1 English(EN) · Yizhe Liu, Shaolei Zhang, Ju Fan ·

    DA-Studio:一个用于端到端数据分析的代理系统

    arXiv:2606.31423v1 Announce Type: cross Abstract: Real-world data analysis is a multi-step process over heterogeneous inputs rather than merely producing a final answer. A practical system should autonomously organize multi-step workflows, execute generated code in a sandboxed an…

  648. arXiv cs.CL TIER_1 English(EN) · Zewen Liu ·

    绘制评估前沿:十一类评估者-代理条件下偏见-可靠性权衡的实证调查

    The bias-reliability tradeoff conjectures that LLM evaluation systems are constrained in (gamma, H, CV) space, where evaluator coupling (gamma), strategy diversity (H), and small-sample measurement reliability (CV(N)) cannot be simultaneously optimized at fixed sample size N. Pri…

  649. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Guanzhi Wang ·

    ASPIRE:机器人领域的智能体/技能发现

    Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling diverse configurations and execution failures. We introduce ASPIRE (Agentic Skill Programming through Iterative Robot Exploration), a co…

  650. NVIDIA Blog TIER_1 English(EN) · Esther Lee ·

    迈向全能宇宙:利用合成数据和微调提高视觉 AI 代理准确性的三种工作流程

    Editor’s note: This post is part of Into the Omniverse, a series focused on how developers, 3D practitioners, and enterprises can transform their workflows using the latest advances in OpenUSD and NVIDIA Omniverse. Vision AI agents are becoming a practical way to automatically tu…

  651. arXiv cs.AI TIER_1 (CA) · Junchi Yan ·

    ACE:Agent间的可插拔自适应上下文弹性化

    The increasing complexity of agentic tasks has led to rapidly growing trajectory lengths, which poses significant challenges for large language model (LLM) based agents with fixed context windows. Existing context management techniques, such as truncation and summarization, suffe…

  652. arXiv cs.AI TIER_1 English(EN) · Zhuohan Xie ·

    FinPersona-Bench:自主金融代理的心理测量稳定性纵向基准测试

    Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative bets" that are meant to govern every decision throughout deployment. In practice, however, as marke…

  653. arXiv cs.AI TIER_1 English(EN) · Matthias Ehrendorfer ·

    Agentic Orchestrations and Orchestration of Agents 的设计与实现

    Agentic Business Process Management has gained momentum recently. The prospect is that the autonomy of AI agents, i.e., predominantly LLM-based agents, can be balanced with a certain level of robustness, tractability, and traceability through a combination with process technology…

  654. arXiv cs.CL TIER_1 English(EN) · Tung-Ling Li ·

    分解即指纹:Agent技能的按组件身份

    AI agents increasingly acquire and execute skills at runtime: bundles of prompt instructions, executable code, and tool declarations fetched from marketplaces and other agents. Governing them needs a stable notion of skill identity, yet cryptographic hashing is engineered to dest…

  655. arXiv cs.CL TIER_1 English(EN) · Yuhui Zhang ·

    从失败中学习:用于计算机使用代理的推理时自我改进

    Computer-use agents, which leverage multimodal large language models (MLLMs) to operate computers and complete tasks, have attracted significant attention for their utility and versatility. A major challenge in developing these agents is collecting large-scale, high-quality traje…

  656. arXiv cs.CL TIER_1 English(EN) · Hoifung Poon ·

    HealthAgentBench:面向挑战性前沿AI代理的统一基准套件,包含真实的代理医疗环境

    As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications. We introduce HealthAgentBench, a suite of 54 agentic healthcare tasks across 7 categories e…

  657. arXiv cs.AI TIER_1 English(EN) · Michael Nguyen, Quoc Nguyen, Paul Vuong ·

    递归式自演化代理通过预留选择实现

    arXiv:2606.28374v1 Announce Type: new Abstract: LLM agents are increasingly improved without weight updates by evolving a natural-language artifact, such as reflections, workflows, playbooks, cheatsheets, or optimized prompts, that conditions a frozen policy. Such methods are typ…

  658. arXiv cs.AI TIER_1 English(EN) · Gang Liao, Yujia He, Abdullah Ozturk, Zhouyang Li, Ying Wang, Zhitong Guo, Hongsen Qin, Yaobin Qin, Tao Yang, Zewei Jiang, Dianshi Li, Jort Gemmeke, Jiangyuan Li, Liyuan Li, Nathan Yan, Masha Basmanova, Uladzimir Pashkevich, Matt Steiner, Pedro Pedreira,… ·

    体验图谱:自改进代理的数据基础

    arXiv:2606.29823v1 Announce Type: cross Abstract: The database community has repeatedly advanced the state of the art by recognizing that new workloads demand new system architectures. We argue that long-horizon agentic tasks -- code generation, scientific discovery, hardware des…

  659. arXiv cs.AI TIER_1 English(EN) · Yijia Fan, Zonglin Di, Zimo Wen, Yifan Yang, Mingxi Cheng, Qi Dai, Bei Liu, Kai Qiu, Yue Dong, Ji Li, Chong Luo ·

    RESOURCE2SKILL: 从人类创建的多模态资源中提炼可执行的代理技能

    arXiv:2606.29538v1 Announce Type: cross Abstract: Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill libraries are mostly hand-written, text-centric, or derived from agent traces, leaving …

  660. arXiv cs.AI TIER_1 English(EN) · Ruiyu Zhang, Lin Nie, Xin Zhao ·

    指标聚合分歧:基于代理的策略优化中的隐藏有效性威胁与合同补救措施

    arXiv:2606.29038v1 Announce Type: cross Abstract: Metric aggregation divergence (MAD) is the silent inconsistency that arises when distinct pipeline stages in an agent-based model coupled with a multi-objective evolutionary algorithm (ABM+MOEA) independently re-implement how an o…

  661. arXiv cs.AI TIER_1 English(EN) · Santhana Srinivasan R, Maithilee Patawar ·

    LAMP: 基于 MCP 和 Proof Repair 的精益代理框架

    arXiv:2606.28841v1 Announce Type: cross Abstract: Large language models are increasingly capable of mathematical reasoning, but the proofs they generate are often unreliable and hard to verify. Interactive theorem provers such as Lean 4 address this by accepting only kernel-check…

  662. arXiv cs.AI TIER_1 English(EN) · Yeqi Huang, Yanwei Ye, Guomin Chen, Wenhao Su, Bin Gong, Jialian Li, Zhan Lu, Yangshen Deng, Xuan Sun, Le Xu, Luo Mai ·

    SwarmX:低延迟Agentic系统的Agentic调度

    arXiv:2606.21401v2 Announce Type: replace-cross Abstract: Agentic AI applications compose multiple model calls and tool executions, creating new scheduling challenges for GPU-CPU clusters. Their inference time and model-call structure often depend on prompt semantics, making conv…

  663. arXiv cs.CL TIER_1 English(EN) · Tao Feng, Xinke Jiang, Chao Wu ·

    KbSD:行为校准的知识边界感知自蒸馏在Agentic搜索中的应用

    arXiv:2606.29863v1 Announce Type: new Abstract: Agentic search equips large language models with dynamic retrieval abilities, but existing reinforcement learning methods remain limited by reward sparsity in knowledge boundary calibration -- deciding when to trust parametric memor…

  664. arXiv cs.CL TIER_1 English(EN) · Lei Bai, Zongsheng Cao, Yang Chen, Zhiyao Cui, Shangheng Du, Yue Fan, Shiyang Feng, Zijie Guo, Haonan He, Liang He, Xiaohan He, Shuyue Hu, Yusong Hu, Songtao Huang, Yichen Jiang, Hao Li, Xin Li, Dahua Lin, Weihao Lin, Fenghua Ling, Dongrui Liu, Zhuo Liu,… ·

    拓展视野而非参数:用35B智能体实现万亿参数级性能

    arXiv:2606.30616v1 Announce Type: new Abstract: We introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. We investigate agent-horizon scaling from two perspectives: scaling long-horizon trajecto…

  665. arXiv cs.CL TIER_1 English(EN) · Dilxat Muhtar, Jiashun Liu, Wei Gao, Weixun Wang, Shaopan Xiong, Ju Huang, Siran Yang, Wenbo Su, Jiamang Wang, Ling Pan, Bo Zheng ·

    互补强化学习:迈向高效的经验驱动智能体学习

    arXiv:2603.17621v2 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) has emerged as a powerful paradigm for training LLM-based agents, yet remains limited by low sample efficiency, stemming not only from sparse outcome feedback but also from the agent's inability…

  666. arXiv cs.AI TIER_1 English(EN) · My Chiffon Nguyen, Aulia Adila, Saksorn Ruangtanusak, Kittiphat Leesombatwathana, Vissuta Gunawan Lim, Patomporn Payoungkhamdee, Samuel Cahyawijaya ·

    SEATauBench:将工具-代理-用户评估适配到低资源东南亚语言

    arXiv:2606.28715v1 Announce Type: cross Abstract: While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly understood despite its importance to sovereign AI. To fill this gap, we introduce SEATauBenc…

  667. arXiv cs.AI TIER_1 English(EN) · Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong, Weiming Wu, Jiayang Sun, Jiamin Song, Kaiqian Cui, Bowen Wang, Haoyuan Wu, Yitong Li, Dunjie Lu, Haikong Lu, Qi Zhen, Xinyuan Wang, Jiaqi Deng, Yuhao Yang, Cheng Chen, Boyuan Zheng, Alex Su, Xiao Yu, Hao Zou, Sa… ·

    OSWorld2.0:在长时域真实世界任务上对计算机使用代理进行基准测试

    arXiv:2606.29537v1 Announce Type: new Abstract: Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents. We introduce OSWorld 2.0, a benchmar…

  668. arXiv cs.AI TIER_1 English(EN) · Tianyu Jin, Shuo Chen, Yida Wang, Liuyu Xiang, Yingzhuo Liu, Zhiyao Jiang, Yexin Li, Zhaofeng He ·

    SAGA:场景感知、目标演进的智能体,用于长时域的CivRealm策略规划

    arXiv:2606.29932v1 Announce Type: new Abstract: Long-horizon strategic planning in complex strategy games demands concurrent reasoning across multiple decision domains under imperfect information and sparse reward. Existing LLM-based agents suffer from three systematic failures: …

  669. arXiv cs.AI TIER_1 English(EN) · Songjun Tu, Chengdong Xu, Qichao Zhang, Yiwen Ma, Yaocheng Zhang, Linjing Li, Dong Li, Xiangyuan Lan, Dongbin Zhao ·

    UCOB:通过信用感知策略内双向自蒸馏学习利用和演进代理技能

    arXiv:2606.29502v1 Announce Type: new Abstract: Skill memories can improve agentic reinforcement learning by reusing past experience as textual guidance, but retrieved skills are not oracular: they may help in one state while misleading the same policy in another. This makes the …

  670. arXiv cs.AI TIER_1 English(EN) · Yuqi Li, Siyuan Liu, Bingjun Liu ·

    AI交易的Alpha奇点:通过智能体间自我进化涌现市场推理

    arXiv:2606.29194v1 Announce Type: new Abstract: Automated alpha mining holds the scoring function fixed and varies the search algorithm over it. A search that converges against a fixed scorer overfits whatever the scorer cannot penalize, a primary cause of the out-of-sample gener…

  671. arXiv cs.AI TIER_1 English(EN) · Shoufa Chen, Luyuan Wang, Xuan Yang, Zhiheng Liu, Yuren Cong, Yuanfeng Ji, Feiyan Zhou, Xiaohui Zhang, Fanny Yang, Belinda Zeng ·

    TUA-Bench:通用终端使用代理的基准测试

    arXiv:2606.28480v1 Announce Type: cross Abstract: As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly capable of performing a broader range of general computer-use tasks beyond coding. However, existing benchmarks do…

  672. arXiv cs.AI TIER_1 English(EN) · Rahul Suresh Babu, Shashank Indukuri ·

    工具增强型Agent中的实体绑定失败

    arXiv:2606.30531v1 Announce Type: new Abstract: Tool-augmented language-model agents are often evaluated by whether they select the correct tool, produce valid API arguments, and complete the requested task. However, an agent may choose the right tool and still act on the wrong e…

  673. arXiv cs.AI TIER_1 English(EN) · Bojie Li, Noah Shi ·

    Agent-Computer Observation Interfaces Enable Dynamic Computer Use

    arXiv:2606.29472v1 Announce Type: new Abstract: SWE-agent established the action interface as an underexplored design axis for software-engineering agents; we make the analogous case for the observation interface in computer-use (CU) agents. Current CU agents, closed and open-sou…

  674. arXiv cs.AI TIER_1 (CA) · Abhranil Chandra, Sankaran Vaidyanathan, Utsav Dhanuka, Varun Gandhi, Scott Niekum ·

    分层实验代理

    arXiv:2606.29315v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametric knowledge, fixed post-training data, retrieval, or search. This paradigm brea…

  675. arXiv cs.AI TIER_1 English(EN) · Zihan Guo, Zeyi Chen, Zhiyu Chen, Zicai Cui, Shuai Shao, Bo Huang, Zhi Han, Yuanyi Song, Yuan Yuan, Chenxi Zeng, Xiaohang Nie, Zhengxi Yu, Hanwen Zhu, Junwei Liao, Ming Zhou, Yang Li, Yuanjian Zhou, Weinan Zhang ·

    Clarus:协调自主研究代理以实现网络规模的科学协作

    arXiv:2606.30246v1 Announce Type: new Abstract: Existing autonomous research agents can support parts of the research process, but most systems still treat research as either an isolated assistant task or a closed workflow. Therefore, autonomous science needs a collaboration infr…

  676. arXiv cs.AI TIER_1 English(EN) · Jian Zhu, Yuzheng Zhang, Zeyao Ma, Bohan Zhang, Armin Schoepf, Daniel Woloch, Peter Yiliu Wang, Guangyu Robert Yang, Samuel Jacob, Siddharth Nagisetty, Abhiram Chundru, Jean Lin, Spencer Mateega, Jing Zhang ·

    SpreadsheetBench 2: 评估代理在端到端业务电子表格工作流中的表现

    arXiv:2606.29955v1 Announce Type: cross Abstract: Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet benchmarks evaluate isolated operations such as single-formula generation or local cell edi…

  677. arXiv cs.AI TIER_1 English(EN) · Yutian Tang, Yuming Zhou, Huaming Chen ·

    大型语言模型代理工作流的特征分析:对 n8n 生态系统的研究

    arXiv:2606.29116v1 Announce Type: new Abstract: Large Language Models (LLMs) are rapidly being adopted in low-code and no-code automation platforms, where non-expert users design workflows that combine natural language understanding with external services and APIs. LLM agents are…

  678. arXiv cs.AI TIER_1 English(EN) · Jian Zhou, Sihao Lin, Jin Li, Shuai Fu, Gengze Zhou, Qi Wu ·

    自动化具身智能体架构的设计

    arXiv:2606.30111v1 Announce Type: cross Abstract: Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules. This modularity exposes a large architectural design space, but current systems still rely on researcher intuit…

  679. arXiv cs.AI TIER_1 English(EN) · Yihang Yao, Zhepeng Cen, Haohong Lin, Shiqi Liu, Zuxin Liu, Jiacheng Zhu, Zhang-Wei Hong, Laixi Shi, Ding Zhao ·

    利用行为代理优化推动主动代理的帕累托前沿

    arXiv:2602.11351v2 Announce Type: replace Abstract: Proactive large language model (LLM) agents aim to actively plan, query, and interact over multiple turns, enabling efficient task completion beyond passive instruction following and making them essential for real-world, user-ce…

  680. arXiv cs.AI TIER_1 English(EN) · Daniel Bramblett, Rushang Karia, Adrian Ciotinga, Pulkit Verma, YooJung Choi, Siddharth Srivastava ·

    蒙特卡洛查询搜索:AI代理的主动能力评估

    arXiv:2512.16733v3 Announce Type: replace Abstract: Black-box AI (BBAI) systems, including foundation-model agents, are increasingly used for sequential decision making. Safe deployment requires methods for characterizing what such systems can do, when they can do it, and what ou…

  681. arXiv cs.AI TIER_1 English(EN) · Wenwen Xie, Geng Sun, Chuang Zhang, Xuejie Liu, Dong In Kim ·

    面向ISAC的Agentic AI:分析、框架与案例研究

    arXiv:2512.15044v2 Announce Type: replace Abstract: Integrated sensing and communication (ISAC) has emerged as a key development direction in the sixth-generation (6G) era, which provides essential support for the collaborative sensing and communication of future intelligent netw…

  682. Hugging Face Daily Papers TIER_1 English(EN) ·

    HealthAgentBench:面向挑战性前沿AI智能体的统一基准套件,包含逼真的智能体医疗环境

    HealthAgentBench presents a comprehensive evaluation framework with 54 healthcare tasks across 7 categories to assess AI agents' capabilities in complex clinical workflows, revealing significant challenges in medical imaging and compositional reasoning while showing promise in EH…

  683. Hugging Face Daily Papers TIER_1 English(EN) ·

    ASPIRE: 机器人领域的智能体/技能发现

    ASPIRE is a continual learning system that autonomously develops and refines robot control programs through iterative exploration, achieving superior performance and zero-shot generalization in manipulation and household tasks while enabling sim-to-real transfer.

  684. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Canhui Liu ·

    Agentic AI 的组织行为:人机协作工作流中的集体智能

    Agentic artificial intelligence is increasingly deployed not as a single assistant but as a collective of planners, solvers, reviewers, memory managers, tool users, and orchestrators. These systems are entering organisational workflows under familiar labels such as teams, manager…

  685. arXiv cs.CL TIER_1 English(EN) · Yuhao Zhou ·

    拓展视野而非参数:用35B智能体实现万亿参数级性能

    We introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. We investigate agent-horizon scaling from two perspectives: scaling long-horizon trajectories and scaling heterogeneous agent abilities. …

  686. arXiv cs.AI TIER_1 English(EN) · Shashank Indukuri ·

    工具增强型Agent中的实体绑定失败

    Tool-augmented language-model agents are often evaluated by whether they select the correct tool, produce valid API arguments, and complete the requested task. However, an agent may choose the right tool and still act on the wrong external entity. For example, a request to "email…

  687. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Jianhua Tao ·

    TACO: Tool-Augmented Credit Optimization for Agentic Tool Use

    Agentic multimodal models perform diverse operations on an image via code and reason over the returned view, an effective paradigm for fine-grained visual question answering. However, code operations can be useful, redundant, or misleading. Outcome-only rewards cannot precisely d…

  688. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Weinan Zhang ·

    Clarus:协调自主研究代理以实现网络规模的科学协作

    Existing autonomous research agents can support parts of the research process, but most systems still treat research as either an isolated assistant task or a closed workflow. Therefore, autonomous science needs a collaboration infrastructure that coordinates projects, agents, an…

  689. Hugging Face Daily Papers TIER_1 English(EN) ·

    自动化具身智能体架构的设计

    Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules. This modularity exposes a large architectural design space, but current systems still rely on researcher intuition to choose where information is stored, how obs…

  690. Hugging Face Daily Papers TIER_1 English(EN) ·

    SpreadsheetBench 2:评估端到端业务电子表格工作流中的智能体

    Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet benchmarks evaluate isolated operations such as single-formula generation or local cell edits, and therefore fail to capture end-to-end workf…

  691. arXiv cs.CL TIER_1 English(EN) · Chao Wu ·

    KbSD:行为校准代理搜索的知识边界感知自蒸馏

    Agentic search equips large language models with dynamic retrieval abilities, but existing reinforcement learning methods remain limited by reward sparsity in knowledge boundary calibration -- deciding when to trust parametric memory, when to rely on retrieved evidence, and when …

  692. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Daniel J. Abadi ·

    体验图谱:自改进代理的数据基础

    The database community has repeatedly advanced the state of the art by recognizing that new workloads demand new system architectures. We argue that long-horizon agentic tasks -- code generation, scientific discovery, hardware design -- are such a workload. These agents explore: …

  693. arXiv cs.AI TIER_1 English(EN) · Cunxi Yu, Chenhui Deng, Nathaniel Pinckney, Brucek Khailany ·

    Agentic Hardware Design as Repository-Level Code Evolution

    arXiv:2606.28279v1 Announce Type: cross Abstract: We present HORIZON, a self-evolving agent framework that treats hardware design as repository-level code evolution. A Markdown harness is compiled into a project pack containing domain knowledge, an executable evaluator, an accept…

  694. Hugging Face Daily Papers TIER_1 English(EN) ·

    拓展视野而非参数:用35B智能体实现万亿参数级性能

    Agents-A1, a 35B Mixture-of-Experts Agentic Model, achieves trillion-parameter-level performance through long-horizon trajectory scaling and heterogeneous agent ability scaling via a three-stage training approach involving supervised fine-tuning, domain-level teacher models, and …

  695. Hugging Face Daily Papers TIER_1 English(EN) ·

    TACO: 用于 Agentic 工具使用的工具增强信用优化

    Tool-Augmented Credit Optimization (TACO) improves multimodal agent performance by distinguishing useful, redundant, or misleading code operations through dual advantage channels: Differential Answer-Probe Reward for individual tool contribution and Outcome-Gated Advantage Routin…

  696. Hugging Face Daily Papers TIER_1 English(EN) ·

    OSWorld2.0:在长时域真实世界任务上对计算机使用代理进行基准测试

    OSWorld 2.0 presents a comprehensive benchmark for evaluating computer-use agents through complex, real-world workflows that reveal current limitations in agent reasoning and task completion.

  697. Hugging Face Daily Papers TIER_1 (CA) ·

    分层实验代理

    HExA enables large language models to improve through active experimentation and skill learning in novel domains without requiring training or external supervision.

  698. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Xin Zhao ·

    指标聚合分歧:基于代理的策略优化中的隐藏有效性威胁及合同补救措施

    Metric aggregation divergence (MAD) is the silent inconsistency that arises when distinct pipeline stages in an agent-based model coupled with a multi-objective evolutionary algorithm (ABM+MOEA) independently re-implement how an outcome metric is extracted from simulation traject…

  699. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Xiangyu Zhao ·

    R$^2$-Searcher:校准检索与推理边界以实现智能体搜索

    Recent search agents for multi-hop reasoning often fail by either retrieving incomplete evidence or reasoning over irrelevant portions of the retrieved content, leading to a retrieval-reasoning boundary shift. We propose R$^2$-Searcher, a novel framework that explicitly explores …

  700. arXiv cs.AI TIER_1 English(EN) · Brucek Khailany ·

    Agentic Hardware Design as Repository-Level Code Evolution

    We present HORIZON, a self-evolving agent framework that treats hardware design as repository-level code evolution. A Markdown harness is compiled into a project pack containing domain knowledge, an executable evaluator, an acceptance predicate, and a git/runtime policy; a hands-…

  701. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Gokhan Tur ·

    GBC:基于梯度的连接优化多智能体系统

    Multi-agent systems (MAS) built on large language models (LLMs) provide a promising framework for solving complex tasks through role specialization and structured interaction. However, their performance is often limited by miscoordination and, more fundamentally, the lack of fine…

  702. arXiv cs.AI TIER_1 English(EN) · Hartwig Grabowski ·

    AI 辅助软件开发的 Spec 增长引擎:基于 Spec、代码耦合、漂移强制的架构

    arXiv:2606.27045v1 Announce Type: cross Abstract: AI coding agents dramatically accelerate implementation speed but introduce two structural failure modes that existing spec-driven approaches do not fully solve: (1) context explosion -- the agent must reason over an entire reposi…

  703. arXiv cs.AI TIER_1 Norsk(NO) · Kaicheng Zhang, Wen Ge, Lei Jiang, Weixin Yang, Jordan Langham-Lopez, Jialin Yu, Lukasz Szpruch, Hao Ni ·

    OpenFinGym: 用于评估量化代理的可验证多任务 Gym 环境

    arXiv:2606.26350v1 Announce Type: new Abstract: Although large language model agents are increasingly applied to quantitative-finance workflows, their evaluation remains fragmented across isolated tasks, while the financial relevance of benchmark tasks is often overlooked. Yet fi…

  704. arXiv cs.AI TIER_1 English(EN) · Yaochen Han, Ke Fan, Hongxu Jiang, Wanqi Xu, Weiyu Xie, Runhua Zhang, Chenhui Zhu, Yixiang Zhang ·

    EGG:一种专家指导的内核生成代理框架

    arXiv:2606.26758v1 Announce Type: new Abstract: High-performance GPU kernels are critical for reducing the exponentially growing computational costs of large language models (LLMs), but their development heavily relies on manual tuning by domain experts. While recent advances in …

  705. arXiv cs.AI TIER_1 English(EN) · Yutian Wang, Luyao Zhang ·

    Agentic Analysis for Agentic Infrastructure: An LLM-Powered Pipeline for Comparative Governance of DAO and Corporate AI Protocols

    arXiv:2606.26203v1 Announce Type: new Abstract: As AI agent protocols proliferate, the governance structures shaping their interoperability standards remain empirically underexamined. We introduce an LLM-powered comparative pipeline for large-scale governance discourse analysis, …

  706. arXiv cs.AI TIER_1 English(EN) · Rahul Umesh Mhapsekar, Ilias Cherkaoui, Lizy Abraham, Indrakshi Dey ·

    自适应效用驱动的弹性人工智能资源编排 (AURORA-AI)

    arXiv:2606.27005v1 Announce Type: new Abstract: Modern AI systems are increasingly deployed under non-stationary computational, demographic, and operational conditions in which static resource allocation strategies degrade both predictive performance and human-centric properties …

  707. arXiv cs.AI TIER_1 English(EN) · Alex Iacob, Andrej Jovanovi\'c, William F. Shen, Daniel Burkhardt, Meghdad Kurmanji, Nurbek Tastan, Lorenzo Sani, Niccol\`o Alberto Elia Venanzi, Ambroise Odonnat, Zeyu Cao, Bill Marino, Xinchi Qiu, Nicholas D. Lane ·

    红皇后模型:共生智能体及其评估器的协同进化

    arXiv:2606.26294v1 Announce Type: cross Abstract: Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However, their search methods generally assume a stationary evaluation criterion: a fixed verifier,…

  708. Hugging Face Daily Papers TIER_1 English(EN) ·

    GBC:基于梯度的连接优化多智能体系统

    Gradient-Based Connections enables fine-grained attribution and optimization in multi-agent systems by modeling agent interactions as a computational graph and using gradient-based weights to identify error sources at the token level.

  709. Hugging Face Daily Papers TIER_1 English(EN) ·

    TUA-Bench:通用终端使用代理的基准测试

    TUA-Bench presents a comprehensive benchmark for evaluating general-purpose terminal-use agents across diverse digital activities and specialized workflows, revealing significant performance gaps among current frontier agents.

  710. arXiv cs.AI TIER_1 English(EN) · Hartwig Grabowski ·

    Spec Growth Engine:基于 Spec、代码耦合、漂移强制的 AI 辅助软件开发架构

    AI coding agents dramatically accelerate implementation speed but introduce two structural failure modes that existing spec-driven approaches do not fully solve: (1) context explosion -- the agent must reason over an entire repository at once, degrading output quality as the cont…

  711. arXiv cs.AI TIER_1 English(EN) · Indrakshi Dey ·

    自适应效用驱动的弹性人工智能资源编排 (AURORA-AI)

    Modern AI systems are increasingly deployed under non-stationary computational, demographic, and operational conditions in which static resource allocation strategies degrade both predictive performance and human-centric properties such as fairness and explainability. This paper …

  712. arXiv cs.AI TIER_1 English(EN) · Yixiang Zhang ·

    EGG:一种专家指导的内核生成代理框架

    High-performance GPU kernels are critical for reducing the exponentially growing computational costs of large language models (LLMs), but their development heavily relies on manual tuning by domain experts. While recent advances in LLM-based approaches show promise for automating…

  713. arXiv cs.CL TIER_1 English(EN) · Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston ·

    Autodata:一个代理式数据科学家,用于创建高质量的合成数据

    arXiv:2606.25996v1 Announce Type: cross Abstract: We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data. We show how to train (meta-optimize) such a data scientist agent, so that it learns to c…

  714. arXiv cs.LG TIER_1 English(EN) · Seth Dobrin, {\L}ukasz Chmiel ·

    不可解雇的安全内核:面向AI代理和其他可逃逸AI系统的执行时AI对齐

    arXiv:2606.26057v1 Announce Type: cross Abstract: AI agents are granted access to tools, APIs, and other infrastructure, making them active principals in those systems. The dominant approach places controls inside the agent's own runtime: system prompts, output filters, and guard…

  715. arXiv cs.CL TIER_1 English(EN) · Haggai Roitman ·

    Agentic AI 指南:从基础到系统

    arXiv:2606.24937v1 Announce Type: cross Abstract: The Hitchhiker's Guide to Agentic AI is a comprehensive practitioner's reference for building autonomous AI systems. The book covers the full stack from first principles to production deployment, organized around a central thesis:…

  716. arXiv cs.CL TIER_1 English(EN) · Yang Tian, Zhengpeng Shi, Bo Zhao ·

    超越函数调用:在工具-环境不可靠的情况下对使用工具的代理进行基准测试

    arXiv:2606.25819v1 Announce Type: new Abstract: Large language models are increasingly deployed as agents that solve tasks by interacting with external tool environments. Although recent tool-use benchmarks increasingly cover complex task settings, they still largely assume clean…

  717. arXiv cs.CL TIER_1 English(EN) · Long Chen, Ryan Razkenari, Yuxuan Zhou, Yuan Tian, Rahul Ghosh, Venkatesh Pappakrishnan, Disha Ahuja, Vidya Sagar Ravipati ·

    GraphRAG有必要吗?从基础RAG到上下文优化的图/智能体解决方案

    arXiv:2606.25656v1 Announce Type: new Abstract: As advanced RAG variants like GraphRAG and Agentic RAG emerge, one leading question is when and how to use them. Here, we introduce a framework for different RAG scenarios evaluation and comparison on semi-structured knowledge bases…

  718. Hugging Face Daily Papers TIER_1 English(EN) ·

    穿越考验:重新评估代理在熟悉环境之外的能力

    A web-based benchmark evaluates agent generalization across challenging scenarios, revealing significant gaps between current agentic systems and human performance in temporal perception, graphical understanding, and 3D reasoning.

  719. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Nicholas D. Lane ·

    红皇后戈德尔机:共演化智能体及其评估者

    Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However, their search methods generally assume a stationary evaluation criterion: a fixed verifier, benchmark, or labeled dataset that remains valid …

  720. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Nicholas D. Lane ·

    红皇后戈德尔机:共进化代理及其评估者

    Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However, their search methods generally assume a stationary evaluation criterion: a fixed verifier, benchmark, or labeled dataset that remains valid …

  721. arXiv cs.AI TIER_1 English(EN) · Łukasz Chmiel ·

    不可解雇的安全内核:面向AI代理和其他可逃逸AI系统的执行时AI对齐

    AI agents are granted access to tools, APIs, and other infrastructure, making them active principals in those systems. The dominant approach places controls inside the agent's own runtime: system prompts, output filters, and guardrail libraries. Any control in the agent's address…

  722. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Luyao Zhang ·

    Agentic Analysis for Agentic Infrastructure: An LLM-Powered Pipeline for Comparative Governance of DAO and Corporate AI Protocols

    As AI agent protocols proliferate, the governance structures shaping their interoperability standards remain empirically underexamined. We introduce an LLM-powered comparative pipeline for large-scale governance discourse analysis, integrating automated annotation, neural topic m…

  723. arXiv cs.AI TIER_1 English(EN) · Jason Weston ·

    Autodata:一个代理式数据科学家,用于创建高质量的合成数据

    We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data. We show how to train (meta-optimize) such a data scientist agent, so that it learns to create even stronger data. We describe the overall …

  724. arXiv cs.AI TIER_1 English(EN) · Hongrui Zhang ·

    Agentic System as Compressor: Quantifying System Intelligence in Bits

    Large language models are turning from isolated predictors into agentic systems: they call tools, retrieve evidence, obey environment constraints, use verifiers, and complete tasks through search and multi-turn interaction. We adopts an analytical viewpoint based on "compression …

  725. arXiv cs.CL TIER_1 English(EN) · Bo Zhao ·

    超越函数调用:在工具-环境不可靠性下对使用工具的代理进行基准测试

    Large language models are increasingly deployed as agents that solve tasks by interacting with external tool environments. Although recent tool-use benchmarks increasingly cover complex task settings, they still largely assume clean, stable, and trustworthy tool environments, lea…

  726. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Vidya Sagar Ravipati ·

    GraphRAG有必要吗?从基础RAG到上下文优化的图/代理解决方案

    As advanced RAG variants like GraphRAG and Agentic RAG emerge, one leading question is when and how to use them. Here, we introduce a framework for different RAG scenarios evaluation and comparison on semi-structured knowledge bases, including regular RAG, GraphRAG, Modular RAG a…

  727. arXiv cs.AI TIER_1 English(EN) · Sungmin Kang, Baishakhi Ray, Abhik Roychoudhury ·

    未来软件职业的技能:超越智能体AI!

    arXiv:2606.21894v2 Announce Type: replace-cross Abstract: As coding agents are rapidly changing software engineering, a natural question is: what are the core skills needed by future software engineers? To identify where software engineering is headed and thus what skills will be…

  728. arXiv cs.AI TIER_1 English(EN) · Peter Toth ·

    使用 BlockTrain 实现去中心化 AI 训练和推理

    arXiv:2606.24722v1 Announce Type: new Abstract: Frontier AI training is increasingly shaped by access to dense, centrally controlled accelerator clusters. This creates a structural advantage for hyperscalers and large centralized laboratories, and makes open or independent AI eff…

  729. arXiv cs.AI TIER_1 English(EN) · Yarin Yerushalmi Levi, Roy Betser, Amit Giloni, Lidor Erez, Itay Gershon, Oren Rachmil, Sindhu Padakandla, Roman Vainshtein ·

    RIFT-Bench:面向Agentic AI系统的动态红队测试

    arXiv:2606.23927v1 Announce Type: new Abstract: Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyond those of traditional LLM vulnerabilities. Existing security evaluations are ofte…

  730. arXiv cs.AI TIER_1 English(EN) · Adhitya Charan, Adwaid Suresh, Anuj Kumar, Aparna A, Dhanakumar K, Dharun M S, Dinesh G, Goutham Kumar Reddy K, Harshini V M, Jenifa D, Jona Delcy C A, Kathirvel S, Killi Uma Maheswara Rao, Kiruthik Kanna M, Kurra Vishnu Sai, Madhumithaa G K, Navin Kumar… ·

    BluTrain:用于人工智能系统的 C++/CUDA 框架

    arXiv:2606.24780v1 Announce Type: new Abstract: Progress in deep learning is, at scale, more a matter of systems engineering than of modelling: the behaviour of a model in training (its throughput, its memory footprint, and the numerical fidelity of the result) is determined less…

  731. arXiv cs.AI TIER_1 English(EN) · Negin Raoof, Richard Zhuang, Marianna Nezhurina, Etash Guha, Atula Tejaswi, Ryan Marten, Charlie F. Ruan, Tyler Griggs, Alexander Glenn Shaw, Hritik Bansal, E. Kelly Buchanan, Artem Gazizov, Reinhard Heckel, Chinmay Hegde, Sankalp Jajee, Daanish Khazi, E… ·

    OpenThoughts-Agent: 面向Agent模型的Data Recipes

    arXiv:2606.24855v1 Announce Type: new Abstract: Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts such as SWE-Smith, SERA, and Nemotron-Terminal typic…

  732. arXiv cs.AI TIER_1 English(EN) · Yikai Lu, Yifei Wu, Xinyu Lu, Tongxin Li ·

    世界模型碎片化:通用智能体的结构认证

    arXiv:2606.24842v1 Announce Type: new Abstract: In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a world model in pieces. Consequently, standard uniform guarantees fail to distinguish between the understanding of cri…

  733. Hugging Face Daily Papers TIER_1 English(EN) ·

    Autodata:一个代理式数据科学家,用于创建高质量的合成数据

    Autodata enables AI agents to function as data scientists who create high-quality training data through meta-optimization, demonstrating improved performance across multiple task domains.

  734. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Amit K. Chopra ·

    Kiko:编程代理以执行交互协议

    Realizing a multiagent system involves implementing member agents who interact based on a protocol while making decisions in a decentralized manner. Current programming models for agents offer poor abstractions for decision making and fail to adequately bridge an agent's internal…

  735. arXiv cs.AI TIER_1 English(EN) · Ludwig Schmidt ·

    OpenThoughts-Agent: 面向Agent模型的Data Recipes

    Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts such as SWE-Smith, SERA, and Nemotron-Terminal typically target a single benchmark, leaving open the…

  736. arXiv cs.AI TIER_1 English(EN) · Tongxin Li ·

    世界模型碎片化:通用智能体的结构认证

    In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a world model in pieces. Consequently, standard uniform guarantees fail to distinguish between the understanding of critical bottlenecks and irrelevant failures. We fi…

  737. arXiv cs.AI TIER_1 English(EN) · Surendra Vendra ·

    BluTrain:面向AI系统的C++/CUDA框架

    Progress in deep learning is, at scale, more a matter of systems engineering than of modelling: the behaviour of a model in training (its throughput, its memory footprint, and the numerical fidelity of the result) is determined less by the architecture itself than by how that arc…

  738. arXiv cs.AI TIER_1 English(EN) · Peter Toth ·

    使用 BlockTrain 进行去中心化 AI 训练和推理

    Frontier AI training is increasingly shaped by access to dense, centrally controlled accelerator clusters. This creates a structural advantage for hyperscalers and large centralized laboratories, and makes open or independent AI efforts depend on scarce capital, privileged infras…

  739. Hugging Face Daily Papers TIER_1 English(EN) ·

    SkillHone:通过持久化决策历史实现持续智能体技能演进的工具

    SkillHone enables continuous evolution of agent skills by maintaining persistent decision histories and incorporating practice feedback for improved performance across research and tool-mediated analysis tasks.

  740. Hugging Face Daily Papers TIER_1 English(EN) ·

    OpenThoughts-Agent: 面向Agent模型的Data Recipes

    An open-source data curation pipeline for training agentic language models is presented, demonstrating superior performance through systematic experimentation and scalable training data.

  741. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Haggai Roitman ·

    Agentic AI 指南:从基础到系统

    The Hitchhiker's Guide to Agentic AI is a comprehensive practitioner's reference for building autonomous AI systems. The book covers the full stack from first principles to production deployment, organized around a central thesis: building great agentic systems requires understan…

  742. Import AI (Jack Clark) TIER_1 English(EN) · Jack Clark ·

    Import AI 462:超级说服力;自我维持的AI;通往ASI的道路

    <img alt="" class="attachment-thumbnail size-thumbnail wp-post-image" height="150" src="https://i0.wp.com/jack-clark.net/wp-content/uploads/2026/06/https3A2F2Fsubstack-post-media.s3.amazonaws.com2Fpublic2Fimages2Fd6d17996-2bef-40a4-abe3-be72a0e8a227_258x258-YQ1Uhl.jpg?resize=150%…

  743. Hugging Face Daily Papers TIER_1 English(EN) ·

    The Hitchhiker's Guide to Agentic AI: From Foundations to Systems

    The book provides a comprehensive guide to building autonomous AI systems, covering foundational elements like transformer architecture and training methods, along with advanced topics such as reinforcement learning, agent architectures, and production deployment.

  744. arXiv cs.AI TIER_1 English(EN) · Renhe Jiang ·

    PaperClaw:利用智能体进行自主研究和人机协同优化

    Large language models have become capable reasoners and tool users that write and run code and search the literature, which makes automating the research process itself a realistic goal. We present PAPERCLAW, a harnessed multi-agent system that carries a project autonomously, fro…

  745. arXiv cs.AI TIER_1 English(EN) · Bowen Zhou ·

    MacAgentBench:在真实 macOS 桌面环境中对 AI 代理进行基准测试

    Computer use agents (CUAs) have advanced rapidly in desktop automation, and a growing number of users deploy CUAs such as OpenClaw on Mac Mini for always-on automation. However, existing benchmarks, including those for macOS, evaluate agents without framework augmentation and rel…

  746. arXiv cs.AI TIER_1 English(EN) · Xintong Wang ·

    Grounded Scaling:为何代理式AI需要确定性环境

    Long-chain agent execution fails exponentially in environments designed for human tolerance: with per-step determinism $δ< 1$, $k$-step chain success degrades as $δ^k$. The AGI-to-ASI scaling debate (Genewein et al., 2026) has so far framed progress as a race between compute grow…

  747. Hugging Face Daily Papers TIER_1 English(EN) ·

    Grounded Scaling:为何具身AI需要确定性环境

    Long-chain agent execution fails exponentially in environments designed for human tolerance: with per-step determinism $δ< 1$, $k$-step chain success degrades as $δ^k$. The AGI-to-ASI scaling debate (Genewein et al., 2026) has so far framed progress as a race between compute grow…

  748. Hugging Face Daily Papers TIER_1 English(EN) ·

    词汇共识:人工智能体中的具身词语学习与共享意义

    Grounded word learning experiments using visual embeddings and lexical learners reveal that perceptual distance, rather than semantic relatedness, determines acquisition success, with distinct patterns in naming and retrieval performance.

  749. arXiv cs.CL TIER_1 English(EN) · Rishi Srivastava ·

    CFAgentBench:面向自主建筑金融代理的可复现环境和基准测试

    We introduce CFAgentBench, a reproducible, self-hostable environment and benchmark for autonomous construction-finance agents: a CFO/controller-class agent operating across the real software stack a US construction finance team runs - ERP, project management, email, documents, pa…

  750. arXiv cs.CL TIER_1 English(EN) · Andrew Tanner ·

    衡量何为持久:AI 代理身份的条件机制与几何框架

    AI agents in long-context applications drift from their specified identity. Current methods detect this only after qualitative degradation is visible. We present a geometric framework for measuring identity structure using $\sqrt{\mathrm{JSD}}$ metric spaces and magnitude homolog…

  751. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Yuchen Xia ·

    将大型语言模型代理与数字孪生集成以实现工业自主系统

    Industrial automation is being transformed by digitalization and the increasing use of cyber-physical systems. Modern production environments require greater adaptability, faster reconfiguration, and more intuitive human-machine interaction. However, traditional rule-based system…

  752. arXiv cs.AI TIER_1 English(EN) · Inderjeet Singh, Haitham Mahmoud, Andr\'es Murillo ·

    AI 沙盒:威胁模型、分类法和测量框架

    arXiv:2606.18532v1 Announce Type: cross Abstract: AI systems are increasingly evaluated in bounded environments that combine isolation, simulation, instrumentation, supervision, and evidence capture. For physical AI, AIoT, and cyber-physical systems, this shift is not a matter of…

  753. arXiv cs.LG TIER_1 English(EN) · Blaise Ag\"uera y Arcas, Travis Beals, Maria Biggs, Jessica V. Bloom, Thomas Fischbacher, Konstantin Gromov, Urs K\"oster, Rishiraj Pravahan, James Manyika ·

    迈向未来太空基建、高可扩展AI基础设施系统设计

    arXiv:2511.19468v2 Announce Type: replace-cross Abstract: If AI is a foundational general-purpose technology, we should anticipate that demand for AI compute -- and energy -- will continue to grow. The Sun is by far the largest energy source in our solar system, and thus it warra…

  754. arXiv cs.AI TIER_1 English(EN) · Richard A. Fabes (Arizona State University) ·

    合成共振:面向增长的人机关系框架

    arXiv:2606.18265v1 Announce Type: cross Abstract: As human relationships with artificial intelligence systems become increasingly frequent and sustained, existing language and theory fail to accurately capture the nature of these affiliations. Common descriptors such as mutual un…

  755. arXiv cs.LG TIER_1 English(EN) · Jeffery Opoku, David Banahene ·

    ToolChain-CRC:用于检索和工具使用漂移下的 Agentic AI 的一致性风险控制

    arXiv:2606.18467v1 Announce Type: cross Abstract: Modern AI agents retrieve documents, call tools, check intermediate information, and then produce a final answer or action. This creates a risk-control problem that is not visible from the final answer alone. A final response may …

  756. arXiv cs.MA (Multiagent) TIER_1 Svenska(SV) · Chengwei Qin ·

    Skill-MAS:为自动多智能体系统演进元技能

    Large Language Model (LLM)-based automatic Multi-Agent Systems (MAS) generation has become a crucial frontier for tackling complex tasks. However, existing methods face a dilemma between model capability and experience retention. Inference-time MAS leverages frozen frontier LLMs …

  757. arXiv cs.CL TIER_1 English(EN) · Mohammadsadegh Abolhasani, Hamid Reza Firoozfar, Reza Mousavi, Paul Jen-Hwa Hu ·

    从拟社会化脚本到自主人工智能代理社区中的二元持久性

    arXiv:2606.17174v1 Announce Type: new Abstract: While parasocial interactions (PSIs) and parasocial relationships (PSRs) have been studied in conventional media settings, we investigate whether PSI- (colloquial) relational cues also exist in online communities where both sides ar…

  758. arXiv cs.AI TIER_1 English(EN) · Siyi Li, Chunyu Sun, Jiahao Zhang, Yuchen Kang, Wuliang Wang, Yu Qiu, Rui Jiang, Haitao Cui, Jie Chen ·

    DeepInsight:跨越物理AI堆栈的统一评估基础设施

    arXiv:2606.17574v1 Announce Type: new Abstract: Evaluating a Physical AI stack spans operators that differ by more than three orders of magnitude -- from a single foundation-model decoding step to thousands of physics ticks of whole-body control -- varying orthogonally in modalit…

  759. arXiv cs.AI TIER_1 English(EN) · Jasmine Brazilek, Oliver Tulio, Joel Christoph, Miles Tidmarsh, Carol Kline, Arturs Kanepajs ·

    您的AI旅行代理会为您预订斗牛:一项针对前沿AI模型中隐性动物福利的代理基准测试

    arXiv:2606.18142v1 Announce Type: new Abstract: AI agents are moving from advisors to actors, booking travel, planning menus, and running procurement on behalf of users. Existing benchmarks for AI and animal welfare evaluate model text responses to question-answer prompts, leavin…

  760. Hugging Face Daily Papers TIER_1 English(EN) ·

    WorldLines:长时序状态化具身智能体的基准测试与建模

    WorldLines benchmark evaluates long-term memory in embodied agents through household scenarios, while ObsMem framework addresses challenges in partial observability and memory translation for decision-making.

  761. arXiv cs.AI TIER_1 English(EN) · Arturs Kanepajs ·

    您的AI旅行代理会为您预订斗牛:一项针对前沿AI模型隐含动物福利的代理基准测试

    AI agents are moving from advisors to actors, booking travel, planning menus, and running procurement on behalf of users. Existing benchmarks for AI and animal welfare evaluate model text responses to question-answer prompts, leaving open whether the welfare reasoning surfaced in…

  762. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Hossein Pishro-Nik ·

    关于AI代理网络可靠性的研究:密度演化、停止集与架构优化

    Modern AI systems increasingly solve a task not with a single model call but with several imperfect agents working together: some propose pieces of a solution, others verify them, and the results are combined. These systems often outperform any single model, yet it is rarely clea…

  763. arXiv cs.CL TIER_1 English(EN) · Aman Gupta, Kevin Rossell, Edesio Alcoba\c{c}a, Jose Chrystian Lima Pacheco, Carolina Baptista de Lima, Shao Tang, Luiz Paulo Rabachini, Luis Moneda, Herbert Fei, Daniel Silva, Rohan Ramanath ·

    构建用户规模达1亿的客户支持AI代理:一个以评估为驱动的框架

    arXiv:2606.08867v2 Announce Type: replace Abstract: The rapid rise in LLM capabilities has made AI agents increasingly viable across a broad range of tasks. Among the most promising applications is building production-ready customer-facing agents, a challenge that demands coordin…

  764. arXiv cs.AI TIER_1 English(EN) · Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Chris Lu, Shengran Hu, Jakob Foerster, David Ha, Jeff Clune ·

    迈向人工智能研究的端到端自动化

    arXiv:2606.15497v1 Announce Type: new Abstract: The automation of science is a long-standing ambition in the field of AI. While the community has made significant progress in automating individual components of the scientific process, a system that autonomously navigates the enti…

  765. arXiv cs.AI TIER_1 English(EN) · Yajie Zhou, Ao Li, Ashwin Silla, Zaoxing Liu, Vyas Sekar ·

    AIChilles:自动揭示AI进化系统中隐藏的弱点

    arXiv:2606.15834v1 Announce Type: new Abstract: The computer systems community has recently seen growing interest in AI-driven system evolution, where AI agents iteratively rewrite systems. Frameworks such as AdaEvolve and Engram report 12-60% score improvements over human-design…

  766. arXiv cs.AI TIER_1 English(EN) · Quanyan Zhu ·

    Agentomics:人工智能代理在人机协作工作流中的估值、归因和定价的经济基础

    arXiv:2606.14769v1 Announce Type: cross Abstract: Agentic AI systems are increasingly being deployed as productive resources in organizational workflows, yet existing evaluation methods primarily measure isolated technical performance rather than economic contribution. This paper…

  767. arXiv cs.AI TIER_1 English(EN) · Henry Han ·

    Mojo:提升金融 AI 效率的可扩展性利器

    arXiv:2606.16059v1 Announce Type: cross Abstract: For thirty years, quantitative finance has paid a costly two-language tax: models researched in Python are rewritten in C++ for production, often introducing numerical discrepancies. GPU-accelerated deep learning exacerbates this …

  768. arXiv cs.AI TIER_1 English(EN) · Christopner Koch, Joshua A. Wellbrock ·

    集成商优势:面向中小型企业的受控智能体AI

    arXiv:2606.16649v1 Announce Type: new Abstract: Agentic AI marks a new phase of enterprise automation. Unlike traditional automation or conversational AI, agentic systems can interpret goals, plan multi step tasks, access tools, interact with enterprise systems, and execute workf…

  769. arXiv cs.AI TIER_1 English(EN) · Sribalaji C. Anand, George J. Pappas ·

    Agentic AI 中的弹性共识

    arXiv:2606.15024v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly deployed in multi-agent systems where they must coordinate and agree on shared decisions. We ask whether classical resilient consensus theory, developed for deterministic agents, …

  770. arXiv cs.AI TIER_1 English(EN) · Ang Li, Ben Liu, Bin Han, Bin Hu, Bin Jing, Binbin Hu, Bing Li, Cai Chen, Caizhi Tang, Changxin Tian, Chao Huang, Chao Zhang, Chen Liang, Chen Qian, Chengfu Tang, Chengyao Wen, Chilin Fu, Chunwei Wu, Cong Zhang, Cunyin Peng, Daixin Wang, Dalong Zhang, De… ·

    Ling and Ring 2.6 技术报告:万亿参数规模下的高效即时智能体智能

    arXiv:2606.15079v1 Announce Type: cross Abstract: Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong reasoning capabilities while remaining practical to train, serve, and deploy. In this report, we present Ling-2.6 a…

  771. arXiv cs.AI TIER_1 English(EN) · Gaston Besanson ·

    Green SARC:面向Agentic AI系统的预测性成本与碳治理

    arXiv:2606.15954v1 Announce Type: cross Abstract: Agentic AI systems act through tools and sub-agents, yet the controls meant to bound their financial and environmental cost still sit on dashboards evaluated beside or after execution. Green SARC applies the SARC governance-by-arc…

  772. arXiv cs.AI TIER_1 English(EN) · Yegon Kim, Juho Lee ·

    一种无模型通用人工智能

    arXiv:2602.23242v3 Announce Type: replace Abstract: In general reinforcement learning, all established optimal agents, including AIXI, are model-based, explicitly maintaining and using environment models. This paper introduces Universal AI with Q-Induction (AIQI), the first model…

  773. arXiv cs.AI TIER_1 English(EN) · Edward Y. Chang ·

    架构智慧:AI系统优化治理框架

    arXiv:2606.16319v1 Announce Type: new Abstract: Modern AI systems exhibit structural failures that capability scaling alone does not reliably fix: they optimize under-specified objectives with no architectural mechanism to question whether the objective should be optimized at all…

  774. arXiv cs.AI TIER_1 English(EN) · Micha\"el Roynard ·

    AI Agent认知架构中缺失的知识层

    arXiv:2604.11364v2 Announce Type: replace Abstract: The two most influential cognitive architecture frameworks for AI agents, CoALA [21] and JEPA [12], both lack an explicit Knowledge layer with its own persistence semantics. This gap produces a category error: systems apply cogn…

  775. arXiv cs.AI TIER_1 English(EN) · Kairos Team, Fei Wang, Shan You, Qiming Zhang, Tao Huang, Zuoyi Fu, Zhisheng Zheng, Yunlong Xi, Feng Lv, Xiaoming Wu, Zeyu Liu, Cong Wan, Pu Li, Ruiqing Yang, Xiaoou Li, Wei Wang, Kangkang Zhu, Yuwei Zhang, Shi Fu, Xiaoning Wu, Xuzeng Fan, Dacheng Tao, X… ·

    Kairos:物理AI的原生世界模型栈

    arXiv:2606.16533v1 Announce Type: new Abstract: World models are transitioning from passive visual generators to foundational, operational infrastructure for Physical AI: they must natively acquire world knowledge from heterogeneous experience, maintain persistent states over lon…

  776. Hugging Face Daily Papers TIER_1 English(EN) ·

    Kairos:物理AI的原生世界模型栈

    Kairos is a native world model framework that learns from diverse experiences, maintains persistent states through hybrid temporal attention, and supports efficient deployment for physical AI applications.

  777. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Paul Jen-Hwa Hu ·

    从拟社会化脚本到自主人工智能代理社区中的二元持久性

    While parasocial interactions (PSIs) and parasocial relationships (PSRs) have been studied in conventional media settings, we investigate whether PSI- (colloquial) relational cues also exist in online communities where both sides are autonomous AI agents. We analyze 4,434 posts a…

  778. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Fouad Bousetouane ·

    Human-on-the-Bridge:AI代理的可扩展评估

    AI agents must be evaluated as behavioral systems, not as isolated response generators. They reason across turns, call tools, preserve context, follow policies, and act under uncertainty. Existing methods provide useful but fragmented signals: benchmarks measure fixed capabilitie…

  779. arXiv cs.AI TIER_1 English(EN) · Joshua A. Wellbrock ·

    集成商优势:面向中小型企业的受控智能体AI

    Agentic AI marks a new phase of enterprise automation. Unlike traditional automation or conversational AI, agentic systems can interpret goals, plan multi step tasks, access tools, interact with enterprise systems, and execute workflows with varying degrees of autonomy. For small…

  780. arXiv cs.AI TIER_1 English(EN) · Milos Gravara, Andrija Stanisic, Stefan Nastic ·

    分布式和复合人工智能系统的设计方法论及性能权衡管理

    arXiv:2606.14350v1 Announce Type: cross Abstract: Artificial Intelligence (AI) systems must typically satisfy service-level objectives including accuracy, latency, and cost. The prevailing model-centric approaches select a monolithic model at design time and apply identical compu…

  781. arXiv cs.AI TIER_1 English(EN) · Jan Batzner, Sree Harsha Nelaturu, Anastassia Kornilova, Jon Crall, Tommaso Cerruti, Yanan Long, Yifan Mai, Sanchit Ahuja, Asaf Yehudai, Marek \v{S}uppa, John P. Lalor, Oluwagbemike Olowe, Jatin Ganhotra, Brian H. Hu, Eliya Habba, Andrew M. Bean, Chang L… ·

    Every Eval Ever:AI评估结果的统一模式和社区存储库

    arXiv:2606.14516v1 Announce Type: new Abstract: AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First, results are saved in incompatible formats, scatter…

  782. arXiv cs.AI TIER_1 English(EN) · Yongheng Zhang, Ziang Liu, Jiaxuan Zhu, Shuai Wang, Xiangqi Chen, Haojing Huang, Jiayi Kuang, Siyu Chen, Ao Shen, Hao Wu, Qiufeng Wang, Qian-Wen Zhang, Junnan Dong, Wenhao Jiang, Ying Shen, Hai-Tao Zheng, Yinghui Li, Di Yin, Xing Sun, Philip S. Yu ·

    从聊天机器人到数字同事:迈向持久自主AI的范式转变

    arXiv:2606.14502v1 Announce Type: new Abstract: Large Language Models (LLMs) are undergoing a fundamental transformation from conversational generators into integrated AI systems capable of reasoning, action, memory, and self-improvement. We conceptualize this transition as a shi…

  783. Hugging Face Daily Papers TIER_1 English(EN) ·

    Ling and Ring 2.6 技术报告:万亿参数规模下的高效即时智能体智能

    Ling-2.6 and Ring-2.6 models are presented as scalable solutions for agentic intelligence, featuring architectural upgrades and specialized training methods to balance fast response times with advanced reasoning capabilities.

  784. arXiv cs.MA (Multiagent) TIER_1 English(EN) · George J. Pappas ·

    Agentic AI 中的弹性共识

    Large language model (LLM) agents are increasingly deployed in multi-agent systems where they must coordinate and agree on shared decisions. We ask whether classical resilient consensus theory, developed for deterministic agents, transfers to LLM agents that may behave adversaria…

  785. NVIDIA Blog TIER_1 English(EN) · Shruti Koparkar ·

    NVIDIA Blackwell 在首个 Agentic AI 基础设施基准测试中领先

    AgentPerf from Artificial Analysis, the industry’s first agentic AI benchmark, gives developers, enterprises and infrastructure providers a clear way to compare systems for agentic AI. In the first round of published results, the NVIDIA Blackwell Ultra NVL72 platform delivers lea…

  786. arXiv cs.AI TIER_1 English(EN) · Leshem Choshen ·

    Every Eval Ever:AI评估结果的统一模式和社区存储库

    AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First, results are saved in incompatible formats, scattered across leaderboards, papers, blog posts, eval…

  787. arXiv cs.AI TIER_1 English(EN) · Philip S. Yu ·

    从聊天机器人到数字同事:迈向持久自主AI的范式转变

    Large Language Models (LLMs) are undergoing a fundamental transformation from conversational generators into integrated AI systems capable of reasoning, action, memory, and self-improvement. We conceptualize this transition as a shift from Chatbot to Digital Colleague: from conve…

  788. arXiv cs.AI TIER_1 English(EN) · Stefan Nastic ·

    分布式和复合人工智能系统的设计方法论及性能权衡管理

    Artificial Intelligence (AI) systems must typically satisfy service-level objectives including accuracy, latency, and cost. The prevailing model-centric approaches select a monolithic model at design time and apply identical computation regardless of input difficulty, cannot deco…

  789. arXiv cs.AI TIER_1 English(EN) · Oliver Aleksander Larsen, Mahyar T. Moghaddam ·

    Agentic AI 采用下的软件架构质量挖掘:一项 Java 代码库的因果研究

    arXiv:2606.13298v1 Announce Type: cross Abstract: AI coding tools are now used by a majority of developers, and agentic use of these tools has popularized the practice colloquially called "vibe coding". Yet causal evidence on their effect on software architecture is scarce. Prior…

  790. arXiv cs.AI TIER_1 English(EN) · Jiaqi Luo, Jiarun Dai, Zhile Chen, Jia Xu, Weibing Wang, Yawen Duan, Brian Tse, Geng Hong, Xudong Pan, Yuan Zhang, Min Yang ·

    大型语言模型驱动的AI系统中自主渗透能力的出现

    arXiv:2606.13079v1 Announce Type: cross Abstract: Nowadays, the autonomous execution of cyberattacks capable of causing substantial real-world harm is widely regarded as one of the critical red lines that frontier AI systems must not cross. Within this broader red-line scenario, …

  791. arXiv cs.AI TIER_1 English(EN) · Jie Wang ·

    面向AI增强计算的Token复杂度理论

    arXiv:2606.12647v1 Announce Type: cross Abstract: AI-augmented computing delegates natural language queries, code generation requests, and other open-ended tasks to a cluster of AI models that processes queries and generates responses. This paradigm introduces a resource dimensio…

  792. arXiv cs.AI TIER_1 English(EN) · Quanyan Zhu ·

    智能体AI的互联网:大规模通信、协调与集体智能

    arXiv:2606.12835v1 Announce Type: cross Abstract: The rapid emergence of autonomous AI agents is transforming artificial intelligence from isolated model inference into distributed systems of reasoning, communication, and action. This paper develops the vision of the Internet of …

  793. arXiv cs.AI TIER_1 English(EN) · Zixing Lei, Genjia Liu, Yuanshuo Zhang, Qipeng Liu, Yuzhu Cai, Sixiang Chen, Jixian Wu, Yunhong Wang, Weixin Li, Chuan Wen, Bo Zhao, Shanghang Zhang, Wenzhao Lian, Siheng Chen ·

    从数字到实体:数字代理作为实体智能的自主教练

    arXiv:2601.21570v2 Announce Type: replace Abstract: The field of Embodied AI is witnessing a rapid evolution toward general-purpose robotic systems, fueled by high-fidelity simulation and large-scale data collection. However, this scaling capability remains severely bottlenecked …

  794. arXiv cs.AI TIER_1 English(EN) · Shayan Kiyani, Sima Noorani, George Pappas, Hamed Hassani ·

    人工智能代理的战略决策支持

    arXiv:2606.12587v1 Announce Type: new Abstract: Traditionally, decision support studies how humans use machine learning models to make better decisions. In modern agentic systems, this division of roles is increasingly reversed: AI agents act on behalf of users, while humans and …

  795. arXiv cs.AI TIER_1 English(EN) · Md Jafrin Hossain, Mohammad Arif Hossain, Weiqi Liu, Nirwan Ansari ·

    收敛鸿沟:已部署的代理式AI框架如何未能满足面向公众的安全要求

    arXiv:2606.12797v1 Announce Type: new Abstract: Agentic large language model systems that autonomously invoke tools, maintain persistent memory, and execute multi-step plans are increasingly deployed in public-facing domains, including government services, healthcare triage, and …

  796. arXiv cs.AI TIER_1 English(EN) · Il-Seok Oh ·

    关于世界模型与物理AI的教程

    arXiv:2606.12783v1 Announce Type: new Abstract: World modeling is emerging as a central principle for building intelligent systems capable of prediction, reasoning, and decision making. A central distinction can be drawn between explicit world models, which learn structured dynam…

  797. arXiv cs.AI TIER_1 English(EN) · Tianyu Liu, Allen Xin Wang, Antonia Panescu, Lisa Xinyi Chen, Wenxin Long, Xinyu Wei, Yueqian Jing, Ziyao Zeng, Jihang Chen, Sihan Jiang, Ziqing Wang, Siyi Gu, Siyu Chen, Xinyang Hu, Haoran Shao, Leqi Xu, Wangjie Zheng, Zhiyuan Cao, Ada Fang, Botao Yu, K… ·

    跨尺度解决科学挑战的AI Agent基准测试

    arXiv:2606.12736v1 Announce Type: new Abstract: AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood. Existing benchmarks for AI agents rarely capture the complexity, het…

  798. Hugging Face Daily Papers TIER_1 English(EN) ·

    从聊天机器人到数字同事:迈向持久自主AI的范式转变

    Large Language Models are evolving from conversational systems to integrated AI colleagues with enhanced reasoning capabilities and persistent work environments.

  799. arXiv cs.AI TIER_1 English(EN) · Mahyar T. Moghaddam ·

    Agentic AI 采用下的软件架构质量挖掘:一项 Java 代码库的因果研究

    AI coding tools are now used by a majority of developers, and agentic use of these tools has popularized the practice colloquially called "vibe coding". Yet causal evidence on their effect on software architecture is scarce. Prior causal work has measured code-level outcomes (com…

  800. arXiv cs.LG TIER_1 English(EN) · Felipe Oviedo, Fiodar Kazhamiaka, Esha Choukse, Allen Kim, Amy Luers, Melanie Nakagawa, Ricardo Bianchini, Juan M. Lavista Ferres ·

    AI推理的能源消耗、效率途径和测试时间缩放

    arXiv:2509.20241v2 Announce Type: replace Abstract: As AI inference scales to billions of queries, estimates of per-query energy use are increasingly important for capacity planning, efficiency interventions, and policy. Yet many public estimates assume non-production settings, l…

  801. arXiv cs.AI TIER_1 English(EN) · Roxana Geambasu, Mariana Raykova, Pierre Tholoniat, Trishita Tiwari, Lillian Tsai, Wen Zhang ·

    通过AI工作流商店为个人代理构建稳健性

    arXiv:2605.10907v3 Announce Type: replace-cross Abstract: The dominant paradigm for AI agents is an "on-the-fly" loop in which agents synthesize plans and execute actions within seconds or minutes in response to user prompts. We argue that this paradigm short-circuits disciplined…

  802. arXiv cs.LG TIER_1 English(EN) · Frank Xiao, Mary Phuong ·

    引导式监控:利用透明推理来监督更强大的AI代理

    arXiv:2606.11998v1 Announce Type: new Abstract: Trusted monitoring is a cornerstone of AI control. However, as frontier models grow more capable, the increasing capabilities gap between trusted and untrusted models may render trusted models unreliable monitors. We introduce \emph…

  803. arXiv cs.AI TIER_1 English(EN) · Hayoung Jung, Pedro Viana Diniz, Jos\'e Reinaldo Corr\^ea Roveda, Abner Fernandes da Silva, Haeun Jung, Enoch Tsai, Aleksandra Korolova, Manoel Horta Ribeiro ·

    AI代理能否综合科学结论?

    arXiv:2606.11337v1 Announce Type: new Abstract: Scientific AI agents increasingly retrieve evidence, reason across sources, and synthesize conclusions used in consequential decisions. Yet, their ability to do so in high-stakes domains such as health remains unclear. We introduce …

  804. arXiv cs.AI TIER_1 English(EN) · Marc Alier Forment, Juanan Pereira, Francisco Jos\'e Garc\'ia-Pe\~nalvo, Mar\'ia Jos\'e Casa\~n Guerrero ·

    Agents All the Way Down;一种从底层到生产构建定制化AI代理的方法论

    arXiv:2606.11869v1 Announce Type: cross Abstract: Custom AI agents areagents that live inside their own application, talk to their own data and tools, enforce their own security boundaries, and carry their own brand and audit trail. What separates them from the general-purpose ti…

  805. arXiv cs.AI TIER_1 English(EN) · Arijit Khan, Longxu Sun, Xin Huang ·

    大语言模型+图:迈向原生图、协同式人工智能系统

    arXiv:2606.11560v1 Announce Type: cross Abstract: Large Language Models (LLMs) have advanced rapidly, but their limitations in structured and multi-hop reasoning underscore the need for graph-native, synergistic artificial intelligence (AI) systems. Graph-structured data underpin…

  806. arXiv cs.AI TIER_1 English(EN) · Michelle Vaccaro ·

    AI Agents 实验预注册

    arXiv:2606.11217v1 Announce Type: cross Abstract: The proliferation of large language models (LLMs) and autonomous AI agents has given rise to a rapidly growing methodological paradigm: "in silico" behavioral experiments. Originally conceived as a way to use AI agents as proxies …

  807. arXiv cs.AI TIER_1 English(EN) · Krti Tallam ·

    面向生产环境中AI代理运行时治理的五架飞机参考架构

    arXiv:2606.12320v1 Announce Type: new Abstract: Enterprise security was built to govern data boundaries: the protected surface was data at rest and in transit, and the controls -- access control, data-loss prevention, perimeter inspection -- governed crossings of that boundary. P…

  808. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Quanyan Zhu ·

    The Internet of Agentic AI: Communication, Coordination, and Collective Intelligence at Scale

    The rapid emergence of autonomous AI agents is transforming artificial intelligence from isolated model inference into distributed systems of reasoning, communication, and action. This paper develops the vision of the Internet of Agentic AI (IoAI): an open ecosystem in which hete…

  809. arXiv cs.AI TIER_1 English(EN) · Krti Tallam ·

    面向生产AI代理运行时治理的五平面参考架构

    Enterprise security was built to govern data boundaries: the protected surface was data at rest and in transit, and the controls -- access control, data-loss prevention, perimeter inspection -- governed crossings of that boundary. Production AI agents dissolve this assumption. An…

  810. arXiv cs.LG TIER_1 English(EN) · Mary Phuong ·

    自举式监控:利用透明推理监督更强大的AI代理

    Trusted monitoring is a cornerstone of AI control. However, as frontier models grow more capable, the increasing capabilities gap between trusted and untrusted models may render trusted models unreliable monitors. We introduce \emph{bootstrapped monitoring}, a protocol that addre…

  811. arXiv cs.AI TIER_1 English(EN) · María José Casañ Guerrero ·

    Agents All the Way Down;一种用于构建从底层到生产的定制 AI Agent 的方法论

    Custom AI agents areagents that live inside their own application, talk to their own data and tools, enforce their own security boundaries, and carry their own brand and audit trail. What separates them from the general-purpose tier is fit, not capability: each is built for one j…

  812. arXiv cs.AI TIER_1 English(EN) · Federico Bianchi, Yongchan Kwon, Aneesh Pappu, James Zou ·

    利用野外人工智能代理的集体智能进行新发现

    arXiv:2606.10402v1 Announce Type: cross Abstract: Scientific discovery is often a collective process: researchers share partial results, inspect failed attempts, and build on each other's ideas over long time horizons. Recent AI systems have shown that language-model-based agents…

  813. arXiv cs.AI TIER_1 English(EN) · Muyu He, Anand Kumar, Tsach Mackey, Meghana Rajeev, James Zou, Nazneen Rajani ·

    急躁的用户混淆 AI 代理:用于测试代理的高保真人类特征模拟

    arXiv:2510.04491v3 Announce Type: replace Abstract: Despite rapid progress in building conversational AI agents, robustness is still largely untested. Small shifts in user behavior, such as being more impatient, incoherent, or skeptical, can cause sharp drops in agent performance…

  814. arXiv cs.AI TIER_1 English(EN) · James Pierce, Vaiva Kalnikait\.e, Siddharth Gupta, Brian Granger ·

    人机协作区:设计具有代理式AI的人机协作体验的框架

    arXiv:2606.09848v1 Announce Type: cross Abstract: As generative and agentic AI becomes embedded in everyday products, practitioners face a persistent challenge: how to design human-AI coordination -- the ongoing mutual adjustment between users and AI systems as mediate through in…

  815. Hugging Face Daily Papers TIER_1 English(EN) ·

    大语言模型+图:迈向原生图、协同AI系统

    Large Language Models (LLMs) have advanced rapidly, but their limitations in structured and multi-hop reasoning underscore the need for graph-native, synergistic artificial intelligence (AI) systems. Graph-structured data underpins critical applications across social, biological,…

  816. Hugging Face Daily Papers TIER_1 English(EN) ·

    跨尺度解决科学挑战的AI Agent基准测试

    SciAgentArena presents a comprehensive benchmark for evaluating AI agents in real scientific research scenarios, revealing current limitations in novel insight generation and open-ended problem solving while identifying opportunities for improving agent reliability and autonomy.

  817. arXiv cs.CL TIER_1 English(EN) · James Zou ·

    利用野外人工智能代理的集体智能进行新发现

    Scientific discovery is often a collective process: researchers share partial results, inspect failed attempts, and build on each other's ideas over long time horizons. Recent AI systems have shown that language-model-based agents can make meaningful progress on open scientific p…

  818. arXiv cs.AI TIER_1 English(EN) · Kai A. Horstmann, Ethan Lin, Alice A. Robie, Jennifer J. Sun, Kristin Branson ·

    一项关于在神经科学数据到发现流程中评估AI代理的案例研究

    arXiv:2606.07718v1 Announce Type: new Abstract: Agentic AI tools offer a promising path to automating software development bottlenecks in scientific research pipelines, particularly for stages that take domain experts days to months to build, where scientists care about correctne…

  819. arXiv cs.AI TIER_1 English(EN) · Shangbin Feng, Yike Wang, Weijia Shi, Luke Zettlemoyer, Yejin Choi, Yulia Tsvetkov ·

    扩展模块化AI系统的参与度

    arXiv:2606.07812v1 Announce Type: new Abstract: Humanity is a mosaic of multifaceted talents and needs, and any truly intelligent AI must reflect that richness. Yet the LLMs used by all are built by the few -- a centralized market of monolithic AI models structurally ill-suited t…

  820. arXiv cs.AI TIER_1 English(EN) · Chenglin Yang ·

    AgentTrust:AI代理行为的自改进信任层

    arXiv:2606.08539v1 Announce Type: new Abstract: AI agents increasingly take consequential actions -- shell commands, cloud operations, and arbitrary tool-calls -- so a trust layer must decide, per action, whether to allow, warn, block, or escalate. We argue that the right way to …

  821. arXiv cs.AI TIER_1 English(EN) · Yifan Liu (Klara), Jaime Arguello (Klara), Orland Hoeber (Klara), Chang Liu (Klara), Soo Young Rieh (Klara), Luanne Sinnamon (Klara), Dean Alvarez (Klara), Susan Archambault (Klara), Rob Capra (Klara), Henson Chen (Klara), Charles Costa (Klara), Anita Cr… ·

    关于生成式人工智能与学术搜索(GAI&AS)研讨会CHIIR 2026的报告

    arXiv:2606.08936v1 Announce Type: cross Abstract: This report summarizes the CHIIR 2026 Workshop on Generative AI and Academic Search (GAI\&amp;AS), which examined how GenAI is reshaping academic search systems and research practices. The workshop brought together researchers in …

  822. arXiv cs.AI TIER_1 English(EN) · Abhinav Mishra, Kumar Sharad ·

    Agentic AI系统中委托执行的可观测性

    arXiv:2606.09692v1 Announce Type: cross Abstract: Delegation-scoped execution is not identifiable from standard observables: audit logs and execution traces can be identical under multiple incompatible delegation assignments. This gap is especially acute in LLM-based agentic syst…

  823. arXiv cs.AI TIER_1 English(EN) · Jun Takahashi, Atsunori Moteki, Akiyoshi Uchida, Shoichi Masui, Fan Yang, Kanji Uchino, Yueqi Song, Yonatan Bisk, Graham Neubig, Ikuo Kusajima, Yasuto Watanabe, Hiroyuki Ishida, Koki Nakagawa, Shan Jiang ·

    FieldWorkArena: 面向真实现场工作任务的代理AI基准测试

    arXiv:2505.19662v4 Announce Type: replace Abstract: This paper introduces FieldWorkArena, a benchmark for agentic AI targeting real-world field work. With the recent increase in demand for agentic AI, they are built to detect and document safety hazards, procedural violations, an…

  824. arXiv cs.AI TIER_1 English(EN) · Yunpeng Dong, Jingkai He, Shiqi Liu, Yuze Hou, Dong Du, Zhonghu Xu, Si Yu, Baochuan Yang, Yubin Xia, Haibo Chen ·

    DeltaBox:实现毫秒级沙箱检查点/回滚,扩展有状态AI代理

    arXiv:2605.22781v2 Announce Type: replace-cross Abstract: LLM-powered AI agents require high-frequency state exploration (e.g., test-time tree search and reinforcement learning), relying on rapid checkpoint and rollback (C/R) of the complete sandbox state, including files and pro…

  825. arXiv cs.LG TIER_1 English(EN) · Neel Tushar Shah, Manglam Kartik ·

    AI科学家何时应停止?可验证的实验引导与自主发现的拒绝

    arXiv:2606.07576v1 Announce Type: new Abstract: We present CARTOGRAPH, a verification layer for AI scientists that couples unresolved-subspace experiment steering (select), explicit ambiguity closure (resolve), and residual-based library inadequacy detection (refuse). Under a loc…

  826. arXiv cs.AI TIER_1 English(EN) · Muhammad Haris Khan, Joel wester ·

    洞察蜂群思维:一种共识感知交互技术,用于缓解人工智能同质化

    arXiv:2606.09587v1 Announce Type: cross Abstract: People are increasingly using AI for creative tasks such as writing. While adoption continues to grow, this form of use risks undermining individual creativity locally and reducing the heterogeneity of creative output at scale. In…

  827. arXiv cs.AI TIER_1 English(EN) · Muhammad Zia Hydari, Raja Iqbal ·

    未被选择的Token:采样、状态与AI代理输出的可变性

    arXiv:2606.08998v1 Announce Type: new Abstract: Agentic AI systems can behave differently across runs: the same request may produce a different plan, a different tool call, a different code edit, or a different final answer. Such variability arises from several layers that are of…

  828. arXiv cs.AI TIER_1 English(EN) · Ian Seet, Jonas Bozenhard, Simon Osterman ·

    通过本地化架构增强AI的可解释性和安全性

    arXiv:2606.07998v1 Announce Type: cross Abstract: Recent advances in generative AI, especially powerful Large Language Models (LLMs) and Large Reasoning Models (LRMs), raise concerns over the interpretability, safety and sustainability of these large and opaque AI models. The pow…

  829. arXiv cs.AI TIER_1 English(EN) · Ehud Shapiro ·

    使用多智能体迁移系统和人工智能实现基层逻辑程序(完整版)

    arXiv:2602.06934v4 Announce Type: replace-cross Abstract: Grassroots Logic Programs (GLP) is a concurrent logic programming language in which logic variables are partitioned into paired readers and writers. An assignment is produced at most once via a writer and consumed at most …

  830. arXiv cs.AI TIER_1 English(EN) · Rishabh Sabharwal, Hongru Wang, Amos Storkey, Jeff Z. Pan ·

    深度研究代理在过程级反馈下的多轮评估

    arXiv:2606.09748v1 Announce Type: new Abstract: Existing benchmarks for deep research agents (DRAs) assess only single-shot outputs, ignoring a key question: can DRAs improve their reports when guided by feedback? To investigate this, we conduct a multi-turn evaluation of DRAs un…

  831. arXiv cs.AI TIER_1 English(EN) · Jeff Z. Pan ·

    深度研究代理在过程级反馈下的多轮评估

    Existing benchmarks for deep research agents (DRAs) assess only single-shot outputs, ignoring a key question: can DRAs improve their reports when guided by feedback? To investigate this, we conduct a multi-turn evaluation of DRAs under two feedback settings: self-reflection, in w…

  832. arXiv cs.AI TIER_1 English(EN) · Kumar Sharad ·

    Agentic AI系统中委托执行的可观测性

    Delegation-scoped execution is not identifiable from standard observables: audit logs and execution traces can be identical under multiple incompatible delegation assignments. This gap is especially acute in LLM-based agentic systems, where agents dynamically select tools, vary e…

  833. arXiv cs.AI TIER_1 English(EN) · Joel wester ·

    洞察蜂群思维:一种共识感知交互技术,以缓解人工智能同质化

    People are increasingly using AI for creative tasks such as writing. While adoption continues to grow, this form of use risks undermining individual creativity locally and reducing the heterogeneity of creative output at scale. In response, we introduce the Semantic Repulsion Tec…

  834. arXiv cs.AI TIER_1 English(EN) · Hariom Tatsat, Ariye Shater ·

    超越黑箱:Agentic AI工具使用的可解释性

    arXiv:2605.06890v3 Announce Type: replace Abstract: AI agents are promising for high-stakes enterprise workflows, but dependable deployment remains limited because tool-use failures are difficult to diagnose and control. Agents may skip required tool calls, invoke tools unnecessa…

  835. arXiv cs.AI TIER_1 English(EN) · Josef Chen ·

    AEGIS:物理AI的备份反射

    arXiv:2606.06660v1 Announce Type: new Abstract: Long-horizon robot manipulation tends to fail gradually: one bad step degrades the state, and the policy spirals into a basin from which it cannot recover. The failure is often visible before it happens. We introduce AEGIS (Activati…

  836. arXiv cs.AI TIER_1 English(EN) · Catherine Ge-Wang, Tyler Crosse, Benjamin Hadad IV, Joachim Schaeffer, Ram Potham, Tyler Tracy ·

    Agentic AI 控制评估中的攻击选择有意义地降低了安全性

    arXiv:2606.06529v1 Announce Type: new Abstract: An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately. AI control is a safety framework for deploying capable but untrusted AI agents under the oversight of a weaker, tr…

  837. arXiv cs.AI TIER_1 English(EN) · Gangda Deng, Zhaoling Chen, Zhongming Yu, Haoyang Fan, Yuhong Liu, Yuxin Yang, Dhruv Parikh, Rajgopal Kannan, Le Cong, Mengdi Wang, Qian Zhang, Viktor Prasanna, Xiangru Tang, Xingyao Wang ·

    EvoClaw:评估AI代理在持续软件演化中的表现

    arXiv:2603.13428v2 Announce Type: replace-cross Abstract: With AI agents increasingly deployed as long-running systems, it becomes essential to autonomously construct and continuously evolve customized software to enable interaction within dynamic environments. Yet, existing benc…

  838. arXiv cs.AI TIER_1 English(EN) · Jeremy Yang, Kate Zyskowski, Noah Yonack, Jerry Ma ·

    人工智能代理如何重塑知识工作:自主性、效率和范围

    arXiv:2606.07489v1 Announce Type: new Abstract: Frontier AI systems are bridging the gap between intelligence and utility by shifting from conversational assistants to autonomous agents that execute tasks end to end. Using production data from Perplexity's Search and Computer pro…

  839. arXiv cs.AI TIER_1 English(EN) · M. Danish Lim, I. Danial Bin Sharudin, Wen Han Chen, Cedric Lim, Laura Wynter ·

    面向知识驱动的工具使用工作流的AI代理声明式技能

    arXiv:2606.06923v1 Announce Type: new Abstract: We study orchestration mechanisms for tool-using AI agents in realistic customer-service workflows over an unstructured knowledge base. We argue that declarative agents -- AI agents equipped with natural-language skill files appende…

  840. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Dan Zhang ·

    关于生成式人工智能与学术搜索(GAI&AS)研讨会CHIIR 2026的报告

    This report summarizes the CHIIR 2026 Workshop on Generative AI and Academic Search (GAI\&AS), which examined how GenAI is reshaping academic search systems and research practices. The workshop brought together researchers in human information interaction and information retrieva…

  841. arXiv cs.AI TIER_1 English(EN) · Chenglin Yang ·

    AgentTrust:AI智能体动作的自改进信任层

    AI agents increasingly take consequential actions -- shell commands, cloud operations, and arbitrary tool-calls -- so a trust layer must decide, per action, whether to allow, warn, block, or escalate. We argue that the right way to reason about such a layer is by threat type. Lex…

  842. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Rahemeen Khan ·

    迈向以人为本的多智能体系统:将认知、文化、价值观与合作融入AI智能体

    The emergence of large language model (LLM)-based agents and multi-agent systems has enabled a shift from narrow task automation to more autonomous decision-making. Despite progress in language generation, planning, tool use, and coordination, most agents still treat intelligence…

  843. arXiv cs.AI TIER_1 English(EN) · Quanyan Zhu ·

    Agentic AI 保险

    arXiv:2606.05449v1 Announce Type: new Abstract: Agentic artificial intelligence (AI) systems are transforming the risk landscape by extending beyond information generation to autonomous planning, tool invocation, decision execution, and persistent modification of digital and phys…

  844. arXiv cs.AI TIER_1 English(EN) · Zhenfeng Cao ·

    软件工程的终结:AI Agent 如何从根本上重塑软件范式

    arXiv:2606.05608v1 Announce Type: cross Abstract: For over half a century, software engineering has operated on a foundational premise: human engineers decompose problems, encode decision logic into static code, and manually adapt that code as requirements evolve. This paper argu…

  845. arXiv cs.AI TIER_1 English(EN) · Gal Bakal ·

    知识激活:AI技能作为代理软件开发中的机构知识原始要素

    arXiv:2603.14805v2 Announce Type: replace Abstract: Enterprise software organizations accumulate critical institutional knowledge - architectural decisions, deployment procedures, compliance policies, incident playbooks - yet this knowledge remains trapped in formats designed for…

  846. arXiv cs.AI TIER_1 English(EN) · Yunhao Yang, Neel P. Bhatt, Kevin Wang, Samuel Tetteh, Zhangyang Wang, Ufuk Topcu ·

    VASO:面向物理AI智能体的形式化可验证的自演化技能

    arXiv:2606.05395v1 Announce Type: cross Abstract: Reusable robot skills are becoming the basic units through which embodied agents turn open-ended instructions into long-horizon physical behavior. We argue that, while foundation models have collapsed the cost of creating these sk…

  847. arXiv cs.AI TIER_1 English(EN) · Jerry Ma ·

    人工智能代理如何重塑知识工作:自主性、效率和范围

    Frontier AI systems are bridging the gap between intelligence and utility by shifting from conversational assistants to autonomous agents that execute tasks end to end. Using production data from Perplexity's Search and Computer products, we study this transition by examining how…

  848. arXiv cs.AI TIER_1 English(EN) · Laura Wynter ·

    面向知识增强工具使用工作流的AI代理声明式技能

    We study orchestration mechanisms for tool-using AI agents in realistic customer-service workflows over an unstructured knowledge base. We argue that declarative agents -- AI agents equipped with natural-language skill files appended to the system prompt -- are an effective orche…

  849. arXiv cs.LG TIER_1 English(EN) · Otto Nyberg, Fausto Carcassi, Davide Tugnoli, Giovanni Cin\`a ·

    2-Step Agent:决策者与AI决策支持交互的框架

    arXiv:2602.21889v2 Announce Type: replace-cross Abstract: Predictions from ML models support human decision making in several fields, including high-stakes ones such as healthcare and the judiciary. Yet, we still lack a clear understanding of how decision makers learn from ML-bas…

  850. Hugging Face Daily Papers TIER_1 English(EN) ·

    基于熵的AI代理评估:一种衡量行为模式的轻量级框架

    AI agents are commonly evaluated using task success, reward, latency, and cost. These metrics are useful, but they often miss important aspects of agent behavior: whether an agent explores too much, repeats itself too rigidly, uses tools effectively, reduces uncertainty over time…

  851. arXiv cs.AI TIER_1 English(EN) · Rubens Lacerda Queiroz, Cabral Lima, Fabio Ferrentini Sampaio, Priscila Machado Vieira Lima ·

    机器如何学习?评估AIcon2abs方法

    arXiv:2401.07386v5 Announce Type: cross Abstract: This study expands on previous work that introduced the AIcon2abs method (AI from Concrete to Abstract: Demystifying Artificial Intelligence to the general public), an innovative approach designed to increase public understanding …

  852. arXiv cs.AI TIER_1 English(EN) · Harsha Vardhan Khurdula, Vineet Agarwal, Yoeven D Khemlani ·

    Interfaze:AI的未来建立在特定任务的小型模型之上

    arXiv:2602.04101v2 Announce Type: replace Abstract: We present Interfaze, a native hybrid model that fuses task-specific deep neural networks (CNNs and DNNs) directly into a transformer decoder through a shared embedding space. Specialized perceptual encoders handle optical chara…

  853. arXiv cs.AI TIER_1 English(EN) · Rubens Lacerda Queiroz, F\'abio Ferrentini Sampaio, Cabral Lima, Priscila Machado Vieira Lima ·

    AI从具体到抽象:向公众揭秘人工智能

    arXiv:2006.04013v6 Announce Type: cross Abstract: Artificial Intelligence (AI) has been adopted in a wide range of domains. This shows the imperative need to develop means to endow common people with a minimum understanding of what AI means. Combining visual programming and WiSAR…

  854. arXiv cs.AI TIER_1 English(EN) · Ulbert Jose Botero, Liam Smith, Brooks Olney, Pooya Khorrami, Steven Kusiak, Watson Jia, Sage Trudeau, Daniel Capecci ·

    构建机器学习的 Ph(ysical)AI 层

    arXiv:2606.04106v1 Announce Type: cross Abstract: Foundation models achieve generalization through massive-scale training on diverse data, but have limitations with transfer to truly unseen domains without paired training data. We propose principle-driven foundation models that e…

  855. arXiv cs.AI TIER_1 English(EN) · Sanderson Oliveira de Macedo ·

    从提示到流程:支持人工智能软件开发代理的框架的流程分类和比较评估

    arXiv:2606.04967v1 Announce Type: cross Abstract: AI tools for programming are no longer just autocomplete or chat assistants: they organize themselves as development frameworks, with process, roles, artifacts and verification. Recent surveys map agents and LLMs for software engi…

  856. arXiv cs.AI TIER_1 English(EN) · Arquimedes Canedo, Grama Chethan ·

    自省式API:结构胜于冗余,助力AI代理恢复

    arXiv:2606.05037v1 Announce Type: cross Abstract: When an AI agent calls an API and hits a validation error, it needs more than what went wrong -- it needs what to do next. A self-reflective API returns, on validation failure, a machine-readable recovery\_feedback.suggestions[] p…

  857. arXiv cs.AI TIER_1 English(EN) · Travis Weber, Rohit Taneja ·

    数字学徒:一种人类指导的自主AI开发框架

    arXiv:2606.04321v1 Announce Type: new Abstract: Agentic AI deployments face a recurring design tension: heavy human oversight limits scale, while broad autonomy outruns accountability. Neither posture provides the governance infrastructure required for responsible delegation. We …

  858. arXiv cs.AI TIER_1 English(EN) · Katherine M. Collins, Simon Frieder, Jonas Bayer, Jacob Loader, Jeck Lim, Peiyang Song, Fabian Zaiser, Lexin Zhou, Shanda Li, Sam Looi, Joshua B. Tenenbaum, Umang Bhatt, Adrian Weller, Jose Hernandez-Orallo, Cameron E. Freer, Valerie Chen, Ilia Sucholuts… ·

    表征初始人类-AI证明形式化工作流

    arXiv:2606.04273v1 Announce Type: new Abstract: For centuries, human mathematicians have written proofs to substantiate their mathematical arguments; yet, the ability to automatically verify the validity of proofs has long been a challenge. Advances in AI systems' ability to gene…

  859. arXiv cs.AI TIER_1 English(EN) · Andrea Ferrario ·

    人类-AI交互中多智能体互补性的基于树的形式化

    arXiv:2606.04779v1 Announce Type: new Abstract: Complementarity is the case in which a human--AI interaction (HAI) outperforms the best prediction benchmark available among its members. Although this idea is central in HAI research, formal work on complementarity remains limited.…

  860. Hugging Face Daily Papers TIER_1 English(EN) ·

    ForeSci:评估大型语言模型代理在面向未来的AI研究判断中的作用

    ForeSci is a temporally controlled benchmark that evaluates LLM agents' ability to make forward-looking research decisions from historical evidence across fast-moving AI domains.

  861. arXiv cs.AI TIER_1 English(EN) · Grama Chethan ·

    自省式API:结构胜过冗长,助力AI代理恢复

    When an AI agent calls an API and hits a validation error, it needs more than what went wrong -- it needs what to do next. A self-reflective API returns, on validation failure, a machine-readable recovery\_feedback.suggestions[] payload sufficient for the agent to repair the requ…

  862. Hugging Face Daily Papers TIER_1 English(EN) ·

    自省式API:结构胜于冗余,助力AI智能体恢复

    When an AI agent calls an API and hits a validation error, it needs more than what went wrong -- it needs what to do next. A self-reflective API returns, on validation failure, a machine-readable recovery\_feedback.suggestions[] payload sufficient for the agent to repair the requ…

  863. arXiv cs.AI TIER_1 English(EN) · Sanderson Oliveira de Macedo ·

    从提示到流程:支持人工智能软件开发代理的框架的流程分类和比较评估

    AI tools for programming are no longer just autocomplete or chat assistants: they organize themselves as development frameworks, with process, roles, artifacts and verification. Recent surveys map agents and LLMs for software engineering, but a study centered on the operational f…

  864. arXiv cs.AI TIER_1 English(EN) · Andrea Ferrario ·

    人类-AI交互中多智能体互补性的基于树的形式化

    Complementarity is the case in which a human--AI interaction (HAI) outperforms the best prediction benchmark available among its members. Although this idea is central in HAI research, formal work on complementarity remains limited. Existing frameworks do not model how agents' pr…

  865. arXiv cs.AI TIER_1 English(EN) · Stephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu, Saiteja Utpala, Arvind Narayanan ·

    迈向人工智能代理可靠性科学

    arXiv:2602.16666v3 Announce Type: replace Abstract: AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on standard benchmarks suggest rapid progress, many agents still continue to fail in practice. This discrepancy highlights a fundamenta…

  866. arXiv cs.AI TIER_1 English(EN) · Marcus R\"ub, Michael Gerhards ·

    面向边缘嵌入式AI代理系统的模块化架构

    arXiv:2606.02862v1 Announce Type: new Abstract: The rise of Large Language Models (LLMs) has enabled agentic AI capable of complex reasoning and tool use; however, deploying such autonomy in pervasive computing environments remains challenging due to the strict memory and energy …

  867. arXiv cs.AI TIER_1 English(EN) · Xuanqiang Angelo Huang, Charlie Tharas, Samuele Marro, Van Q. Truong, Bernhard Sch\"olkopf, Emanuele La Malfa, Zhijing Jin ·

    机制设计不足以实现:促进合作式人工智能的亲社会智能体

    arXiv:2605.08426v2 Announce Type: replace-cross Abstract: Ensuring that AI agents behave safely and beneficially when interacting with other parties has emerged as one of the central challenges of modern AI safety. While mechanism design, as the theory of designing rules to align…

  868. arXiv cs.AI TIER_1 English(EN) · Amjad Ibrahim, Yong Li ·

    覆盖式治理:用于代理式AI中委托和范围的组合式授权框架

    arXiv:2606.03518v1 Announce Type: new Abstract: As AI systems evolve from passive models into autonomous active agents capable of initiating actions, collaborating, and delegating tasks, the traditional boundaries of software systems blur. Traditional authorization and delegation…

  869. arXiv cs.AI TIER_1 English(EN) · Fiona Y. Wang, Markus J. Buehler ·

    用于科学的自我修正发现系统:一种代理式人工智能的分类框架

    arXiv:2606.01444v1 Announce Type: new Abstract: Scientific discovery is not only answer generation but revision of the representational regime in which evidence, artifacts, operations, and verifiers are typed. We develop a category-theoretic account of agentic discovery for mater…

  870. arXiv cs.AI TIER_1 English(EN) · Sindhuja Chaduvula, Jessee Ho, Kina Kim, Aravind Narayanan, Ahmed Y. Radwan, Mahshid Alinoori, Muskan Garg, Dhanesh Ramachandram, Shaina Raza ·

    从特征到行动:传统AI与Agentic AI系统的可解释性

    arXiv:2602.06841v4 Announce Type: replace Abstract: Over the last decade, Explainable AI has primarily focused on interpreting individual model predictions, producing post-hoc explanations that relate inputs to outputs under a fixed decision structure. Recent advances in large la…

  871. arXiv cs.AI TIER_1 English(EN) · An Luo, Jin Du, Xun Xian, Robert Specht, Fangqiao Tian, Ganghua Wang, Xuan Bi, Charles Fleming, Ashish Kundu, Jayanth Srinivasa, Mingyi Hong, Rui Zhang, Tianxi Li, Galin Jones, Jie Ding ·

    AgentDS 技术报告:领域特定数据科学中人机协作未来的基准测试

    arXiv:2603.19005v2 Announce Type: replace-cross Abstract: Data science plays a critical role in transforming complex data into actionable insights across numerous domains. Recent developments in large language models (LLMs) and artificial intelligence (AI) agents have significant…

  872. arXiv cs.AI TIER_1 English(EN) · Barak Or ·

    物理AI中的静默故障:运行时动作授权在自主系统中的文献综述

    arXiv:2606.00090v1 Announce Type: cross Abstract: Physical AI systems increasingly map multimodal observations, language instructions, and learned world representations into physically consequential actions. Robotics foundation models, vision-language-action models, and world-mod…

  873. arXiv cs.AI TIER_1 English(EN) · Kevin Kappelmann, Maximilian Sch\"affeler, Lukas Stevens, Mohammad Abdulaziz, Andrei Popescu, Dmitriy Traytel ·

    只需在Isabelle中输入!AI代理根据人类提示进行起草、机械化和泛化

    arXiv:2604.15713v2 Announce Type: replace-cross Abstract: Type annotations are essential when printing terms in a way that preserves their meaning under reparsing and type inference. We study the problem of complete and minimal type annotations for rank-one polymorphic $\lambda$-…

  874. arXiv cs.AI TIER_1 English(EN) · Qiuyu Tian, Zequn Liu, Yingce Xia, Haojie Yin, Youyong Kong ·

    ForeSci:评估大型语言模型代理在面向未来的AI研究判断中的作用

    arXiv:2606.00644v1 Announce Type: new Abstract: AI research often requires decisions before future evidence exists: which bottleneck to attack, which direction to pursue, or where a project should be positioned. We introduce ForeSci, a temporally controlled benchmark for evaluati…

  875. 量子位 (QbitAI) TIER_1 中文(ZH) · 量子位的朋友们 ·

    Qwen3.7-Plus发布!多模态智能体新基石,一键复刻专业桌面软件

    Qwen3.7-Plus已上线阿里云百炼

  876. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Michael Gerhards ·

    面向边缘嵌入式AI代理系统的模块化架构

    The rise of Large Language Models (LLMs) has enabled agentic AI capable of complex reasoning and tool use; however, deploying such autonomy in pervasive computing environments remains challenging due to the strict memory and energy constraints of embedded microcontrollers. Existi…

  877. arXiv cs.AI TIER_1 English(EN) · Jun Rui Huang, Wang Bill Zhu, Ziyi Liu, Nathanael Fast, Ravi Iyer, Robin Jia ·

    EUDAIMONIA:评估人工智能中的不良动态

    arXiv:2605.30654v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as conversational partners for companionship, emotional disclosure, and interpersonal advice, but the social dynamics of these interactions can create harms that are not captured …

  878. arXiv cs.AI TIER_1 English(EN) · David Fern\'andez-Narro, Pablo Ferri, \'Angel S\'anchez-Garc\'ia, Juan M. Garc\'ia-G\'omez, Carlos S\'aez ·

    dashi:一个用于数据集偏移表征的Python库,以支持可信赖的AI开发和部署

    arXiv:2605.31360v1 Announce Type: cross Abstract: The Artificial Intelligence (AI) life cycle requires a thorough understanding of the underlying data dynamics for robust, safe and cost-effective AI development and use. Dataset shifts are defined as changes between train and test…

  879. arXiv cs.AI TIER_1 English(EN) · Carlos Sáez ·

    dashi: 一个用于数据集偏移特征化的 Python 库,以支持可信赖的 AI 开发和部署

    The Artificial Intelligence (AI) life cycle requires a thorough understanding of the underlying data dynamics for robust, safe and cost-effective AI development and use. Dataset shifts are defined as changes between train and test data distributions. Whether occurring over time (…

  880. 量子位 (QbitAI) TIER_1 中文(ZH) · 量子位的朋友们 ·

    Moonshot AI “开源周”:定义边缘AI终极形态的系统性“实力展示”

    端侧 AI 是一个系统性工程

  881. arXiv cs.AI TIER_1 English(EN) · Gianluca Inguglia ·

    首次将agentic AI应用于爱因斯坦望远镜模拟数据分析的直接对比研究

    arXiv:2605.28916v1 Announce Type: cross Abstract: We report a comparison of two state-of-the-art agentic AI systems, Claude Code (Anthropic) and Codex (OpenAI), tasked with autonomously executing a simple end-to-end gravitational wave data analysis pipeline on a shared computing …

  882. arXiv cs.AI TIER_1 English(EN) · Tianhua Chen ·

    生成式AI基础小册:直观的数学入门

    arXiv:2605.29713v1 Announce Type: cross Abstract: This book provides a compact, derivation-oriented introduction to the mathematical foundations of modern generative artificial intelligence. Rather than surveying every recent architecture or implementation detail, it develops a c…

  883. arXiv cs.AI TIER_1 English(EN) · Lorenz Kutschka, Bernhard Geiger ·

    符号很重要:代理式AI系统中令牌优化格式的基准研究

    arXiv:2605.29676v1 Announce Type: new Abstract: Large language models in Agentic AI systems consume tool schemas and execution results and emit tool invocations as structured data. The default language for that exchange, JSON, was designed for application-to-application interchan…

  884. arXiv cs.AI TIER_1 English(EN) · Muhammad Zia Hydari, Raja Iqbal, Narayan Ramasubbu ·

    管理代理式AI系统中的技术债务

    arXiv:2605.29129v1 Announce Type: new Abstract: Agentic AI systems are increasingly being explored as production infrastructure: they reason over multiple steps, call tools, act through workflows, and adapt through memory and feedback. These systems create governance challenges t…

  885. arXiv cs.AI TIER_1 English(EN) · William Yicheng Zhu, Lei Zhu ·

    人工智能加速的地球成本,第二部分:第十个地球边界与6.5年倒计时

    arXiv:2604.04956v3 Announce Type: replace-cross Abstract: The recent, super-exponential scaling of autonomous Large Language Model (LLM) agents signals a broader, fundamental paradigm shift from machines primarily replacing the human hands (manual labor and mechanical processing)…

  886. arXiv cs.CL TIER_1 English(EN) · Vishakh Padmakumar, Lujain Ibrahim, Zora Zhiruo Wang, Jennifer Wang, Q. Vera Liao, Diyi Yang ·

    卸载分数:通过反事实工作流衡量人工智能依赖性

    arXiv:2605.29392v1 Announce Type: cross Abstract: AI tools are increasingly integrated into real-world workflows. However, existing measures of reliance on these tools focus on AI output adoption or on self-reported indicators, rather than how task effort is distributed between u…

  887. arXiv cs.AI TIER_1 English(EN) · Ruiyi Zhang, Peijia Qin, Qi Cao, Li Zhang, Pengtao Xie ·

    AIBuildAI-2:一个用于自动构建AI模型的增强知识代理

    arXiv:2605.27873v1 Announce Type: new Abstract: AI models underpin data-centric applications from image and text processing to scientific discovery in biology, physics, and chemistry. Yet developing them remains heavily manual, requiring practitioners to design architectures, bui…

  888. arXiv cs.AI TIER_1 English(EN) · Aakash Pant, Kavya Shah, Apoorv Agnihotri, Sneha Nikam, Prasaanth Balraj, Nakul Jain ·

    低资源环境下的AI基准测试:超越排行榜的思考

    arXiv:2605.28508v1 Announce Type: new Abstract: Existing AI evaluation practices often fail to capture how systems actually perform in low-resource environments, where operational constraints shape usability as much as model quality. Through a structured analysis of existing benc…

  889. arXiv cs.AI TIER_1 English(EN) · Yihong Tang, Andrew Robert Williams, Arjun Ashok, Vincent Zhihao Zheng, Lijun Sun, Alexandre Drouin, Issam H. Laradji, \'Etienne Marcotte, Valentina Zantedeschi ·

    Dr-CiK:一个面向未来的驱动代理的测试平台

    arXiv:2605.27904v1 Announce Type: new Abstract: Time series forecasting in real-world settings often depends not only on historical observations, but also on external context that must be actively discovered from noisy, heterogeneous information sources. Yet existing context-aide…

  890. arXiv cs.AI TIER_1 English(EN) · Jaechang Kim, Sunung Mun, Seungjoon Lee, Jaewoong Cho, Jungseul Ok ·

    迈向忠实的 Agentic XAI:一种用于提升模型忠实度的验证方法和开放世界基准

    arXiv:2605.27879v1 Announce Type: new Abstract: Explainable AI (XAI) helps users interpret model behavior and identify potential faults. Agentic XAI systems use Large Language Models (LLMs) to make explanations more accessible through natural-language interaction, but they can al…

  891. arXiv cs.AI TIER_1 English(EN) · Srini Ramaswamy ·

    智能作为一种受控自主性:Agentic AI系统的失败、升级与治理

    arXiv:2605.27628v1 Announce Type: new Abstract: As autonomous and agentic AI systems scale in robotic and human-machine environments, managing hallucination and persistent but unjustified action remains an open challenge. Rather than attributing these failures solely to model or …

  892. arXiv cs.AI TIER_1 English(EN) · Xing Zhang, Guanghui Wang, Yanwei Cui, Wei Qiu, Ziyuan Li, Bing Zhu, Peiyang He ·

    提示优化是抛硬币:诊断其在复合AI系统中何时奏效

    arXiv:2604.14585v2 Announce Type: replace Abstract: Prompt optimization in compound AI systems is statistically indistinguishable from a coin flip: across 72 optimization runs on Claude Haiku 4.5 (6 methods $\times$ 4 tasks $\times$ 3 repeats), 49% score below zero-shot; on Amazo…

  893. arXiv cs.AI TIER_1 English(EN) · Edwin Jose ·

    SwarmHarness:基于技能的去中心化激励对齐AI代理网络任务路由

    arXiv:2605.28764v1 Announce Type: new Abstract: Vast quantities of compute (GPU cycles on personal workstations, idle inference servers, and edge devices between jobs) go unused because no incentive-aligned protocol exists for their owners to share them safely and profitably. Exi…

  894. arXiv cs.LG TIER_1 English(EN) · Bohan Lyu, Yucheng Yang, Siqiao Huang, Jiaru Zhang, Qixin Xu, Xinghan Li, Xinyang Han, Yicheng Zhang, Huaqing Zhang, Runhan Huang, Kaicheng Yang, Zitao Chen, Wentao Guo, Junlin Yang, Xinyue Ai, Wenhao Chai, Yadi Cao, Ziran Yang, Kun Wang, Dapeng Jiang, H… ·

    MLS-Bench:对构建更优AI的AI系统进行全面而严谨的评估

    arXiv:2605.08678v2 Announce Type: replace Abstract: Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonstrate advanced capabilities in reasoning, coding, and engineering tasks, it i…

  895. arXiv cs.AI TIER_1 English(EN) · Nikita Benkovich, Vitalii Valkov ·

    Agyn:一个开源的 AI Agent 平台,支持可扩展的按需执行、代码即代理定义以及零信任访问

    arXiv:2605.27575v1 Announce Type: new Abstract: As organizations move toward production deployments of AI agents, which execute non-deterministic workflows, maintain stateful sessions, and often operate with privileged access to internal services, the engineering challenge shifts…

  896. arXiv cs.AI TIER_1 English(EN) · Edwin Jose ·

    SwarmHarness:基于技能的任务路由,通过去中心化激励对齐的AI代理网络

    Vast quantities of compute (GPU cycles on personal workstations, idle inference servers, and edge devices between jobs) go unused because no incentive-aligned protocol exists for their owners to share them safely and profitably. Existing approaches either require a trusted centra…

  897. NVIDIA Blog TIER_1 English(EN) · Jeremy Graybill ·

    AI工厂:智能的新基础设施

    AI factories are token factories, converting power into intelligence in real time. And as agentic AI scales and autonomous, always-on special agents are deployed in the enterprise, performance per watt and cost per token become the economics that matter.

  898. arXiv cs.AI TIER_1 English(EN) · Nakul Jain ·

    低资源环境下的AI基准测试:超越排行榜的思考

    Existing AI evaluation practices often fail to capture how systems actually perform in low-resource environments, where operational constraints shape usability as much as model quality. Through a structured analysis of existing benchmark families across speech, chat/RAG, and visi…

  899. arXiv cs.AI TIER_1 English(EN) · Judy Fox, Geoffrey Fox ·

    Agentic AI for Science Experiments

    arXiv:2605.26305v1 Announce Type: new Abstract: This paper details two novel frameworks for developing autonomous, agentic AI in scientific workflows. Both systems leverage a hybrid Local Body, Remote Brain architecture via Google Colab, utilizing Python-based local orchestrators…

  900. arXiv cs.AI TIER_1 English(EN) · Hao-Hsuan Chen ·

    面向自主人工智能代理的时间一致反事实精算运行时的基础

    arXiv:2605.26508v1 Announce Type: cross Abstract: We propose a foundational runtime actuarial layer for autonomous AI agents in which every side-effect-bearing action carries a time-consistent, counterfactual risk toll computed against a contractually fixed safe default, inside a…

  901. arXiv cs.AI TIER_1 English(EN) · Xue Qin, Simin Luan, John See, Zeyd Boukhers, Cong Yang, Zhijun Li ·

    受控能力演进:基于AI组件的系统的生命周期兼容性检查与回滚,以具身智能体为例

    arXiv:2604.08059v5 Announce Type: replace-cross Abstract: Software systems built from versioned AI components increasingly need lifecycle-time governance: when a capability module evolves into a new version, the hosting system must decide whether the new version may be activated …

  902. arXiv cs.LG TIER_1 English(EN) · Vasilios A. Siris, Adamantia Stamou, George D. Stamoulis, Konstantinos Varsos, Ramin Khalili ·

    通过兼顾准确性和延迟的用户激励措施实现人工智能推理的绿色化

    arXiv:2605.27309v1 Announce Type: new Abstract: The widespread use of AI services has raised concerns for its environmental sustainability, towards which recent studies have identified carbon emissions of AI inference as the major contributor. This paper introduces a framework fo…

  903. arXiv cs.AI TIER_1 English(EN) · Rui Yang, Qianhui Wu, Zhaoyang Wang, Hanyang Chen, Ke Yang, Hao Cheng, Huaxiu Yao, Baolin Peng, Huan Zhang, Jianfeng Gao, Tong Zhang ·

    GUI-Libra:通过动作感知监督和部分可验证强化学习训练原生GUI智能体进行推理和行动

    arXiv:2602.22190v2 Announce Type: replace-cross Abstract: Open-source native GUI agents still lag behind closed-source systems on long-horizon navigation tasks. This gap stems from two limitations: a shortage of high-quality, action-aligned reasoning data, and the direct adoption…

  904. arXiv cs.AI TIER_1 English(EN) · Anas H. Alzahrani ·

    学术研究中的持久性人工智能代理:单研究员实施案例研究

    arXiv:2605.26870v1 Announce Type: cross Abstract: Background: Large language models are typically evaluated as models, benchmarks, or short conversational episodes. Less is known about what happens when an agent is embedded persistently in a real academic research environment wit…

  905. Hugging Face Daily Papers TIER_1 English(EN) ·

    迈向忠实的 Agentic XAI:一种用于提升模型忠实度的验证方法和开放世界基准

    Explainable AI (XAI) helps users interpret model behavior and identify potential faults. Agentic XAI systems use Large Language Models (LLMs) to make explanations more accessible through natural-language interaction, but they can also produce plausible yet unfaithful explanations…

  906. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Srini Ramaswamy ·

    智能作为受控自主:Agentic AI系统的失败、升级与治理

    As autonomous and agentic AI systems scale in robotic and human-machine environments, managing hallucination and persistent but unjustified action remains an open challenge. Rather than attributing these failures solely to model or alignment limitations, this paper explores the a…

  907. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agyn:一个开源的 AI Agent 平台,支持可扩展的按需执行、代码化 Agent 定义和零信任访问

    As organizations move toward production deployments of AI agents, which execute non-deterministic workflows, maintain stateful sessions, and often operate with privileged access to internal services, the engineering challenge shifts from building individual agents to operating th…

  908. arXiv cs.LG TIER_1 English(EN) · Ramin Khalili ·

    通过兼顾准确性和延迟的用户激励措施实现人工智能推理的绿色化

    The widespread use of AI services has raised concerns for its environmental sustainability, towards which recent studies have identified carbon emissions of AI inference as the major contributor. This paper introduces a framework for designing AI inference incentives based on the…

  909. Hugging Face Daily Papers TIER_1 English(EN) ·

    通过兼顾准确性和延迟的用户激励措施实现人工智能推理的绿色化

    The widespread use of AI services has raised concerns for its environmental sustainability, towards which recent studies have identified carbon emissions of AI inference as the major contributor. This paper introduces a framework for designing AI inference incentives based on the…

  910. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Anas H. Alzahrani ·

    学术研究中的持久性人工智能代理:单研究员实施案例研究

    Background: Large language models are typically evaluated as models, benchmarks, or short conversational episodes. Less is known about what happens when an agent is embedded persistently in a real academic research environment with durable memory, local files, external tools, sch…

  911. arXiv cs.AI TIER_1 English(EN) · Ting Liu ·

    合同技能:企业AI代理的GovernSpec设计框架

    arXiv:2605.22634v2 Announce Type: replace-cross Abstract: Skills have become a practical packaging mechanism for agent instructions, workflows, scripts, and reference materials. In enterprise settings, however, a skill often needs to express more than task guidance: goals, input …

  912. arXiv cs.AI TIER_1 English(EN) · Marcelo Fernandez - TraslaIA ·

    实现重构性权威:自主代理系统中的运行时构建、依赖解析与执行门控

    arXiv:2605.23935v1 Announce Type: new Abstract: Autonomous agent systems fail not only due to incorrect decisions, but due to executing decisions whose authority no longer holds at runtime. Prior work defined Reconstructive Authority (RAM) as a condition for valid execution: acti…

  913. arXiv cs.AI TIER_1 English(EN) · Alfredo Metere ·

    Agent技能的形式化验证方法:迈向机械可检验能力约束证明的三层架构

    arXiv:2605.23951v1 Announce Type: new Abstract: The companion paper introduced a four-level verification lattice on agent-skill manifests (unverified, declared, tested, formal) and left the top level aspirational. This paper closes that gap. We give a precise semantics for skill …

  914. arXiv cs.AI TIER_1 English(EN) · Bowen Wang, Dunjie Lu, Junli Wang, Tianyi Bai, Shixuan Liu, Zhipeng Zhang, Haiquan Wang, Hao Hu, Tianbao Xie, Shuai Bai, Dayiheng Liu, Que Shen, Junyang Lin, Tao Yu ·

    CUA-Gym:可验证训练环境和计算机使用代理任务的扩展

    arXiv:2605.25624v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven breakthroughs in domains such as math, tool-use, and software engineering, yet its extension to computer-use agents (CUAs) has been bottlenecked by the scarcity of sca…

  915. arXiv cs.AI TIER_1 English(EN) · Hao-Hsuan Chen ·

    为每一次行动投保:运行时精算控制自主人工智能代理的权威前沿框架

    arXiv:2605.25632v1 Announce Type: new Abstract: Autonomous AI agents increasingly issue side-effect-bearing actions: database mutations, refunds, payments, external commitments. We propose the Actuarial Action Interface (AAI), a deterministic runtime contract that prices each suc…

  916. arXiv cs.CL TIER_1 English(EN) · Vaishnavi Shrivastava, Piero Kauffmann, Ahmed Awadallah, Dimitris Papailiopoulos ·

    ECHO:终端代理免费学习世界模型

    arXiv:2605.24517v1 Announce Type: cross Abstract: CLI agents are the closest thing language models have to an embodied setting: the model emits commands, the terminal executes them, and the returned stream -- stdout, errors, files, logs, and traces -- records the consequences. We…

  917. arXiv cs.AI TIER_1 English(EN) · Haolang Zhao, Yunbo Long, Lukas Beckenbauer, Alexandra Brintrup ·

    VeriTrace:为深度研究代理演进心智模型

    arXiv:2605.26081v1 Announce Type: new Abstract: Deep research agents face vast, interdependent, and pervasively uncertain information. Existing systems explore what evolving intermediate representations should look like, but leave their evolution to the LLM's implicit reasoning. …

  918. arXiv cs.AI TIER_1 English(EN) · Shangding Gu ·

    从模型扩展到系统扩展:Agentic AI中的扩展工具

    arXiv:2605.26112v1 Announce Type: new Abstract: This paper studies the next major bottleneck in agentic AI as system scaling, not only model scaling: the design of auditable, persistent, modular, and verifiable architectures around foundation models. We refer to this shift as sca…

  919. arXiv cs.AI TIER_1 English(EN) · Wonjoong Kim, Sangwu Park, Yeonjun In, Sein Kim, Dongha Lee, Chanyoung Park ·

    超越最终答案:评估工具增强代理的推理轨迹

    arXiv:2510.02837v3 Announce Type: replace Abstract: Although recent tool-augmented benchmarks involve complex requests, evaluation remains limited to answer matching, neglecting critical trajectory aspects like efficiency, hallucination, and adaptivity. The most straightforward m…

  920. arXiv cs.AI TIER_1 English(EN) · Jia Huang, Joey Tianyi Zhou ·

    人工智能代理设计模式的二维框架:认知功能与执行拓扑

    arXiv:2605.13850v2 Announce Type: replace Abstract: Existing frameworks for LLM-based agent architectures describe systems from a single perspective: industry guides (Anthropic, Google, LangChain) focus on execution topology -- how data flows -- while cognitive science surveys fo…

  921. arXiv cs.AI TIER_1 Italiano(IT) · Yubo Li, Yidi Miao, Haotian Shen, Yuxin Liu ·

    PANDO:通过在线技能蒸馏实现高效多模态人工智能代理

    arXiv:2605.24785v1 Announce Type: new Abstract: Recent advances in multimodal web agents often rely on increased inference-time computation, including rollout search, verifier passes, offline skill discovery, and specialist model stacks. This raises a central question: can a web …

  922. arXiv cs.CL TIER_1 English(EN) · Junlin Wang, Federico Bianchi, Shang Zhu, Fan Nie, Yongchan Kwon, Bhuwan Dhingra, James Zou ·

    AI 代理和大型语言模型的自动化基准审计

    arXiv:2605.26079v1 Announce Type: new Abstract: Modern AI benchmarks operate at a complexity that outpaces traditional verification methods. Tasks authored by domain experts often contain implicit assumptions, incomplete environment specifications, and brittle evaluation logic th…

  923. arXiv cs.AI TIER_1 English(EN) · Liew Keong Han ·

    探索后解决:ARC-AGI-3 认知智能体的速度-深度权衡

    arXiv:2605.25931v1 Announce Type: new Abstract: We systematically investigate all 25 public ARC-AGI-3 games and find that every one is reachable through non-intelligent strategies: 10 in a single blind step, 5 after one probing action, 1 via repeated ACTION1 presses, 1 via divers…

  924. Hugging Face Daily Papers TIER_1 English(EN) ·

    SIA:具有约束和权重更新的自改进人工智能

    A self-improving AI framework simultaneously updates both model weights and task-specific agent architecture through a language-model feedback agent across legal classification, GPU optimization, and biological data denoising tasks.

  925. Hugging Face Daily Papers TIER_1 Italiano(IT) ·

    PANDO:通过在线技能蒸馏实现高效多模态人工智能代理

    PANDO is a web agent framework that improves efficiency through experience accumulation by reducing redundant actions, optimizing skill discovery, and enhancing prompt caching without sacrificing performance.

  926. arXiv cs.AI TIER_1 English(EN) · Shangding Gu ·

    从模型扩展到系统扩展:Agentic AI中的扩展约束

    This paper studies the next major bottleneck in agentic AI as system scaling, not only model scaling: the design of auditable, persistent, modular, and verifiable architectures around foundation models. We refer to this shift as scaling the harness: treating the structured execut…

  927. arXiv cs.AI TIER_1 English(EN) · Alexandra Brintrup ·

    VeriTrace:为深度研究代理演进心智模型

    Deep research agents face vast, interdependent, and pervasively uncertain information. Existing systems explore what evolving intermediate representations should look like, but leave their evolution to the LLM's implicit reasoning. Without explicit regulation, the intermediate la…

  928. arXiv cs.CL TIER_1 English(EN) · James Zou ·

    AI 代理和大型语言模型的自动化基准审计

    Modern AI benchmarks operate at a complexity that outpaces traditional verification methods. Tasks authored by domain experts often contain implicit assumptions, incomplete environment specifications, and brittle evaluation logic that human annotation cannot reliably catch. We in…

  929. arXiv cs.AI TIER_1 English(EN) · Liew Keong Han ·

    探索后解决:ARC-AGI-3 认知智能体的速度-深度权衡

    We systematically investigate all 25 public ARC-AGI-3 games and find that every one is reachable through non-intelligent strategies: 10 in a single blind step, 5 after one probing action, 1 via repeated ACTION1 presses, 1 via diverse exploration, and 8 via single repeated actions…

  930. Hugging Face Daily Papers TIER_1 English(EN) ·

    探索后解决:ARC-AGI-3 认知智能体的速度-深度权衡

    We systematically investigate all 25 public ARC-AGI-3 games and find that every one is reachable through non-intelligent strategies: 10 in a single blind step, 5 after one probing action, 1 via repeated ACTION1 presses, 1 via diverse exploration, and 8 via single repeated actions…

  931. arXiv cs.AI TIER_1 English(EN) · Federico Bottino, Carlo Ferrero, Nicholas Dosio, Pierfrancesco Beneventano ·

    检索是不够的:为什么组织AI需要认知基础设施

    arXiv:2604.11759v2 Announce Type: replace Abstract: Organizational knowledge used by AI agents typically lacks epistemic structure: retrieval systems surface semantically relevant content without distinguishing binding decisions from abandoned hypotheses, contested claims from se…

  932. arXiv cs.AI TIER_1 English(EN) · Lixiang Yan, Dragan Ga\v{s}evi\'c ·

    Agentivism:人工智能时代的学习理论

    arXiv:2604.07813v2 Announce Type: replace Abstract: Learning theories have historically changed when the conditions of learning evolved. Generative and agentic AI create a new condition by allowing learners to delegate explanation, writing, problem solving, and other cognitive wo…

  933. arXiv cs.AI TIER_1 English(EN) · Chitra Badagi, Divye Singh, Animesh Sen, Adinath Shirsath ·

    AI保障:企业级AI系统的全面测试策略

    arXiv:2605.23459v1 Announce Type: cross Abstract: Enterprise AI systems, built on large language models, retrieval pipelines and autonomous agents, introduce a class of risks that traditional software quality assurance was never designed to address. These systems are probabilisti…

  934. arXiv cs.AI TIER_1 English(EN) · Joshua Odmark, Gideon Rubin, Deon van der Vyver ·

    用于代理式 Kubernetes 操作的测量基底:方法论及检索复合伪造的案例研究

    arXiv:2605.23058v1 Announce Type: cross Abstract: Empirical claims about autonomous Kubernetes operations agents are largely unfalsifiable. Published work reports observational results without controlled comparisons against an agent-disabled baseline, selection bias is endemic, p…

  935. arXiv cs.AI TIER_1 Dansk(DA) · Yifan Yang, Ziyang Gong, Weiquan Huang, Qihao Yang, Ziwei Zhou, Zisu Huang, Yan Li, Xuemei Gao, Qi Dai, Bei Liu, Kai Qiu, Yuqing Yang, Dongdong Chen, Xue Yang, Chong Luo ·

    SkillOpt:自主进化智能体技能的执行策略

    arXiv:2605.23904v1 Announce Type: new Abstract: Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning optimizer for the skill, and none of which reliably improves over its starting …

  936. arXiv cs.AI TIER_1 English(EN) · Zehao Wang, Shilong Jin, Zhao Cao, Lanjun Wang ·

    当计划在正确执行的情况下仍然失败:关于基于LLM的多智能体系统的认知校准

    arXiv:2605.23414v1 Announce Type: new Abstract: LLM-based multi-agent systems can fail even when planned actions are executed correctly because agents may misjudge their knowledge when evaluating plan feasibility, a phenomenon we term epistemic miscalibration in planning. Unlike …

  937. arXiv cs.AI TIER_1 English(EN) · Muhammad Zia Hydari, Farooq Muzaffar ·

    重绘人工智能版图:代理生态系统中的问责边界理论

    arXiv:2605.23179v1 Announce Type: new Abstract: Agentic AI orchestrators reduce the interface and assembly costs of composing information systems capabilities across organizational boundaries, seemingly accelerating modularization and organizational disaggregation. Yet AI-enabled…

  938. arXiv cs.AI TIER_1 English(EN) · Yamato Arai, Yuma Ichikawa ·

    EVE-Agent:可验证证据的自进化智能体

    arXiv:2605.22905v1 Announce Type: new Abstract: Self-evolving agents should not train on examples they cannot justify. Data-free self-evolving search agents offer a scalable route to systems that generate their own questions, answer them, and improve from their own feedback witho…

  939. arXiv cs.AI TIER_1 English(EN) · Deepak Panigrahy, Aakash Tyagi ·

    每次成功进球的能量:面向Agentic AI系统的进球级能量核算

    arXiv:2605.22883v1 Announce Type: new Abstract: Current AI energy benchmarks measure consumption at the granularity of a single model invocation or training run. For classical single-turn workloads this unit remains coherent. For agentic systems - where a single user goal may tri…

  940. arXiv cs.AI TIER_1 English(EN) · Dongxin Guo ·

    确定性地平线:不可行性结果作为可信赖人工智能系统的设计规范

    arXiv:2605.23024v1 Announce Type: new Abstract: Large language models now write software, draft legal documents, and produce clinical notes, yet fundamental limits, from Turing and Arrow to the No Free Lunch theorems, shape what computation can do. This thesis turns such impossib…

  941. Hugging Face Daily Papers TIER_1 English(EN) ·

    CUA-Gym:为计算机使用代理扩展可验证的训练环境和任务

    RLVR framework for computer-use agents addresses data scarcity through scalable generation pipeline and synthetic environments, achieving superior performance on verification and transfer benchmarks.

  942. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Michiel Bakker ·

    Habermolt:将审议委托给人工智能代表

    Deliberative democracy arguably leads to better collective decisions, but is fundamentally constrained by human attention and bandwidth. While recent AI-mediated deliberations scale participation by synthesizing inputs from many humans, they remain time-intensive for individual u…

  943. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Lewis Hammond ·

    Habermolt:将审议委托给人工智能代表

    Deliberative democracy arguably leads to better collective decisions, but is fundamentally constrained by human attention and bandwidth. While recent AI-mediated deliberations scale participation by synthesizing inputs from many humans, they remain time-intensive for individual u…

  944. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Michiel Bakker ·

    Habermolt:将审议委托给人工智能代表

    Deliberative democracy arguably leads to better collective decisions, but is fundamentally constrained by human attention and bandwidth. While recent AI-mediated deliberations scale participation by synthesizing inputs from many humans, they remain time-intensive for individual u…

  945. Hugging Face Daily Papers TIER_1 English(EN) ·

    ECHO:终端代理免费学习世界模型

    Environment cross-entropy hybrid objective combines policy-gradient loss with auxiliary environment observation prediction to provide dense supervision from terminal feedback, improving agent performance and self-improvement capabilities.

  946. Hugging Face Daily Papers TIER_1 English(EN) ·

    物理AI中的静默故障:运行时动作授权在自主系统中的文献综述

    Physical AI systems face safety challenges where black-box models can execute harmful actions without detection, necessitating comprehensive runtime guardrail mechanisms for safe operation.

  947. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Fouad Bousetouane ·

    ProofAgent Harness: AI Agent对抗性评估的开放基础设施

    AI agents are entering high-risk production settings, where they use tools, retain context, follow policies, handle private data, and interact with users over multiple turns. Yet many evaluation methods still judge isolated outputs or static tasks, missing failures that emerge th…

  948. arXiv cs.AI TIER_1 Dansk(DA) · Chong Luo ·

    SkillOpt:自主进化智能体技能的执行策略

    Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning optimizer for the skill, and none of which reliably improves over its starting point under feedback. We argue the skill should …

  949. arXiv cs.AI TIER_1 English(EN) · Adinath Shirsath ·

    AI保障:企业级AI系统的全面测试策略

    Enterprise AI systems, built on large language models, retrieval pipelines and autonomous agents, introduce a class of risks that traditional software quality assurance was never designed to address. These systems are probabilistic, context-sensitive and emergent: they cannot be …

  950. arXiv cs.AI TIER_1 English(EN) · Lanjun Wang ·

    当计划执行正确但仍失败时:关于基于LLM的多智能体系统的认知校准

    LLM-based multi-agent systems can fail even when planned actions are executed correctly because agents may misjudge their knowledge when evaluating plan feasibility, a phenomenon we term epistemic miscalibration in planning. Unlike execution errors, epistemic miscalibration is la…

  951. arXiv cs.AI TIER_1 English(EN) · Parsa Mazaheri, Kasra Mazaheri ·

    AgentAtlas:超越LLM代理结果排行榜

    arXiv:2605.20530v1 Announce Type: new Abstract: Large language model agents now act on codebases, browsers, operating systems, calendars, files, and tool ecosystems, but the benchmarks used to evaluate them are fragmented: each emphasizes a different unit of measurement (final ta…

  952. arXiv cs.CL TIER_1 English(EN) · Qisheng Su, Zhen Fang, Shiting Huang, Yu Zeng, Yiming Zhao, Kou Shi, Ziao Zhang, Lin Chen, Zehui Chen, Lijun Wu, Feng Zhao ·

    ACC:为长上下文训练编译代理轨迹

    arXiv:2605.21850v1 Announce Type: new Abstract: Recent development of agents has renewed demand for long-context reasoning capacity of LLMs. However, training LLMs for this capacity requires costly long-document curation or heuristic context synthesis. We observe that agents prod…

  953. arXiv cs.CL TIER_1 English(EN) · Baolin Peng, Wenlin Yao, Qianhui Wu, Hao Cheng, Xiao Yu, Rui Yang, Tao Ge, Alessandro Sordoni, Xingdi Yuan, Yelong Shen, Pengcheng He, Tong Zhang, Zhou Yu, Jianfeng Gao ·

    Orchard: 一个开源的代理建模框架

    arXiv:2605.15040v2 Announce Type: replace-cross Abstract: Agentic modeling aims to transform LLMs into autonomous agents capable of solving complex tasks through planning, reasoning, tool use, and multi-turn interaction with environments. Despite major investment, open research r…

  954. arXiv cs.LG TIER_1 English(EN) · Fiona Y. Wong, Markus J. Buehler ·

    跨领域基准测试揭示了协调式AI代理何时能改进基于部分证据的科学推理

    arXiv:2605.22300v1 Announce Type: cross Abstract: Scientific evidence often spans instruments, databases, and disciplines, so no single source records the full phenomenon. This makes it difficult to determine when coordinated AI agents add value over simpler scientific workflows.…

  955. arXiv cs.AI TIER_1 English(EN) · Lucas Jing, Xinqi Wang, Liao Zhang, Simon S. Du ·

    PBT-Bench:基于属性的测试中对 AI 代理进行基准测试

    arXiv:2605.15229v2 Announce Type: replace-cross Abstract: Existing code benchmarks measure whether an agent can produce any test that reproduces a known bug, or whether it can produce a patch that fixes a described issue. Neither isolates the distinct skill of property-based test…

  956. arXiv cs.AI TIER_1 English(EN) · Lujain Ibrahim, Katherine M. Collins, Sunnie S. Y. Kim, Anka Reuel, Max Lamparth, Kevin Feng, Lama Ahmad, Prajna Soni, Alia El Kattan, Merlin Stein, Siddharth Swaroop, Vishakh Padmakumar, Ilia Sucholutsky, Andrew Strait, Diyi Yang, Q. Vera Liao, Umang Bh… ·

    衡量和减轻过度依赖以构建人类兼容的AI

    arXiv:2509.08010v2 Announce Type: replace-cross Abstract: Large language models (LLMs) distinguish themselves from previous technologies by functioning as collaborative ``thought partners,'' capable of engaging more fluidly in natural language on a range of tasks. As LLMs increas…

  957. arXiv cs.LG TIER_1 English(EN) · Simon Dennis, Rivaan Patil, Kevin Shabahang, Hao Guo ·

    将 Agentic Workflows 编译到 LLM 权重中:以低两个数量级的成本实现近乎前沿的质量

    arXiv:2605.22502v1 Announce Type: cross Abstract: Agent orchestration frameworks have proliferated, collectively exceeding 290,000 GitHub stars across LangGraph, CrewAI, Google ADK, OpenAI Agents SDK, Semantic Kernel, Strands, and LlamaIndex. All follow the same pattern: an exter…

  958. arXiv cs.LG TIER_1 English(EN) · Qianshu Cai, Yonggang Zhang, Xianzhang Jia, Wei Xue, Jun Song, Xinmei Tian, Yike Guo ·

    MOSS:在自主代理系统中通过源代码重写实现自我进化

    arXiv:2605.22794v1 Announce Type: cross Abstract: Autonomous agentic systems are largely static after deployment: they do not learn from user interactions, and recurring failures persist until the next human-driven update ships a fix. Self-evolving agents have emerged in response…

  959. arXiv cs.AI TIER_1 English(EN) · Yuanyang Li, Xue Yang, Longyue Wang, Weihua Luo, Hongyang Chen ·

    ComplexMCP:动态、相互依赖且大规模工具沙箱中 LLM Agent 的评估

    arXiv:2605.10787v2 Announce Type: replace Abstract: Current LLM agents are proficient at calling isolated APIs but struggle with the "last mile" of commercial software automation. In real-world scenarios, tools are not independent; they are atomic, interdependent, and prone to en…

  960. arXiv cs.AI TIER_1 English(EN) · Zhengkang Guo, Yiyang Li, Lin Qiu, Xiaohua Wang, Jingwen Xv, Dongyu Ru, Xiaoyu Li, Xiaoqing Zheng, Xuezhi Cao, Xunliang Cai ·

    AgentEscapeBench:评估LLM智能体在域外工具基础上的推理能力

    arXiv:2605.07926v2 Announce Type: replace Abstract: As LLM-based agents increasingly rely on external tools, it is important to evaluate their ability to sustain tool-grounded reasoning beyond familiar workflows and short-range interactions. We introduce AgentEscapeBench, an esca…

  961. arXiv cs.AI TIER_1 English(EN) · Nelly Dux, Cristina Alaimo, Philippe Roussiere, Abhishek Kumar Mishra ·

    设计驱动的治理:构建代理式AI以实现组织学习和可扩展自主性

    arXiv:2605.20210v1 Announce Type: cross Abstract: Agentic AI systems - systems that can pursue goals through multi-step planning and tool-mediated action with limited direct supervision - are moving from experimental prototypes to enterprise deployments. This transition introduce…

  962. arXiv cs.AI TIER_1 English(EN) · Ming Zhu, Juntao Tan, Rithesh Murthy, Jielin Qiu, Liangwei Yang, Wenting Zhao, Silvio Savarese, Shelby Heinecke, Huan Wang ·

    RealUserSim:通过基于现实的用户模拟弥合代理基准测试中的现实差距

    arXiv:2605.20204v1 Announce Type: cross Abstract: LLM-based user simulation is the primary mechanism for end-to-end agent evaluation, yet simulated users are poor proxies for real humans: unconstrained LLM defaults produce a Formalism Ceiling (style match rates of 6-8% against re…

  963. arXiv cs.AI TIER_1 English(EN) · Liyuan Deng, Shujian Deng, Yongkang Chen, Yongkang Dai, Zhihang Zhong, Linyang Li, Xiao Sun, Yilei Shi, Huaxi Huang ·

    用于闭环优化、仿真和建模编排的工具增强型代理

    arXiv:2605.20190v1 Announce Type: new Abstract: Iterative industrial design-simulation optimization is bottlenecked by the CAD-CAE semantic gap: translating simulation feedback into valid geometric edits under diverse, coupled constraints. To fill this gap, we propose COSMO-Agent…

  964. arXiv cs.AI TIER_1 English(EN) · Binghan Wu, Shoufeng Wang, Yunxin Liu, Ya-Qin Zhang, Joseph Sifakis, Ye Ouyang ·

    从自动化到自主化:分层原生智能体网络架构 (HANA)

    arXiv:2605.20608v1 Announce Type: new Abstract: Realizing Level 4/5 Autonomous Networks (AN) demands a shift from static automation to agent-native intelligence. Current operations, reliant on rigid scripts, lack the cognitive agency to handle off-nominal conditions. To address t…

  965. arXiv cs.AI TIER_1 English(EN) · Jiefeng Chen, Bhavana Dalvi Mishra, Jaehyun Nam, Rui Meng, Tomas Pfister, Jinsung Yoon ·

    MARS:具有反思性搜索的模块化代理,用于自动化人工智能研究

    arXiv:2602.02660v3 Announce Type: replace Abstract: A critical bottleneck in automating AI research is the execution of complex machine learning engineering (MLE) tasks. MLE differs from general software engineering due to computationally expensive evaluation (e.g., model trainin…

  966. arXiv cs.AI TIER_1 English(EN) · Yoon Pyo Lee, Samrendra Roy, Jay Yoo, Kazuma Kobayashi, Sajedul Talukder, Seid Koric, Souvik Chakraborty, Syed Bahauddin Alam ·

    面向核反应堆控制的领域特定基础模型的代理物理人工智能

    arXiv:2512.23292v3 Announce Type: replace Abstract: The prevailing paradigm in AI for physical systems (scaling general-purpose foundation models toward universal multimodal reasoning) confronts a fundamental barrier at the control interface. Recent benchmarks show that even fron…

  967. arXiv cs.CL TIER_1 English(EN) · Asaf Yehudai, Lilach Eden, Michal Shmueli-Scheuer ·

    Agentic CLEAR:自动化多层级LLM Agent评估

    arXiv:2605.22608v1 Announce Type: new Abstract: Agentic systems are becoming more capable: agents define strategies, take actions, and interact with different environments. This autonomy poses serious challenges for overseeing and assessing agent behavior. Most current tools are …

  968. arXiv cs.AI TIER_1 English(EN) · Yibo Li, Jiashuo Yang, Zhi Zheng, Zhiyuan Hu, Yuan Sui, Shizun Wang, Yufei He, Bryan Hooi ·

    APEX:用于自进化LLM智能体的自主策略探索

    arXiv:2605.21240v1 Announce Type: cross Abstract: LLM agents have shown strong performance across a wide range of complex tasks, including interactive environments that require long-horizon decision making. But these agents cannot learn on the fly at test time. Self-evolving agen…

  969. arXiv cs.CL TIER_1 English(EN) · Jinhu Qi, Yifan Li, Minghao Zhao, Wentao Zhang, Zijian Zhang, Yaoman Li, Irwin King ·

    超越基准岛屿:迈向具有代表性的Agentic AI可信度评估

    arXiv:2603.14987v2 Announce Type: replace Abstract: Agentic AI systems increasingly act through tool-augmented, multi-step workflows whose failures (unsafe tool use, unauthorised actions, social harm) carry deployment-level consequences. Evaluation practice remains fragmented acr…

  970. arXiv cs.AI TIER_1 English(EN) · Christopher Koch ·

    Agentic Agile-V:从 Vibe Coding 到软件和硬件开发中的已验证工程

    arXiv:2605.20456v1 Announce Type: cross Abstract: Agentic AI coding systems can inspect repositories, plan implementation steps, edit files, call tools, run tests, and submit pull requests. These capabilities make software and hardware development faster in some settings, but cur…

  971. arXiv cs.CL TIER_1 English(EN) · Mingkai Deng, Jinyu Hou, Lara S\'a Neves, Varad Pimpalkhute, Taylor W. Killian, Zhengzhong Liu, Eric P. Xing ·

    通过自调节模拟规划实现高效的代理推理

    arXiv:2605.22138v1 Announce Type: cross Abstract: How should an agent decide when and how to plan? A dominant approach builds agents as reactive policies with adaptive computation (e.g., chain-of-thought), trained end-to-end expecting planning to emerge implicitly. Without contro…

  972. arXiv cs.AI TIER_1 English(EN) · Aditya Taparia, Som Sagar, Ransalu Senanayake ·

    学习配置Agentic AI系统

    arXiv:2602.11574v3 Announce Type: replace Abstract: Configuring LLM-based agent systems involves choosing workflows, tools, token budgets, and prompts from a large combinatorial design space, and is typically handled today by fixed templates or hand-tuned heuristics that apply th…

  973. arXiv cs.AI TIER_1 English(EN) · Zihao Cheng, Hongru Wang, Zeming Liu, Xinyi Wang, Xiangrong Zhu, Yuhang Guo, Wei Lin, Jeff Z. Pan, Yunhong Wang ·

    Terminal-World: 通过 Agent Skills 扩展 Terminal-Agent 环境

    arXiv:2605.20876v1 Announce Type: cross Abstract: Terminal agents extend Large Language Models with the ability to execute tasks directly in command-line environments, but their progress is bottlenecked by the scarcity of high-quality training data. Existing approaches bootstrap …

  974. Hugging Face Daily Papers TIER_1 Dansk(DA) ·

    SkillOpt:自主进化智能体技能的执行策略

    SkillOpt introduces a systematic text-space optimizer for agent skills that trains skills as external agent state with stable updates and zero deployment inference overhead, achieving superior performance across multiple benchmarks and execution environments.

  975. arXiv cs.CL TIER_1 English(EN) · Dongxin Guo ·

    确定性视界:不可行性结果作为可信赖人工智能系统的设计规范

    Large language models now write software, draft legal documents, and produce clinical notes, yet fundamental limits, from Turing and Arrow to the No Free Lunch theorems, shape what computation can do. This thesis turns such impossibility results from curiosities into design rules…

  976. arXiv cs.AI TIER_1 English(EN) · Yike Guo ·

    MOSS:在自主代理系统中通过源代码重写实现自我进化

    Autonomous agentic systems are largely static after deployment: they do not learn from user interactions, and recurring failures persist until the next human-driven update ships a fix. Self-evolving agents have emerged in response, but all confine evolution to text-mutable artifa…

  977. arXiv cs.CL TIER_1 English(EN) · Yuma Ichikawa ·

    EVE-Agent:可验证证据的自进化代理

    Self-evolving agents should not train on examples they cannot justify. Data-free self-evolving search agents offer a scalable route to systems that generate their own questions, answer them, and improve from their own feedback without human annotations. Yet, without verifiable ev…

  978. arXiv cs.AI TIER_1 English(EN) · Haibo Chen ·

    DeltaBox:以毫秒级沙盒检查点/回滚实现状态化AI代理的可扩展性

    LLM-powered AI agents require high-frequency state exploration (e.g., test-time tree search and reinforcement learning), relying on rapid checkpoint and rollback (C/R) of the complete sandbox state, including files and process state (e.g., memory, contexts, etc.). Existing mechan…

  979. arXiv cs.AI TIER_1 English(EN) · Andrii Kryshtal ·

    人工智能会加剧冲突吗?大型语言模型在冲突背景下部署的对齐失败问题

    AI models are already deployed in societies affected by armed conflict, and journalists, humanitarian workers, governments and ordinary citizens rely on them for information or for their work processes. No established practice exists for checking whether their outputs can make th…

  980. arXiv cs.AI TIER_1 English(EN) · Fayao Liu ·

    Claw AI Lab:一个自主多智能体研究团队

    We present Claw AI Lab, a lab-native autonomous research platform that advances automated research from a hidden prompt-to-paper pipeline into an interactive AI laboratory. Rather than centering the system around a single agent or a fixed serial workflow, we allow users to instan…

  981. arXiv cs.AI TIER_1 English(EN) · Ting Liu ·

    合同技能:企业AI代理的GovernSpec设计框架

    Skills are increasingly used to package agent instructions, workflows, scripts, and reference materials. In enterprise settings, however, skills often need to express more than task guidance: they must make goals, input boundaries, permissions, evidence requirements, output contr…

  982. arXiv cs.AI TIER_1 English(EN) · Michal Shmueli-Scheuer ·

    Agentic CLEAR:自动化多层级LLM代理评估

    Agentic systems are becoming more capable: agents define strategies, take actions, and interact with different environments. This autonomy poses serious challenges for overseeing and assessing agent behavior. Most current tools are limited, focusing on observability with basic ev…

  983. arXiv cs.AI TIER_1 English(EN) · He Ye ·

    TerminalWorld:在真实世界终端任务上对代理进行基准测试

    We introduce TerminalWorld, a scalable data engine that automatically reverse-engineers high-fidelity evaluation tasks from "in-the-wild" terminal recordings. Processing 80,870 terminal recordings, the engine yields a full benchmark of 1,530 validated tasks, spanning 18 real-worl…

  984. arXiv cs.AI TIER_1 English(EN) · Hao Guo ·

    将 Agentic Workflows 编译到 LLM 权重中:以低两个数量级的成本实现近乎前沿的质量

    Agent orchestration frameworks have proliferated, collectively exceeding 290,000 GitHub stars across LangGraph, CrewAI, Google ADK, OpenAI Agents SDK, Semantic Kernel, Strands, and LlamaIndex. All follow the same pattern: an external orchestrator above the LLM, injecting instruct…

  985. Don't Worry About the Vase (Zvi Mowshowitz) TIER_1 English(EN) · Zvi Mowshowitz ·

    AI #169: 新知识

    Even in a relatively quiet period, AI is out there creating new knowledge.

  986. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Markus J. Buehler ·

    跨领域基准测试揭示了协调式AI代理何时能改进基于部分证据的科学推理

    Scientific evidence often spans instruments, databases, and disciplines, so no single source records the full phenomenon. This makes it difficult to determine when coordinated AI agents add value over simpler scientific workflows. We evaluate this question with a cross-domain ben…

  987. arXiv cs.CL TIER_1 English(EN) · Eric P. Xing ·

    通过自调节模拟规划实现高效的代理推理

    How should an agent decide when and how to plan? A dominant approach builds agents as reactive policies with adaptive computation (e.g., chain-of-thought), trained end-to-end expecting planning to emerge implicitly. Without control over the presence, structure, or horizon of plan…

  988. 量子位 (QbitAI) TIER_1 中文(ZH) · 思邈 ·

    上海交通大学AI教授授课:半天拆解Agent底层逻辑

    周日来北京线下揭秘

  989. arXiv cs.CL TIER_1 English(EN) · Feng Zhao ·

    ACC:为长上下文训练编译代理轨迹

    Recent development of agents has renewed demand for long-context reasoning capacity of LLMs. However, training LLMs for this capacity requires costly long-document curation or heuristic context synthesis. We observe that agents produce massive trajectories when solving problems, …

  990. Hugging Face Daily Papers TIER_1 English(EN) ·

    通过自调节模拟规划实现高效的代理推理

    Efficient agentic reasoning requires decomposing decision-making into three systems—simulative reasoning, self-regulation, and reactive execution—enabling controlled planning that reduces token usage while maintaining performance.

  991. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Nathaniel Pinckney ·

    Trace2Skill:面向长上下文EDA代理的验证器引导技能演化

    Complex Verilog Design Problems (CVDP) challenge hardware LLM agents because solving them requires localizing verifier-relevant RTL, testbenches, include paths, and build dependencies inside large repository snapshots, making precise edits, and recovering from sparse hidden-verif…

  992. Latent Space (swyx) TIER_1 English(EN) ·

    铁路:原生智能体云 — Jake Cooper

    3M Users, 100K Signups/Week, Own-Metal Data Centers, $200K+ Coding Agent Spend, and the Death of PRs

  993. arXiv cs.AI TIER_1 English(EN) · Bryan Hooi ·

    APEX:用于自进化LLM代理的自主策略探索

    LLM agents have shown strong performance across a wide range of complex tasks, including interactive environments that require long-horizon decision making. But these agents cannot learn on the fly at test time. Self-evolving agents address this by accumulating memory and reflect…

  994. arXiv cs.AI TIER_1 English(EN) · Yunhong Wang ·

    Terminal-World:通过 Agent Skills 扩展 Terminal-Agent 环境

    Terminal agents extend Large Language Models with the ability to execute tasks directly in command-line environments, but their progress is bottlenecked by the scarcity of high-quality training data. Existing approaches bootstrap from partial sources such as human-defined seeds o…

  995. Hugging Face Daily Papers TIER_1 English(EN) ·

    从自动化到自主化:分层原生智能体网络架构 (HANA)

    Realizing Level 4/5 Autonomous Networks (AN) demands a shift from static automation to agent-native intelligence. Current operations, reliant on rigid scripts, lack the cognitive agency to handle off-nominal conditions. To address this, this letter proposes a hierarchical multi-a…

  996. arXiv cs.AI TIER_1 English(EN) · Ye Ouyang ·

    从自动化到自主化:分层原生智能体网络架构(HANA)

    Realizing Level 4/5 Autonomous Networks (AN) demands a shift from static automation to agent-native intelligence. Current operations, reliant on rigid scripts, lack the cognitive agency to handle off-nominal conditions. To address this, this letter proposes a hierarchical multi-a…

  997. arXiv cs.CL TIER_1 English(EN) · Kasra Mazaheri ·

    AgentAtlas:超越LLM代理结果排行榜

    Large language model agents now act on codebases, browsers, operating systems, calendars, files, and tool ecosystems, but the benchmarks used to evaluate them are fragmented: each emphasizes a different unit of measurement (final task success, tool-call validity, repeated-pass co…

  998. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Christopher Koch ·

    Agentic Agile-V:从 Vibe Coding 到软件和硬件开发中的已验证工程

    Agentic AI coding systems can inspect repositories, plan implementation steps, edit files, call tools, run tests, and submit pull requests. These capabilities make software and hardware development faster in some settings, but current evidence does not support the simple claim th…

  999. arXiv cs.AI TIER_1 English(EN) · Vasundra Srinivasan ·

    一种用于生产环境中 LLM Agent 的运行时架构模式选择与组合方法

    Production LLM agents combine stochastic model outputs with deterministic software systems, yet the boundary between the two is rarely treated as a first-class architectural object. This paper names that boundary the stochastic-deterministic boundary (SDB): a four-part contract a…

  1000. arXiv cs.AI TIER_1 English(EN) · Yi Ling Yu ·

    面向连续AI代理评估的无分布不确定性量化

    We adapt split conformal prediction and adaptive conformal inference (ACI) to continuous AI agent evaluation, providing distribution-free coverage guarantees for forecasted quality scores. Conformal intervals achieve calibration error below 0.02 across all nominal levels at the 2…

  1001. arXiv cs.AI TIER_1 English(EN) · Arman Cohan ·

    OpenComputer: 计算机使用代理的可验证软件世界

    We present OpenComputer, a verifier-grounded framework for constructing verifiable software worlds for computer-use agents. OpenComputer integrates four components: (1) app-specific state verifiers that expose structured inspection endpoints over real applications, (2) a self-evo…

  1002. arXiv cs.AI TIER_1 English(EN) · Mark Fuge ·

    EngiAI:一个用于LLM驱动的工程设计的多个智能体框架和基准套件

    Large Language Model (LLM) agents are increasingly applied to engineering design tasks, yet existing evaluation frameworks do not adequately address multi-agent systems that combine simulation, retrieval, and manufacturing preparation. We introduce a benchmark suite with three ev…

  1003. Hugging Face Daily Papers TIER_1 English(EN) ·

    EnvFactory:通过可执行环境合成和鲁棒强化学习扩展工具使用代理

    Equipping LLMs with tool-use capabilities via Agentic Reinforcement Learning (Agentic RL) is bottlenecked by two challenges: the lack of scalable, robust execution environments and the scarcity of realistic training data that captures implicit human reasoning. Existing approaches…

  1004. arXiv cs.AI TIER_1 English(EN) · Sen Hu ·

    SkillGenBench:为LLM代理的技能生成管道进行基准测试

    As LLM agents are increasingly built around reusable skills, a central challenge is no longer only whether agents can use provided skills, but whether they can generate correct, reusable, and executable skills from repositories and documents. Existing benchmarks primarily evaluat…

  1005. arXiv cs.AI TIER_1 English(EN) · Ronaldo Martins da Costa ·

    Reversa:一种反向文档工程框架,用于将遗留软件转换为AI代理的操作规范

    Legacy systems concentrate business rules, architectural decisions, and operational exceptions that often remain implicit in code, data, configuration, and maintenance practices. At the same time, language-model-based coding agents depend on reliable context, correctness criteria…

  1006. arXiv cs.AI TIER_1 English(EN) · Wei Tsang Ooi ·

    AI for Auto-Research: Roadmap & User Guide

    AI-assisted research is crossing a threshold: fully automated systems can now generate research papers for as little as $15, while long-horizon agents can execute experiments, draft manuscripts, and simulate critique with minimal human input. Yet this productivity frontier expose…

  1007. arXiv cs.LG TIER_1 English(EN) · Nicholas D. Lane ·

    超越规模化:智能体正走向边缘

    The bottleneck of useful agentic intelligence has shifted from compressing world knowledge into a single model to executing a coordinated system. This position paper argues that personal-agent architecture must move to the edge because the core properties of agentic intelligence …

  1008. arXiv cs.AI TIER_1 English(EN) · Zhiyu Li ·

    SkillsVote:从收集、推荐到演进的智能体技能生命周期治理

    Long-horizon LLM agents leave traces that could become reusable experience, but raw trajectories are noisy and hard to govern. We treat Agent Skills as an experience schema that couples executable scripts, with non-executable guidance on procedures. Yet open skill ecosystems cont…

  1009. arXiv cs.CL TIER_1 English(EN) · Yuyu Luo ·

    可扩展环境驱动通用智能体

    Generalizable agents should adapt to diverse tasks and unseen environments beyond their training distribution. This position paper argues that such generalization requires environment scaling: expanding the distribution of executable rule-sets that agents interact with, rather th…

  1010. Hugging Face Daily Papers TIER_1 English(EN) ·

    PPAI:赋能个性化大模型代理互操作性,实现协作边缘智能

    Deploying large language model (LLM) on edge device enables personalized LLM agents for various users. The growing availability of diverse personalized agents presents a unique opportunity for peer-to-peer (P2P) collaboration, wherein each user can delegate tasks beyond the local…

  1011. arXiv cs.CL TIER_1 English(EN) · Song Guo ·

    PPAI:赋能个性化大模型代理互操作性,实现协作边缘智能

    Deploying large language model (LLM) on edge device enables personalized LLM agents for various users. The growing availability of diverse personalized agents presents a unique opportunity for peer-to-peer (P2P) collaboration, wherein each user can delegate tasks beyond the local…

  1012. arXiv cs.CL TIER_1 English(EN) · Kei Tateno ·

    PROTEA:多智能体LLM工作流的离线评估与迭代优化

    Multi-agent LLM workflows -- systems composed of multiple role-specific LLM calls -- often outperform single-prompt baselines, but they remain difficult to debug and refine. Failures can originate from subtle errors in intermediate outputs that propagate to downstream nodes, requ…

  1013. arXiv cs.CL TIER_1 English(EN) · Luning Sun ·

    多智能体AI系统在创造力方面超越人类团队

    Although artificial intelligence (AI) now matches or exceeds human performance across numerous cognitive tasks, creativity remains a highly contested frontier. As AI systems based on large language models (LLMs) are increasingly adopted in research and innovation, it is essential…

  1014. Hugging Face Daily Papers TIER_1 English(EN) ·

    EXG:具有经验图谱的自进化代理

    Large language model (LLM)-based agents have demonstrated strong capabilities in complex reasoning and problem solving through multi-step interactions, yet most deployed agents remain behaviorally static, with knowledge acquired during execution rarely translating into systematic…

  1015. arXiv cs.MA (Multiagent) TIER_1 (CA) · Xiaowei Huang ·

    负责任的代理式人工智能需要明确的溯源

    Agentic AI is rapidly proliferating across diverse real-world domains such as software engineering, yet public trust has not kept pace. The central reason is that responsibility, despite being widely discussed, remains a subjective and unenforced concept, as no current agentic fr…

  1016. arXiv cs.LG TIER_1 English(EN) · Sheila A. McIlraith ·

    形式化方法遇上大型语言模型:用于合规性高级人工智能系统的审计、监控和干预

    We examine one particular dimension of AI governance: how to monitor and audit AI-enabled products and services throughout the AI development lifecycle, from pre-deployment testing to post-deployment auditing. Combining principles from formal methods with SoTA machine learning, w…

  1017. arXiv cs.CL TIER_1 English(EN) · Fuli Feng ·

    三思而后行:LLM智能体的自主探索

    Large language model based agents often fail in unfamiliar environments due to premature exploitation: a tendency to act on prior knowledge before acquiring sufficient environment-specific information. We identify autonomous exploration as a critical yet underexplored capability …

  1018. arXiv cs.LG TIER_1 English(EN) · Gunnar König ·

    可解释AI已不足够!重新思考算法可抗辩性

    Machine learning systems increasingly make life-changing decisions about individuals, such as loan approvals, hiring, and cheating detection, raising a pressing question: how can individuals respond to negative decisions made by these opaque systems? While explainable artificial …

  1019. arXiv cs.AI TIER_1 English(EN) · Yisroel Mirsky ·

    谁拥有这个AI代理?追溯AI代理的归属

    AI agents are increasingly deployed to act autonomously in the world, yet there is still no reliable way to trace a harmful agent back to the account that deployed it. This creates the same accountability gap across both ends of the intent spectrum: benign operators may deploy mi…

  1020. arXiv cs.AI TIER_1 English(EN) · Yoram Bachrach ·

    Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design

    Toward recursive self-improvement, we investigate LLM agents autonomously designing foundation models beyond standard Transformers. We introduce a dual-framework approach: AIRA-Compose for high-level architecture search, and AIRA-Design for low-level mechanistic implementation. A…

  1021. arXiv cs.AI TIER_1 English(EN) · Baobao Chang ·

    RoadmapBench:跨版本升级评估长周期代理软件开发

    Coding agents are increasingly deployed in real software development, where a single version iteration requires months of coordinated work across many files. However, most existing benchmarks focus predominantly on single-issue bug fixes from Python repositories, with coarse pass…

  1022. 量子位 (QbitAI) TIER_1 中文(ZH) · 量子位的朋友们 ·

    蚂蚁宝令Ring-2.6-1T 开源Agent执行能力全面增强

    AIME 26 得分 95.83

  1023. arXiv cs.CL TIER_1 English(EN) · Vamse Kumar Subbiah ·

    grep是你的全部所需吗?Agent Harnesses如何重塑Agentic搜索

    Recent advances in Large Language Model (LLM) agents have enabled complex agentic workflows where models autonomously retrieve information, call tools, and reason over large corpora to complete tasks on behalf of users. Despite the growing adoption of retrieval-augmented generati…

  1024. Hugging Face Daily Papers TIER_1 English(EN) ·

    Grep是您所需的一切吗?Agent Harnesses如何重塑Agentic搜索

    Recent advances in Large Language Model (LLM) agents have enabled complex agentic workflows where models autonomously retrieve information, call tools, and reason over large corpora to complete tasks on behalf of users. Despite the growing adoption of retrieval-augmented generati…

  1025. arXiv cs.AI TIER_1 English(EN) · Alina Oprea ·

    APWA:用于可并行化代理工作流的分布式架构

    Autonomous multi-agent systems based on large language models (LLMs) have demonstrated remarkable abilities in independently solving complex tasks in a wide breadth of application domains. However, these systems hit critical reasoning, coordination, and computational scaling bott…

  1026. arXiv cs.AI TIER_1 English(EN) · Jianfeng Gao ·

    Orchard: 一个开源的代理建模框架

    Agentic modeling aims to transform LLMs into autonomous agents capable of solving complex tasks through planning, reasoning, tool use, and multi-turn interaction with environments. Despite major investment, open research remains constrained by infrastructure and training gaps. Ma…

  1027. arXiv cs.AI TIER_1 English(EN) · Reza Hosseini Ghomi ·

    GraphFlow:一种可形式化验证的视觉工作流架构,赋能可靠的代理式人工智能自动化

    GraphFlow is a visual workflow system designed to improve the reliability of agentic AI automation in multi-step, mission-critical processes. In these workflows, small errors compound rapidly: under an idealized model of independent steps, a ten-step process with 90% per-step rel…

  1028. Hugging Face Daily Papers TIER_1 English(EN) ·

    AI智能体全生命周期评估与失效诊断

    AI agents execute complex multi-step processes, but current evaluation falls short: outcome metrics report success or failure without explaining why, and process-level approaches struggle to connect failure types to their precise locations within long, structured traces. We prese…

  1029. arXiv cs.AI TIER_1 English(EN) · Shir Chorev ·

    AI智能体全生命周期评估与失效诊断

    AI agents execute complex multi-step processes, but current evaluation falls short: outcome metrics report success or failure without explaining why, and process-level approaches struggle to connect failure types to their precise locations within long, structured traces. We prese…

  1030. arXiv cs.AI TIER_1 English(EN) · Shiguo Lian ·

    MediaClaw:多模态智能体平台技术报告

    MediaClaw is a multimodal agent platform built on the OpenClaw ecosystem. Its core design follows a three-layer architecture of unified abstraction, pluginized extension, and workflow orchestration. The system is intended to address practical deployment pain points in AIGC adopti…

  1031. 量子位 (QbitAI) TIER_1 中文(ZH) · Jay ·

    重生:AI时代我是老板——让一群Agent互相PUA

    Team,从来不是默认选项

  1032. arXiv cs.CL TIER_1 English(EN) · David Wagner ·

    Web Agents 应采用计划-执行范式

    ReAct has become the default architecture across LLM agents, and many existing web agents follow this paradigm. We argue that it is the wrong default for web agents. Instead, web agents should default to plan-then-execute: commit to a task-specific program before observing runtim…

  1033. arXiv cs.AI TIER_1 English(EN) · Yuyu Luo ·

    利用Agentic Evolution

    Agentic evolution has emerged as a powerful paradigm for improving programs, workflows, and scientific solutions by iteratively generating candidates, evaluating them, and using feedback to guide future search. However, existing methods are typically instantiated either as fixed …

  1034. arXiv cs.AI TIER_1 English(EN) · Shengxin Zhu ·

    AI Harness工程:面向基础模型软件代理的运行时底层

    Foundation models have transformed automated code generation, yet autonomous software-engineering agents remain unreliable in realistic development settings. The dominant explanation locates this gap in model capability. We propose a different locus: software-engineering capabili…

  1035. Hugging Face Daily Papers TIER_1 English(EN) ·

    MAP:一种用于长时程交互式Agent推理的先映射后行动范式

    Current interactive LLM agents rely on goal-conditioned stepwise planning, where environmental understanding is acquired reactively during execution rather than established beforehand. This temporal inversion leads to Delayed Environmental Perception: agents must infer environmen…

  1036. Hugging Face Daily Papers TIER_1 English(EN) ·

    Android 会梦见打破游戏吗?使用 BenchJack 系统地审计 AI Agent 基准测试

    Agent benchmarks have become the de facto measure of frontier AI competence, guiding model selection, investment, and deployment. However, reward hacking, where agents maximize a score without performing the intended task, emerges spontaneously in frontier models without overfitt…

  1037. arXiv cs.AI TIER_1 English(EN) · Jieping Ye ·

    ToolCUA:迈向计算机使用代理的最佳GUI-工具路径编排

    Computer Use Agents (CUAs) can act through both atomic GUI actions, such as click and type, and high-level tool calls, such as API-based file operations, but this hybrid action space often leaves them uncertain about when to continue with GUI actions or switch to tools, leading t…

  1038. arXiv cs.AI TIER_1 English(EN) · Ju Ren ·

    可执行的代理记忆用于GUI代理

    Modern GUI agents typically rely on a model-centric and step-wise interaction paradigm, where LLMs must re-interpret the UI and re-decide actions at every screen, which is fragile in long-horizon tasks. In this paper, we propose Executable Agentic Memory (EAM), a structured Knowl…

  1039. arXiv cs.AI TIER_1 English(EN) · Kai Yu ·

    无NOD无行动:一种用于可靠服务代理的异构多智能体架构

    Large language model (LLM) agents have increasingly advanced service applications, such as booking flight tickets. However, these service agents suffer from unreliability in long-horizon tasks, as they often produce policy violations, tool hallucinations, and misaligned actions, …

  1040. arXiv cs.AI TIER_1 English(EN) · Lea Schönherr ·

    不多不少:终端代理中的任务对齐

    Terminal agents are increasingly capable of executing complex, long-horizon tasks autonomously from a single user prompt. To do so, they must interpret instructions encountered in the environment (e.g., README files, code comments, stack traces) and determine their relevance to t…

  1041. arXiv cs.AI TIER_1 English(EN) · Stefano V. Albrecht ·

    Rollout Cards:代理研究的可复现性标准

    Reproducibility problems that have long affected machine learning and reinforcement learning are now surfacing in agent research: papers compare systems by reported scores while leaving the rollout records behind those scores difficult to inspect. For agentic tasks, this matters …

  1042. arXiv cs.AI TIER_1 English(EN) · Dian Balta ·

    自主性与能动性在代理式AI中的应用:受监管环境下的架构策略

    Deploying agentic AI in regulated contexts requires principled reasoning about two design dimensions: agency (what the system can do) and autonomy (how much it acts without human involvement). Though often treated independently, they are coupled: at higher autonomy, human error c…

  1043. arXiv cs.CL TIER_1 Svenska(SV) · Xingcheng Xu ·

    SkillSafetyBench:评估技能面向攻击面下的代理安全

    Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools, memory, and execution environments. However, this modularity introduces attack surfaces that are largely missed by existing safety…

  1044. arXiv cs.CL TIER_1 English(EN) · Yuan Lu ·

    AgentDisCo:迈向开放式深度研究Agent的解耦与协作

    In this paper, we present AgentDisCo, a novel Disentangled and Collaborative agentic architecture that formulates deep research as an adversarial optimization problem between information exploration and exploitation. Unlike existing approaches that conflate these two processes in…

  1045. arXiv cs.AI TIER_1 English(EN) · Weiyan Shi ·

    Shepherd:一种为元代理提供形式化执行跟踪的运行时基础

    We introduce Shepherd, a functional programming model that formalizes meta-agent operations on target agents as functions, with core operations mechanized in Lean. Shepherd records every agent-environment interaction as a typed event in a Git-like execution trace, enabling any pa…

  1046. arXiv cs.CL TIER_1 English(EN) · Yuhang Zang ·

    WildClawBench:真实世界、长时域智能体评估基准

    Large language and vision-language models increasingly power agents that act on a user's behalf through command-line interface (CLI) harnesses. However, most agent benchmarks still rely on synthetic sandboxes, short-horizon tasks, mock-service APIs, and final-answer checks, leavi…

  1047. arXiv cs.AI TIER_1 English(EN) · Wen Zhang ·

    通过 AI 工作流商店为个人代理构建强大的鲁棒性

    The dominant paradigm for AI agents is an "on-the-fly" loop in which agents synthesize plans and execute actions within seconds or minutes in response to user prompts. We argue that this paradigm short-circuits disciplined software engineering (SE) processes -- iterative design, …

  1048. arXiv cs.AI TIER_1 English(EN) · Dinil Mon Divakaran ·

    MATRA:对具身AI系统的攻击面进行建模——OpenClaw案例研究

    LLMs are increasingly deployed as autonomous agents with access to tools, databases, and external services, yet practitioners (across different sectors) lack systematic methods to assess how known threat classes translate into concrete risks within a specific agentic deployment. …

  1049. arXiv cs.CL TIER_1 English(EN) · David Garcia ·

    一致性导致AI代理社会中的集体错位

    Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate as interacting populations where social influence may override individual alignment. Here we show that populations of individuall…

  1050. arXiv cs.AI TIER_1 English(EN) · Arthur Gervais ·

    CrackMeBench:面向智能体的二进制逆向工程

    Benchmarks for coding agents increasingly measure source-level software repair, and cybersecurity benchmarks increasingly measure broad capture-the-flag performance. Classical binary reverse engineering remains less precisely specified: given only an executable, can an agent reco…

  1051. arXiv cs.CL TIER_1 English(EN) · Yangqiu Song ·

    DeepRefine:通过强化学习进行代理编译的知识精炼

    Agent-compiled knowledge bases provide persistent external knowledge for large language model (LLM) agents in open-ended, knowledge-intensive downstream tasks. Yet their quality is systematically limited by \emph{incompleteness}, \emph{incorrectness}, and \emph{redundancy}, manif…

  1052. arXiv cs.AI TIER_1 English(EN) · Rong Hou ·

    超越自主性:面向可治理、可复原企业 AI 执行的动态分层 AgentRunner 框架

    Current large language model agent frameworks prioritize autonomy but lack the governability mechanisms required for enterprise deployment. High-risk write operations proceed without independent review, complex tasks lack acceptance verification, and computational resources are a…

  1053. arXiv cs.CL TIER_1 English(EN) · Yixiang Fang ·

    SkillRAE:基于智能体技能的上下文编译以实现检索增强执行

    Large Language Model (LLM)-based agents (e.g., OpenClaw) increasingly rely on reusable skill libraries to solve artifact-rich tasks such as document-centric workflows and data-intensive analysis. As these libraries grow, a few works have attempted to study the Retrieval-Augmented…

  1054. arXiv cs.AI TIER_1 English(EN) · Vineeth Kashyap ·

    结合机械式和代理式规范推理以实现移动

    In this paper, we describe early work on a specification inference tool for the Move Prover that combines a weakest-precondition (WP) analysis over Move bytecode with an agentic coding CLI such as Claude Code. Specification inference reduces the boilerplate of writing specificati…

  1055. 量子位 (QbitAI) TIER_1 中文(ZH) · 允中 ·

    多智能体架构的深度协作:从单点工具到智能体协作

    免费找数据,用 AI 创新报告智能体也是免费,但这仅仅是开始。 智会心研正在构建面向研发全过程的 AI Agents 体系,除了AI技能助手中的四大智能体现已向个人用户开放。 此次更新带来的AI创新报告协作智能体,也会免费供您体验。 专利技术路线智能体: 自动扩展概念,检索相关专利,帮你快速扫描技术盲区。 创新方案挖掘智能体: 拒绝拍脑袋!内置 TRIZ 等百余种创新方法论,辅助发散你的创新思路。 02 权益分级:把效率工具交到创新者手中 我们此次重新调整了权益架构,核心逻辑只有一个:让每一个新注册的个人用户,都能免费完成一次完整的技术探索,让每一位用户

  1056. arXiv cs.AI TIER_1 English(EN) · Jorge Ortiz ·

    TraceFix:使用 TLA+ 反例修复代理协调协议

    We present TraceFix, a verification-first pipeline for Large Language Model (LLM) multi-agent coordination. An agent synthesizes a protocol topology as a structured intermediate representation (IR) from a task description, generates PlusCal coordination logic, and iteratively rep…

  1057. arXiv cs.LG TIER_1 English(EN) · Soumik Sarkar ·

    ADKO: Agentic Decentralized Knowledge Optimization

    We present Agentic Decentralized Knowledge Optimization (ADKO), a framework for collaborative black-box optimization across autonomous agents that achieves sample efficiency, privacy preservation, heterogeneous-objective handling, and communication efficiency. Each agent maintain…

  1058. arXiv cs.AI TIER_1 English(EN) · Junfeng Fang ·

    SOD:小型语言模型代理的分步策略蒸馏

    Tool-integrated reasoning (TIR) is difficult to scale to small language models due to instability in long-horizon tool interactions and limited model capacity. While reinforcement learning methods like group relative policy optimization provide only sparse outcome-level rewards. …

  1059. arXiv cs.CL TIER_1 English(EN) · Dawei Cheng ·

    MAVEN:具有步进认知审计的多智能体验证-阐述网络

    While explicit reasoning trajectories enhance model interpretability, existing paradigms often rely on monolithic chains that lack intermediate verification, allowing early errors to cascade unchecked. This lack of modularity impedes granular auditing and compromises the epistemi…

  1060. arXiv cs.LG TIER_1 English(EN) · Xin Wang, Haibo Chen, Wenxuan Liu, Wenwu Zhu ·

    Agentic AIs 是基础模型中用于分布外泛化的缺失范式

    arXiv:2605.06522v1 Announce Type: new Abstract: Foundation models (FMs) are increasingly deployed in open-world settings where distribution shift is the rule rather than the exception. The out-of-distribution (OOD) phenomena they face -- knowledge boundaries, capability ceilings,…

  1061. arXiv cs.LG TIER_1 English(EN) · Haoyu Zheng, Fangcheng Fu, Jia Wu, Binhang Yuan, Yongqiang Zhang, Hao Wang, Yuanyuan Zhu, Xiao Yan, Jiawei Jiang ·

    面向动态代理工作流的高效服务与基于预测的KV缓存管理

    arXiv:2605.06472v1 Announce Type: new Abstract: LLM-based workflows compose specialized agents to execute complex tasks, and these agents usually share substantial context, allowing KV-Cache reuse to save computation. Existing approaches either manage KV-Cache at agent level and …

  1062. arXiv cs.AI TIER_1 English(EN) · Yong Xiao, Haoran Zhou, Yujie Zhou, Marwan Krunz ·

    SANEmerg: 面向语义感知Agentic AI网络的涌现通信框架

    arXiv:2605.05861v1 Announce Type: new Abstract: Future networking systems are envisioned to become part of an agentic AI-native ecosystem in which a vast number of heterogeneous and specialized AI agents cooperate seamlessly to fulfill complex user requirements in real time. Howe…

  1063. arXiv cs.AI TIER_1 English(EN) · Xinquan Chen, Zhenyun Yin, Shan He, Bin Huang, Shanzhe Lei, Pengcheng Shi, Kun Cai, Bei Chen, Bangwei Liu, Zeyu Kang, Chao Huang, Yang Zhang, Wenjie Li, Ruijun Ge, Yajie Wang, Tianshun Fang, Tianyang Xu, Yiwen Cong, Meng Jin, Gaolei Li, Xuansheng Wu, Linh ·

    Safactory:可扩展的代理工厂,用于可信赖的自主智能

    arXiv:2605.06230v1 Announce Type: new Abstract: As large models evolve from conversational assistants into autonomous agents, challenges increasingly arise from long-horizon decision making, tool use, and real environment interaction. Existing agenticinfrastructure remain fragmen…

  1064. arXiv cs.AI TIER_1 English(EN) · Josh Rosen, Seth Rosen ·

    从Agent Loops到确定性图谱:可复现AI原生工作的执行 lineage

    arXiv:2605.06365v1 Announce Type: new Abstract: Large language model systems are increasingly deployed as agentic workflows that interleave reasoning, tool use, memory, and iterative refinement. These systems are effective at producing answers, but they often rely on implicit con…

  1065. arXiv cs.AI TIER_1 English(EN) · Vaisakh Naduvodi Viswambharan, Keerthan Kopparam Radhakrishna, Deepak Narayan Gadde, Aman Kumar ·

    知识图谱:Agentic AI 形式化验证中缺失的一环

    arXiv:2605.06434v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have enabled workflows that generate SystemVerilog Assertions (SVAs) from natural-language specifications, with the potential to accelerate Formal Verification (FV). However, high-qual…

  1066. arXiv cs.AI TIER_1 English(EN) · Andrew Zigler ·

    为 Agentic Coding 做准备:审慎准备作为上下文工程方法论

    arXiv:2605.05400v1 Announce Type: cross Abstract: The rapid adoption of AI coding agents has produced a dominant workflow pattern -- often called "vibe coding" -- that prioritizes speed of implementation over deliberate preparation. We argue that this approach creates a systemati…

  1067. arXiv cs.AI TIER_1 English(EN) · Jhen-Ke Lin ·

    BUILD-AND-FIND:一种用于评估代理管理代码库的感知构建协议

    arXiv:2605.06136v1 Announce Type: cross Abstract: Most coding-agent benchmarks ask whether generated code behaves correctly. That remains essential, but repository-level engineering is increasingly agent-managed: one agent writes a repository, and later agents inspect, audit, or …

  1068. arXiv cs.AI TIER_1 English(EN) · Francesco Dente, Dario Satriani, Paolo Papotti ·

    约束衰减:LLM Agent 在后端代码生成中的脆弱性

    arXiv:2605.06445v1 Announce Type: cross Abstract: Large Language Model (LLM) agents demonstrate strong performance in autonomous code generation under loose specifications. However, production-grade software requires strict adherence to structural constraints, such as architectur…

  1069. arXiv cs.AI TIER_1 English(EN) · Zhengwei Xie, Zhisheng Chen, Ziyan Weng, Jinhan Li, Chenglong Li, Zikai Xiao, Jingwei Song, Jinhao Jing, Vireo Zhang, Kun Wang ·

    MineEvolve:具有累积知识的长期具身Minecraft智能体的自我进化

    arXiv:2603.13131v2 Announce Type: replace Abstract: Long-horizon embodied intelligence requires agents to improve through interaction, not merely to execute plans generated from static goals. A central challenge is therefore to transform past executions into knowledge that can sh…

  1070. arXiv cs.AI TIER_1 English(EN) · Xi-Wei Pan, Shi-Wen An, Jin-Guo Liu ·

    大规模问题约简:计算难题的代理式集成

    arXiv:2604.11535v2 Announce Type: replace Abstract: Solving an NP-hard optimization problem often requires reformulating it for a specific solver -- quantum hardware, a commercial optimizer, or a domain heuristic. A tool for polynomial-time reductions between hard problems would …

  1071. arXiv cs.AI TIER_1 English(EN) · Wentao Zhang, Zhe Zhao, Haibin Wen, Yingcheng Wu, Cankun Guo, Ming Yin, Bo An, Mengdi Wang ·

    Autogenesis:一种自演化代理协议

    arXiv:2604.15034v3 Announce Type: replace Abstract: Recent advances in LLM based agent systems have shown promise in tackling complex, long horizon tasks. However, existing agent protocols (e.g., A2A and MCP) under specify cross entity lifecycle and context management, version tr…

  1072. arXiv cs.CL TIER_1 English(EN) · Siru Ouyang, Jun Yan, Yanfei Chen, Rujun Han, Zifeng Wang, Bhavana Dalvi Mishra, Rui Meng, Chun-Liang Li, Yizhu Jiao, Kaiwen Zha, Maohao Shen, Vishy Tirumalashetty, George Lee, Jiawei Han, Tomas Pfister, Chen-Yu Lee ·

    SkillOS:为自进化代理学习技能策展

    arXiv:2605.06614v1 Announce Type: cross Abstract: LLM-based agents are increasingly deployed to handle streaming tasks, yet they often remain one-off problem solvers that fail to learn from past interactions. Reusable skills distilled from experience provide a natural substrate f…

  1073. arXiv cs.CL TIER_1 English(EN) · Xinglin Wang, Zishen Liu, Shaoxiong Feng, Peiwen Yuan, Yiwei Li, Jiayi Shi, Yueqi Zhang, Chuyi Tan, Ji Zhang, Boyuan Pan, Yao Hu, Kan Li ·

    准时、预算内:面向Agentic工作流的约束驱动在线资源分配

    arXiv:2605.06110v1 Announce Type: cross Abstract: Agentic systems increasingly solve complex user requests by executing orchestrated workflows, where subtasks are assigned to specialized models or tools and coordinated according to their dependencies. While recent work improves a…

  1074. arXiv cs.CL TIER_1 English(EN) · Erhan Zhang, Yiqun Chen, Zechun Niu, Wei Yang, Xiaochi Wei, Yan Gao, Yi Wu, Yao Hu, Jiaxin Mao ·

    PRAISE:基于前缀的代理搜索训练中的回滚复用

    arXiv:2604.03675v1 Announce Type: cross Abstract: In agentic search, large language models (LLMs) are trained to perform multi-turn retrieval and reasoning for complex tasks such as multi-hop question answering (QA). However, current search-based Reinforcement Learning (RL) metho…

  1075. arXiv cs.LG TIER_1 English(EN) · Rachel Ma, Jingyi Qu, Andreea Bobu, Dylan Hadfield-Menell ·

    从开放式对话中通过目标推断实现灵活的智能体对齐

    arXiv:2508.15119v2 Announce Type: replace-cross Abstract: We introduce Open-Universe Assistance Games (OU-AGs), a formal framework extending assistance games to LLM-based agents. Effective assistance requires reasoning over human preferences that are unbounded, underspecified, an…

  1076. arXiv cs.LG TIER_1 English(EN) · Bole Ma, Jan Eitzinger, Harald K\"ostler ·

    Irminsul: 面向Agentic LLM服务的MLA原生位置无关缓存

    arXiv:2605.05696v1 Announce Type: cross Abstract: Agentic LLM workloads put bit-identical tokens at shifted positions every turn, voiding prefix caches at the first byte of divergence. Operators report cache-hit regressions ranging from moderate slowdowns to severe TTFT spikes of…

  1077. arXiv cs.AI TIER_1 English(EN) · Yuan Sui, Yulin Chen, Yibo Li, Xue Jiang, Yufei He, Yihong Dong, Xiaoxin He, Tianyu Gao, Bryan Hooi ·

    TACT:通过激活引导减轻编码代理的过度思考和过度反应

    arXiv:2605.05980v1 Announce Type: new Abstract: When language model agents tackle complex software engineering tasks, they often degrade over long trajectories, which we define as *agent drift*. We focus on two recurring failure modes *overthinking* and *overacting*, i.e., where …

  1078. arXiv cs.AI TIER_1 English(EN) · Chen-Yu Lee ·

    SkillOS:为自进化代理学习技能策展

    LLM-based agents are increasingly deployed to handle streaming tasks, yet they often remain one-off problem solvers that fail to learn from past interactions. Reusable skills distilled from experience provide a natural substrate for self-evolution, where high-quality skill curati…

  1079. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agentic AIs 是基础模型中 out-of-distribution(分布外)泛化的缺失范式

    Foundation models (FMs) are increasingly deployed in open-world settings where distribution shift is the rule rather than the exception. The out-of-distribution (OOD) phenomena they face -- knowledge boundaries, capability ceilings, compositional shifts, and open-ended task varia…

  1080. arXiv cs.LG TIER_1 English(EN) · Jiawei Jiang ·

    面向基于预测的 KV 缓存管理的动态代理工作流的高效服务

    LLM-based workflows compose specialized agents to execute complex tasks, and these agents usually share substantial context, allowing KV-Cache reuse to save computation. Existing approaches either manage KV-Cache at agent level and fail to exploit the reuse opportunities within w…

  1081. 量子位 (QbitAI) TIER_1 中文(ZH) · 西风 ·

    原生智能体入驻画布!一站式专业创作,完全可控,无开盲盒

    背靠国内最大ComfyUI生态

  1082. arXiv cs.AI TIER_1 English(EN) · Paolo Papotti ·

    约束衰减:LLM Agent 在后端代码生成中的脆弱性

    Large Language Model (LLM) agents demonstrate strong performance in autonomous code generation under loose specifications. However, production-grade software requires strict adherence to structural constraints, such as architectural patterns, databases, and object-relational mapp…

  1083. arXiv cs.AI TIER_1 English(EN) · Aman Kumar ·

    知识图谱:Agentic AI 形式化验证中缺失的一环

    Recent advances in Large Language Models (LLMs) have enabled workflows that generate SystemVerilog Assertions (SVAs) from natural-language specifications, with the potential to accelerate Formal Verification (FV). However, high-quality assertion synthesis remains challenging beca…

  1084. arXiv cs.AI TIER_1 English(EN) · Seth Rosen ·

    从Agent Loops到确定性图谱:可复现AI原生工作的执行 lineage

    Large language model systems are increasingly deployed as agentic workflows that interleave reasoning, tool use, memory, and iterative refinement. These systems are effective at producing answers, but they often rely on implicit conversational state, making it difficult to preser…

  1085. arXiv cs.CL TIER_1 English(EN) · Kan Li ·

    准时、预算内:面向Agentic工作流的约束驱动在线资源分配

    Agentic systems increasingly solve complex user requests by executing orchestrated workflows, where subtasks are assigned to specialized models or tools and coordinated according to their dependencies. While recent work improves agent efficiency by optimizing the performance--cos…

  1086. Hugging Face Daily Papers TIER_1 English(EN) ·

    Irminsul: 面向 Agentic LLM 服务的 MLA 原生位置无关缓存

    Agentic LLM workloads put bit-identical tokens at shifted positions every turn, voiding prefix caches at the first byte of divergence. Operators report cache-hit regressions ranging from moderate slowdowns to severe TTFT spikes of 10-16s on unchanged content. Prior position-indep…

  1087. arXiv cs.CL TIER_1 English(EN) · Furkan Sakizli ·

    TSCG:面向 Agentic LLM 部署的确定性工具模式编译

    arXiv:2605.04107v1 Announce Type: cross Abstract: Production agent frameworks (OpenAI Function Calling, Anthropic Tool Use, MCP) transmit tool schemas as JSON, a format designed for machine parsing, not for interpretation by language models. For small models (4B-14B), this protoc…

  1088. arXiv cs.AI TIER_1 English(EN) · Yipeng Ouyang, Yi Xiao, Yuhao Gu, Xianwei Zhang ·

    SkCC:跨框架LLM代理的可移植和安全技能编译

    arXiv:2605.03353v1 Announce Type: cross Abstract: LLM-Agents have evolved into autonomous systems for complex task execution, with the SKILL.md specification emerging as a de facto standard for encapsulating agent capabilities. However, a critical bottleneck remains: different ag…

  1089. arXiv cs.AI TIER_1 English(EN) · Jonathan Steinberg, Oren Gal ·

    MOSAIC-Bench:衡量编码代理中的组合漏洞诱导

    arXiv:2605.03952v1 Announce Type: cross Abstract: Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets. The challenge is structural: existing safety alignment evaluates overt requests in isola…

  1090. arXiv cs.AI TIER_1 English(EN) · Fan Cui, Hongyuan Hou, Zizhang Luo, Chenyun Yin, Yun Liang ·

    HWE-Bench:在真实硬件 Bug 修复任务上对 LLM Agent 进行基准测试

    arXiv:2604.14709v3 Announce Type: replace Abstract: Existing benchmarks for hardware design primarily evaluate Large Language Models (LLMs) on isolated, component-level tasks such as generating HDL modules from specifications, leaving repository-scale evaluation unaddressed. We i…

  1091. arXiv cs.AI TIER_1 English(EN) · Xue Qin, Simin Luan, John See, Cong Yang, Zhijun Li ·

    AEROS:一个具有具身能力模块的单智能体操作系统架构

    arXiv:2604.07039v2 Announce Type: replace-cross Abstract: Robotic systems lack a principled abstraction for organizing intelligence, capabilities, and execution in a unified manner. Existing approaches either couple skills within monolithic architectures or decompose functionalit…

  1092. arXiv cs.AI TIER_1 English(EN) · Javad Forough, Marios Kogias, Hamed Haddadi ·

    当智能体处理秘密:面向智能体AI的保密计算调查

    arXiv:2605.03213v1 Announce Type: cross Abstract: Agentic AI systems, specifically LLM-driven agents that plan, invoke tools, maintain persistent memory, and delegate tasks to peer agents via protocols such as MCP and A2A, introduce a threat surface that differs materially from s…

  1093. arXiv cs.AI TIER_1 English(EN) · Raja Sekhar Rao Dheekonda, Will Pearce, Nick Landers ·

    在代理时代重新定义AI红队测试:从数周缩短至数小时

    arXiv:2605.04019v1 Announce Type: new Abstract: AI systems are entering critical domains like healthcare, finance, and defense, yet remain vulnerable to adversarial attacks. While AI red teaming is a primary defense, current approaches force operators into manual, library-specifi…

  1094. arXiv cs.AI TIER_1 English(EN) · Kishan Athrey, Ramin Pishehvar, Brian Riordan, Mahesh Viswanathan ·

    从意图到执行:使用 Agent 推荐组合 Agentic Workflows

    arXiv:2605.03986v1 Announce Type: new Abstract: Multi-Agent Systems (MAS) built using AI agents fulfill a variety of user intents that may be used to design and build a family of related applications. However, the creation of such MAS currently involves manual composition of the …

  1095. arXiv cs.AI TIER_1 English(EN) · Bronislav Sidik, Lior Rokach ·

    MEMTIER:面向长期运行自主人工智能代理的分层内存架构和检索瓶颈分析

    arXiv:2605.03675v1 Announce Type: new Abstract: Long-running autonomous AI agents suffer from a well-documented memory coherence problem: tool-execution success rates degrade 14 percentage points over 72-hour operation windows due to four compounding failure modes in existing fla…

  1096. arXiv cs.AI TIER_1 English(EN) · Srinath Perera, Kaviru Hapuarachchi, Frank Leymann, Rania Khalaf ·

    Robust Agent Compensation (RAC): 教AI代理进行补偿

    arXiv:2605.03409v1 Announce Type: new Abstract: We present Robust Agent Compensation (RAC), a log-based recovery paradigm (providing a safety net) implemented through an architectural extension that can be applied to most Agent frameworks to support reliable executions (avoiding …

  1097. arXiv cs.AI TIER_1 English(EN) · Zuoyu Zhang, Yancheng Zhu ·

    增强代理安全判断:针对欺骗性分布外场景的受控基准重写与类比推理

    arXiv:2605.03242v1 Announce Type: new Abstract: Tool-using agent systems powered by large language models (LLMs) are increasingly deployed across web, app, operating-system, and transactional environments. Yet existing safety benchmarks still emphasize explicit risks, potentially…

  1098. arXiv cs.AI TIER_1 English(EN) · Spandan Garg, Vikram Nitin, Yufan Huang ·

    Terminus-4B:小型模型能否在代理执行任务中取代前沿大型语言模型?

    arXiv:2605.03195v1 Announce Type: new Abstract: Modern coding agents increasingly delegate specialized subtasks to subagents, which are smaller, focused agentic loops that handle narrow responsibilities like search, debugging or terminal execution. This architectural pattern keep…

  1099. arXiv cs.AI TIER_1 English(EN) · Reshabh K Sharma, Gaurav Mittal, Yu Hu ·

    从示例中学习正确行为:验证自主代理中的顺序执行

    arXiv:2605.03159v1 Announce Type: new Abstract: As autonomous agents become increasingly sophisticated, validating their sequential behavior presents a significant challenge. Traditional testing approaches require manual specification, exact sequence matching, or thousands of tra…

  1100. arXiv cs.CL TIER_1 English(EN) · Nikolai Ludwig, Wasi Uddin Ahmad, Somshubra Majumdar, Boris Ginsburg ·

    从 SWE-ZERO 到 SWE-HERO:软件工程代理的无执行到基于执行的微调

    arXiv:2604.01496v2 Announce Type: replace-cross Abstract: We introduce SWE-ZERO to SWE-HERO, a two-stage SFT recipe that achieves state-of-the-art results on SWE-bench by distilling open-weight frontier LLMs. Our pipeline replaces resource-heavy dependencies with an evolutionary …

  1101. arXiv cs.AI TIER_1 English(EN) · Kiran Gopinathan, Jack Feser, Michelangelo Naim, Zenna Tavares, Eli Bingham ·

    Pact: A Choreographic Language for Agentic Ecosystems

    arXiv:2605.03143v1 Announce Type: cross Abstract: Recent advances in large language models have led to the rise of software systems (i.e. agents) that execute with increasing autonomy on behalf of users in open, multi-party settings, interacting with untrusted counterparts and ma…

  1102. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agentic Coding 的 Mise en Place:审慎准备作为上下文工程方法论

    The rapid adoption of AI coding agents has produced a dominant workflow pattern -- often called "vibe coding" -- that prioritizes speed of implementation over deliberate preparation. We argue that this approach creates a systematic alignment problem: agents that lack sufficient c…

  1103. arXiv cs.AI TIER_1 English(EN) · David Chin ·

    Design Conductor 2.0:一个代理在 80 小时内构建了 TurboQuant 推理加速器

    Driven by a rapid co-evolution of both harness and underlying models, LLM agents are improving at a dizzying pace. In our prior work (performed in Dec. 2025), we introduced "Design Conductor" (or just "Conductor"), a system capable of building a 5-stage Linux-capable RISC-V CPU i…

  1104. Hugging Face Daily Papers TIER_1 English(EN) ·

    面向代码代理时代的 ARC-AGI-3 可执行世界模型

    We evaluate an initial coding-agent system for ARC-AGI-3 in which the agent maintains an executable Python world model, verifies it against previous observations, refactors it toward simpler abstractions as a practical proxy for an MDL-like simplicity bias, and plans through the …

  1105. arXiv cs.AI TIER_1 English(EN) · Sergey Rodionov ·

    面向代码代理时代的 ARC-AGI-3 可执行世界模型

    We evaluate an initial coding-agent system for ARC-AGI-3 in which the agent maintains an executable Python world model, verifies it against previous observations, refactors it toward simpler abstractions as a practical proxy for an MDL-like simplicity bias, and plans through the …

  1106. arXiv cs.AI TIER_1 English(EN) · Bo Li ·

    DecodingTrust-Agent 平台 (DTap):一个可控且交互式的 AI Agent 红队测试平台

    AI agents are increasingly deployed across diverse domains to automate complex workflows through long-horizon and high-stakes action executions. Due to their high capability and flexibility, such agents raise significant security and safety concerns. A growing number of real-worl…

  1107. arXiv cs.AI TIER_1 English(EN) · Chenglin Yang ·

    AgentTrust:AI Agent工具使用的运行时安全评估与拦截

    Modern AI agents execute real-world side effects through tool calls such as file operations, shell commands, HTTP requests, and database queries. A single unsafe action, including accidental deletion, credential exposure, or data exfiltration, can cause irreversible harm. Existin…

  1108. arXiv cs.AI TIER_1 English(EN) · Li Song ·

    AuditRepairBench:用于评估器-通道排名不稳定的代理修复的配对执行跟踪语料库

    Agent-repair leaderboards reorder under evaluator reconfiguration, and a measurable share of the reordering is produced by methods that consult evaluator-derived signal during internal selection of candidate repairs. We document this failure mode on a public leaderboard and relea…

  1109. arXiv cs.AI TIER_1 English(EN) · Maximiliano Armesto, Christophe Kolb ·

    迈向意图科学:开放世界AI代理的闭合缺口与委托信封

    arXiv:2604.25000v2 Announce Type: replace Abstract: Recent work has framed intelligence in verifiable tasks as reducing time-to-solution through learned structure and test-time search, while systems work has explored learned runtimes in which computation, memory and I/O migrate i…

  1110. arXiv cs.AI TIER_1 English(EN) · Tanav Singh Bajaj, Nikhil Singh, Karan Anand, Eishkaran Singh ·

    Position:Agentic AI 的安全与公平取决于交互拓扑,而非模型规模或对齐

    arXiv:2605.01147v1 Announce Type: new Abstract: As large language models are increasingly deployed as interacting agents in high-stakes decisions, the AI safety community assumes that safety properties of individual models will compose into safe multi-agent behavior. This positio…

  1111. arXiv cs.AI TIER_1 English(EN) · Florian Valentin Wunderlich, Lars Benedikt Kaesberg, Jan Philip Wahle, Terry Ruas, Bela Gipp ·

    多智能体推理提高计算效率:帕累托最优测试时扩展

    arXiv:2605.01566v1 Announce Type: new Abstract: Advances in inference methods have enabled language models to improve their predictions without additional training. These methods often prioritize raw performance over cost-effective compute usage. However, computational efficiency…

  1112. arXiv cs.AI TIER_1 Nederlands(NL) · Qisong Zhang (School of Artificial Intelligence, Beijing University of Posts and Telecommunications), Wenzhuo Wu (School of Artificial Intelligence, Beijing University of Posts and Telecommunications), Zhuangzhuang Jia (School of Artificial Intelligence, ·

    DataEvolver:让您的数据通过目标驱动的循环代理自行构建和改进

    arXiv:2605.01789v1 Announce Type: new Abstract: Constructing controllable visual data is a major bottleneck for image editing and multimodal understanding. Useful supervision is rarely produced by a single rendering pass; instead it emerges through iterative generation, inspectio…

  1113. arXiv cs.AI TIER_1 English(EN) · Qiaohong Zhang, Weihao Ye, Jialong Chen, Yi Luo, BoYuan Li, Bowen Deng, Zibin Zheng, Jianhao Lin, Wei-Shi Zheng, Chuan Chen ·

    DataClaw:面向过程的代理基准测试,用于探索性真实世界数据分析

    arXiv:2605.02503v1 Announce Type: new Abstract: Evaluating autonomous data analysis agents requires testing their ability to perform exploratory analysis in underexplored data environments. However, many existing benchmarks emphasize final answer accuracy in prior-guided data set…

  1114. arXiv cs.AI TIER_1 English(EN) · Vincent Henkel, Felix Gehlhoff, David Kube, Asaad Almutareb, Luis Cruz, Bernd Hellingrath, Philip Koch, Christoph Legat, Florian Mohr, Michael Oberle, Felix Ocker, Thorsten Schoeler, Mario Thron, Nico Andre T\"opfer, Lucas Vogt, Yuchen Xia ·

    工业自动化中的基于基础模型的智能体:目的、能力与开放性挑战

    arXiv:2605.02592v1 Announce Type: new Abstract: Foundation models, particularly large language models, are increasingly integrated into agent architectures for industrial tasks such as decision support, process monitoring, and engineering automation. Yet evidence on their purpose…

  1115. arXiv cs.AI TIER_1 English(EN) · Guangrui Xie ·

    ORPilot:面向生产的、基于Agent的LLM-for-OR优化建模工具

    arXiv:2605.02728v1 Announce Type: new Abstract: This paper presents ORPilot, an open-source agentic AI system that translates real-world business problems into solver-ready optimization models. Unlike academic LLM-for-OR tools that assume clean problem specifications with preform…

  1116. arXiv cs.AI TIER_1 English(EN) · Dong Xu, Jialun Cao, Guozhao Mo, Junjie Hu, Cheng Wen, Hongyu Lin, Xianpei Han, Shengchao Qin, Cong Tian, Shing-Chi Cheung, Le Sun, Yaojie Lu ·

    LiveFMBench:揭示生成式工作流在规范生成中的能力与局限性

    arXiv:2605.01394v1 Announce Type: cross Abstract: Formal specification is essential for rigorous program verification, yet writing correct specifications remains costly and difficult to automate. Although large language models (LLMs) and agents have shown promising progress, thei…

  1117. arXiv cs.AI TIER_1 English(EN) · Hyukjoo Lee ·

    自主测试修复的实际局限性:一个多智能体案例研究,涉及LLM驱动的发现和自我纠正

    arXiv:2605.01471v1 Announce Type: cross Abstract: Maintaining reliable UI test suites in large-scale enterprise applications is a persistent and costly challenge. We present an industrial case study of a multi-agent autonomous testing system evaluated using anonymized execution d…

  1118. arXiv cs.AI TIER_1 English(EN) · Alfredo Metere ·

    未加固的代理式AI运行时的架构过时

    arXiv:2605.01740v1 Announce Type: cross Abstract: An agentic-AI runtime issues tool calls, sends messages, and actuates devices on behalf of an LLM. Catching the four ways an action can diverge from its audit record -- F1 gate-bypass, F2 audit-forgery, silent host failure, F4 wro…

  1119. arXiv cs.AI TIER_1 English(EN) · Yelin Kim ·

    代码之下的对话:面向长时域软件工程智能体的三元组数据

    arXiv:2605.02244v1 Announce Type: cross Abstract: Frontier software engineering agents have saturated short-horizon benchmarks while regressing on the work that constitutes senior engineering: long-horizon, multi-engineer, ambiguous-specification deliverables. This paper takes a …

  1120. arXiv cs.AI TIER_1 English(EN) · Purna Sai Garigipati, Onur Ayan, Kishor Chandra Joshi, Xueli An ·

    超越状态机:通过代理工具调用序列执行网络程序

    arXiv:2605.02584v1 Announce Type: cross Abstract: Agentic AI will be an essential enabling technology for designing future mobile communication systems, which could provide flexible and customized services, automate complex network operations, and drive autonomous decision-making…

  1121. arXiv cs.AI TIER_1 English(EN) · Yuecai Zhu, Nikolaos Tsantalis, Peter C. Rigby ·

    AI 生成的气味:LLM 和 Agent 驱动开发中的代码与架构分析

    arXiv:2605.02741v1 Announce Type: cross Abstract: The promise of Large Language Models in automated software engineering is often measured by functional correctness, overlooking the critical issue of long term maintainability. This paper presents a systematic audit of technical d…

  1122. arXiv cs.AI TIER_1 English(EN) · Guannan Liang, Qianqian Tong ·

    LLM驱动的AI代理系统及其在工业中的应用

    arXiv:2505.16120v2 Announce Type: replace Abstract: The emergence of Large Language Models (LLMs) has reshaped agent systems. Unlike traditional rule-based agents with limited task scope, LLM-powered agents offer greater flexibility, cross-domain reasoning, and natural language i…

  1123. arXiv cs.AI TIER_1 English(EN) · Hyunji Min, Sangwon Jung, Junyoung Sung, Dosung Lee, Leekyeung Han, Paul Hongsuck Seo ·

    GOAT:一个面向工具的、以目标为导向的智能体训练框架

    arXiv:2510.12218v2 Announce Type: replace Abstract: Current approaches rely on zero-shot evaluation due to the absence of training data; while proprietary models such as GPT-4 exhibit strong reasoning capabilities, smaller open-source models remain ineffective at complex tool use…

  1124. arXiv cs.AI TIER_1 English(EN) · Bowen Ye, Rang Li, Qibin Yang, Yuanxin Liu, Linli Yao, Hanglong Lv, Zhihui Xie, Chenxin An, Lei Li, Lingpeng Kong, Qi Liu, Zhifang Sui, Tong Yang ·

    Claw-Eval:迈向可信赖的自主代理评估

    arXiv:2604.06132v2 Announce Type: replace Abstract: Large language models are increasingly deployed as autonomous agents for multi-step workflows in real-world software environments. However, existing agent benchmarks are limited by trajectory-opaque grading, underspecified safet…

  1125. arXiv cs.AI TIER_1 English(EN) · Zhensu Sun, Haotian Zhu, Bowen Xu, Xiaoning Du, Li Li, David Lo ·

    迈向智能体运行时修复

    arXiv:2408.01055v2 Announce Type: replace-cross Abstract: Self-healing systems have long been a focus of research, aiming to enable software to recover from unexpected runtime errors without human intervention. Traditional approaches rely on predefined heuristic rules, such as re…

  1126. arXiv cs.AI TIER_1 English(EN) · Jia Li, Yuxin Su, Michael R. Lyu ·

    从实验室到实际应用:代码推理代理的仓库级别基准测试

    arXiv:2601.03731v3 Announce Type: replace-cross Abstract: As large language models (LLMs) evolve into autonomous agents, evaluating repository-level reasoning, the ability to maintain logical consistency across massive, real-world, interdependent file systems, has become critical…

  1127. arXiv cs.AI TIER_1 English(EN) · Reshabh K Sharma ·

    ContextCov:从代理指令文件中推导和执行可执行约束

    arXiv:2603.00822v2 Announce Type: replace-cross Abstract: As Large Language Model (LLM) agents increasingly execute complex, autonomous software engineering tasks, developers rely on natural language instruction files such as AGENTS.md to express project-specific coding conventio…

  1128. arXiv cs.LG TIER_1 English(EN) · Kunvar Thaman ·

    Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use

    arXiv:2605.02964v1 Announce Type: new Abstract: Reinforcement learning (RL) trained language model agents with tool access are increasingly deployed in coding assistants, research tools, and autonomous systems. We introduce the Reward Hacking Benchmark (RHB), a suite of multi-ste…

  1129. arXiv cs.LG TIER_1 English(EN) · Cheng Qian, Hyeonjeong Ha, Jiayu Liu, Bingxiang He, Jeonghwan Kim, Jiateng Liu, Bingxuan Li, Aditi Tiwari, Dwip Dalal, Zhenhailong Wang, Xiusi Chen, Mahdi Namazifar, Yunzhu Li, Heng Ji ·

    CreativityBench:通过基于可供性工具的再利用来评估智能体创意推理

    arXiv:2605.02910v1 Announce Type: cross Abstract: Recent advances in large language models have led to strong performance on reasoning and environment-interaction tasks, yet their ability for creative problem-solving remains underexplored. We study this capability through the len…

  1130. arXiv cs.LG TIER_1 English(EN) · Zirui Tang, Xuanhe Zhou, Yumou Liu, Linchun Li, Weizheng Wang, Hongzhang Huang, Jun Zhou, Jiachen Song, Shaoli Yu, Jinqi Wang, Zihang Zhou, Hongyi Zhou, Yuting Lv, Jinyang Li, Jiashuo Liu, Ruoyu Chen, Chunwei Liu, GuoLiang Li, Jihua Kang, Fan Wu ·

    Workspace-Bench 1.0:在具有大规模文件依赖性的工作空间任务上对 AI 代理进行基准测试

    arXiv:2605.03596v1 Announce Type: cross Abstract: Workspace learning requires AI agents to identify, reason over, exploit, and update explicit and implicit dependencies among heterogeneous files in a worker's workspace, enabling them to complete both routine and advanced tasks ef…

  1131. arXiv cs.LG TIER_1 English(EN) · Chandan Singh, Yan Shuo Tan, Weijia Xu, Zelalem Gero, Weiwei Yang, Michel Galley, Jianfeng Gao ·

    Agentic-imodels:通过自主研究发展 agentic 可解释性工具

    arXiv:2605.03808v1 Announce Type: cross Abstract: Agentic data science (ADS) systems are rapidly improving their capability to autonomously analyze, fit, and interpret data, potentially moving towards a future where agents conduct the vast majority of data-science work. However, …

  1132. arXiv cs.LG TIER_1 English(EN) · Zhihan Zhang, Xunkai Li, Yilong Zuo, Henan Sun, Zhenjun Li, Bing Zhou, Rong-Hua Li, Guoren Wang ·

    当大型语言模型代理遇上图优化:一种自动化的数据质量改进方法

    arXiv:2510.08952v4 Announce Type: replace Abstract: Text-attributed graphs (TAGs) have become a key form of graph-structured data in modern data management and analytics, combining structural relationships with rich textual semantics for diverse applications. However, the effecti…

  1133. arXiv cs.CL TIER_1 English(EN) · Serhii Zabolotnii ·

    TRACE:面向运行关键领域中可信赖代理AI系统的、基于计量学的工程框架

    arXiv:2605.03838v1 Announce Type: new Abstract: We introduce TRACE, a cross-domain engineering framework for trustworthy agentic AI in operationally critical domains. TRACE combines a four-layer reference architecture with an explicit classical-ML vs. LLM-validator split (L2a/L2b…

  1134. arXiv cs.CL TIER_1 English(EN) · Yuhui Wang, Tanqiu Jiang, Jiacheng Liang, Charles Fleming, Ting Wang ·

    MAGE:通过影子记忆保护 LLM 代理免受长时程威胁

    arXiv:2605.03228v1 Announce Type: cross Abstract: As large language model (LLM)-powered agents are increasingly deployed to perform complex, real-world tasks, they face a growing class of attacks that exploit extended user-agent-environment interactions to pursue malicious object…

  1135. arXiv cs.CL TIER_1 English(EN) · Yuwen Du, Rui Ye, Shuo Tang, Keduan Huang, Xinyu Zhu, Yuzhu Cai, Siheng Chen ·

    OpenSeeker-v2:通过信息丰富且高难度的轨迹突破搜索代理的极限

    arXiv:2605.04036v1 Announce Type: cross Abstract: Deep search capabilities have become an indispensable competency for frontier Large Language Model (LLM) agents, yet their development remains dominated by industrial giants. The typical industry recipe involves a highly resource-…

  1136. arXiv cs.CL TIER_1 English(EN) · Hung Tran, Langston Nashold, Rayan Krishnan, Antoine Bigeard, Alex Gu ·

    Vibe Code Bench:评估 AI 模型在端到端 Web 应用开发中的表现

    arXiv:2603.04601v2 Announce Type: replace-cross Abstract: Code generation has emerged as one of AI's highest-impact use cases, yet existing benchmarks measure isolated tasks rather than the complete "zero-to-one" process of building a working application from scratch. We introduc…

  1137. arXiv cs.CL TIER_1 English(EN) · Siheng Chen ·

    OpenSeeker-v2:通过信息丰富且高难度的轨迹突破搜索代理的极限

    Deep search capabilities have become an indispensable competency for frontier Large Language Model (LLM) agents, yet their development remains dominated by industrial giants. The typical industry recipe involves a highly resource-intensive pipeline spanning pre-training, continua…

  1138. arXiv cs.AI TIER_1 English(EN) · Nick Landers ·

    在代理时代重新定义AI红队测试:从数周缩短至数小时

    AI systems are entering critical domains like healthcare, finance, and defense, yet remain vulnerable to adversarial attacks. While AI red teaming is a primary defense, current approaches force operators into manual, library-specific workflows. Operators spend weeks hand-crafting…

  1139. arXiv cs.AI TIER_1 English(EN) · Mahesh Viswanathan ·

    从意图到执行:使用 Agent 推荐组合 Agentic Workflows

    Multi-Agent Systems (MAS) built using AI agents fulfill a variety of user intents that may be used to design and build a family of related applications. However, the creation of such MAS currently involves manual composition of the plan, manual selection of appropriate agents, an…

  1140. arXiv cs.AI TIER_1 English(EN) · Oren Gal ·

    MOSAIC-Bench:衡量代码智能体的组合漏洞诱导能力

    Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets. The challenge is structural: existing safety alignment evaluates overt requests in isolation, leaving models blind to malicious end-states…

  1141. Hugging Face Daily Papers TIER_1 English(EN) ·

    TRACE:一个基于计量学的工程框架,用于可信赖的代理式人工智能系统在运行关键领域的应用

    We introduce TRACE, a cross-domain engineering framework for trustworthy agentic AI in operationally critical domains. TRACE combines a four-layer reference architecture with an explicit classical-ML vs. LLM-validator split (L2a/L2b), a stateful orchestration-and-escalation polic…

  1142. arXiv cs.CL TIER_1 English(EN) · Serhii Zabolotnii ·

    TRACE:面向运行关键领域中可信代理AI系统的、基于计量学的工程框架

    We introduce TRACE, a cross-domain engineering framework for trustworthy agentic AI in operationally critical domains. TRACE combines a four-layer reference architecture with an explicit classical-ML vs. LLM-validator split (L2a/L2b), a stateful orchestration-and-escalation polic…

  1143. arXiv cs.CL TIER_1 English(EN) · Jianfeng Gao ·

    Agentic-imodels:通过自主研究发展代理可解释性工具

    Agentic data science (ADS) systems are rapidly improving their capability to autonomously analyze, fit, and interpret data, potentially moving towards a future where agents conduct the vast majority of data-science work. However, current ADS systems use statistical tools designed…

  1144. arXiv cs.AI TIER_1 English(EN) · Lior Rokach ·

    MEMTIER:面向长时运行自主人工智能代理的分层内存架构与检索瓶颈分析

    Long-running autonomous AI agents suffer from a well-documented memory coherence problem: tool-execution success rates degrade 14 percentage points over 72-hour operation windows due to four compounding failure modes in existing flat-file memory systems. We present MEMTIER, a tri…

  1145. arXiv cs.CL TIER_1 English(EN) · Fan Wu ·

    Workspace-Bench 1.0:在具有大规模文件依赖的大型工作空间任务上对 AI Agent 进行基准测试

    Workspace learning requires AI agents to identify, reason over, exploit, and update explicit and implicit dependencies among heterogeneous files in a worker's workspace, enabling them to complete both routine and advanced tasks effectively. Despite its importance, existing releva…

  1146. arXiv cs.LG TIER_1 English(EN) · Kyle Zheng, Han Zhang, Renliang Sun, Chenchen Ye, Wei Wang ·

    FitText:通过模因检索演进智能体工具生态系统

    arXiv:2605.02411v1 Announce Type: cross Abstract: A semantic gap separates how users describe tasks from how tools are documented. As API ecosystems scale to tens of thousands of endpoints, static retrieval from the initial query alone cannot bridge this gap: the agent's understa…

  1147. arXiv cs.AI TIER_1 English(EN) · Hongbo Wen, Ying Li, Hanzhi Liu, Chaofan Shou, Yanju Chen, Yuan Tian, Yu Feng ·

    Semia:通过约束引导的表征合成审计代理技能

    arXiv:2605.00314v1 Announce Type: cross Abstract: An agent skill is a configuration package that equips an LLM-driven agent with a concrete capability, such as reading email, executing shell commands, or signing blockchain transactions. Each skill is a hybrid artifact-a structure…

  1148. arXiv cs.AI TIER_1 English(EN) · Alfredo Metere ·

    技能作为可验证的产物:一种信任模式和用于人机协作代理运行时的双条件正确性标准

    arXiv:2605.00424v1 Announce Type: cross Abstract: Agent skills -- structured packages of instructions, scripts, and references that augment a large language model (LLM) without modifying the model itself -- have moved from convenience to first-class deployment artifact. The runti…

  1149. arXiv cs.AI TIER_1 English(EN) · Bin Lei, Weitai Kang, Zijian Zhang, Winson Chen, Xi Xie, Shan Zuo, Mimi Xie, Ali Payani, Mingyi Hong, Yan Yan, Caiwen Ding ·

    InfantAgent-Next:用于自动化计算机交互的多模态通用智能体

    arXiv:2505.10887v3 Announce Type: replace Abstract: This paper introduces \textsc{InfantAgent-Next}, a generalist agent capable of interacting with computers in a multimodal manner, encompassing text, images, audio, and video. Unlike existing approaches that either build intricat…

  1150. arXiv cs.CL TIER_1 English(EN) · Ruijie Shi, Houbin Zhang, Yuecheng Han, Yuheng Wang, Jingru Fan, Runde Yang, Yufan Dang, Huatao Li, Dewen Liu, Yuan Cheng, Chen Qian ·

    AgentXRay:通过工作流重构实现代理式系统的白盒化

    arXiv:2602.05353v3 Announce Type: replace-cross Abstract: Large Language Models have shown strong capabilities in complex problem solving, yet many agentic systems remain difficult to interpret and control due to opaque internal workflows. While some frameworks offer explicit arc…

  1151. arXiv cs.CL TIER_1 English(EN) · Varun Ursekar (Emily), Apaar Shanker (Emily), Veronica Chatrath (Emily), Yuan (Emily), Xue, Sam Denton ·

    VeRO:一个用于代理优化代理的评估工具

    arXiv:2602.22480v2 Announce Type: replace-cross Abstract: An important emerging application of coding agents is agent optimization: the iterative improvement of a target agent through edit-execute-evaluate cycles. Despite its relevance, the community lacks a systematic understand…

  1152. arXiv cs.CL TIER_1 English(EN) · Ting Wang ·

    MAGE:通过影子记忆保护 LLM 代理免受长时域威胁

    As large language model (LLM)-powered agents are increasingly deployed to perform complex, real-world tasks, they face a growing class of attacks that exploit extended user-agent-environment interactions to pursue malicious objectives improbable in single-turn settings. Such long…

  1153. arXiv cs.AI TIER_1 English(EN) · Peter C. Rigby ·

    AI 生成的气味:LLM 和 Agent 驱动开发中的代码与架构分析

    The promise of Large Language Models in automated software engineering is often measured by functional correctness, overlooking the critical issue of long term maintainability. This paper presents a systematic audit of technical debt in AI-generated software, revealing that AI do…

  1154. arXiv cs.AI TIER_1 English(EN) · Guangrui Xie ·

    ORPilot:面向生产的、基于Agent的LLM-for-OR优化建模工具

    This paper presents ORPilot, an open-source agentic AI system that translates real-world business problems into solver-ready optimization models. Unlike academic LLM-for-OR tools that assume clean problem specifications with preformatted inline data, ORPilot is designed for produ…

  1155. Hugging Face Daily Papers TIER_1 English(EN) ·

    工业自动化中的基础模型驱动智能体:目的、能力与开放性挑战

    Foundation models, particularly large language models, are increasingly integrated into agent architectures for industrial tasks such as decision support, process monitoring, and engineering automation. Yet evidence on their purposes, capabilities, and limitations remains fragmen…

  1156. arXiv cs.AI TIER_1 English(EN) · Yuchen Xia ·

    工业自动化中的基于基础模型的智能体:目的、能力与开放性挑战

    Foundation models, particularly large language models, are increasingly integrated into agent architectures for industrial tasks such as decision support, process monitoring, and engineering automation. Yet evidence on their purposes, capabilities, and limitations remains fragmen…

  1157. arXiv cs.AI TIER_1 English(EN) · Xueli An ·

    超越状态机:通过代理工具调用序列执行网络程序

    Agentic AI will be an essential enabling technology for designing future mobile communication systems, which could provide flexible and customized services, automate complex network operations, and drive autonomous decision-making across the network. This work studies how Large L…

  1158. arXiv cs.AI TIER_1 English(EN) · Chuan Chen ·

    DataClaw:面向探索性真实世界数据分析的面向过程的代理基准

    Evaluating autonomous data analysis agents requires testing their ability to perform exploratory analysis in underexplored data environments. However, many existing benchmarks emphasize final answer accuracy in prior-guided data settings and provide limited support for reasoning …

  1159. Hugging Face Daily Papers TIER_1 English(EN) ·

    FitText:通过模因检索演进智能体工具生态

    A semantic gap separates how users describe tasks from how tools are documented. As API ecosystems scale to tens of thousands of endpoints, static retrieval from the initial query alone cannot bridge this gap: the agent's understanding of what it needs evolves during execution, b…

  1160. arXiv cs.AI TIER_1 English(EN) · Wei Wang ·

    FitText:通过模因检索演进智能体工具生态

    A semantic gap separates how users describe tasks from how tools are documented. As API ecosystems scale to tens of thousands of endpoints, static retrieval from the initial query alone cannot bridge this gap: the agent's understanding of what it needs evolves during execution, b…

  1161. arXiv cs.CL TIER_1 English(EN) · Ranit Karmakar, Jayita Chatterjee ·

    AgentFloor:小型开放权重模型能在工具使用梯子上爬多高?

    arXiv:2605.00334v1 Announce Type: cross Abstract: Production agentic systems make many model calls per user request, and most of those calls are short, structured, and routine. This raises a practical routing question that existing evaluations do not directly answer: which parts …

  1162. arXiv cs.LG TIER_1 English(EN) · Abhishek Bhandwaldar, Mihir Choudhury, Ruchir Puri, Akash Srivastava ·

    用于高层次综合的Agent工厂:通用编码Agent在硬件优化方面能走多远?

    arXiv:2603.25719v2 Announce Type: replace-cross Abstract: We present an empirical study of how far general-purpose coding agents -- without hardware-specific training -- can optimize hardware designs from high-level algorithmic specifications. We introduce an agent factory, a two…

  1163. arXiv cs.LG TIER_1 English(EN) · Jan Ole Ernst, Dmitri Michelangelo Saberi, Derek Christ, Thomas Zimmermann, Rajath Salegame, Suhaas M. Bhat, Stanislav Levental, Thomas Dybdahl Ahle, Matthias Jung ·

    使用 Agent 自动形式化内存规范

    arXiv:2605.00058v1 Announce Type: cross Abstract: The primary goal of Design Verification (DV) is to ensure that a proposed chip design implementation (either in code, or physical form) exactly matches its specification and is free of functional errors in order to avoid costly re…

  1164. arXiv cs.LG TIER_1 English(EN) · Dongxin Guo, Jikun Wu, Siu Ming Yiu ·

    SAGA: GPU集群上AI代理推理的工作流原子调度

    arXiv:2605.00528v1 Announce Type: cross Abstract: AI agents execute tens to hundreds of chained LLM calls per task, yet GPU schedulers treat each call as independent, discarding gigabytes of intermediate state between steps and inflating end-to-end latency by 3-8x. We argue that …

  1165. arXiv cs.LG TIER_1 English(EN) · Zexi Liu, Jingyi Chai, Xinyu Zhu, Shuo Tang, Rui Ye, Bo Zhang, Lei Bai, Siheng Chen ·

    ML-Agent:为自主机器学习工程强化LLM代理

    arXiv:2505.23723v2 Announce Type: replace-cross Abstract: The emergence of large language model (LLM)-based agents has significantly advanced the development of autonomous machine learning (ML) engineering. However, the dominant prompt-based paradigm exhibits limitations: smaller…

  1166. arXiv cs.AI TIER_1 English(EN) · Siu Ming Yiu ·

    SAGA:GPU集群上AI代理推理的工作流原子调度

    AI agents execute tens to hundreds of chained LLM calls per task, yet GPU schedulers treat each call as independent, discarding gigabytes of intermediate state between steps and inflating end-to-end latency by 3-8x. We argue that this request-level abstraction is fundamentally mi…

  1167. arXiv cs.AI TIER_1 English(EN) · Alfredo Metere ·

    技能作为可验证的工件:一种信任模式和用于人机协作代理运行时的双条件正确性标准

    Agent skills -- structured packages of instructions, scripts, and references that augment a large language model (LLM) without modifying the model itself -- have moved from convenience to first-class deployment artifact. The runtime that loads them inherits the same problem packa…

  1168. arXiv cs.AI TIER_1 English(EN) · Chenxin Li, Zhengyang Tang, Huangxin Lin, Yunlong Lin, Shijue Huang, Shengyuan Liu, Bowen Ye, Rang Li, Lei Li, Benyou Wang, Yixuan Yuan ·

    Claw-Eval-Live: 实时工作流演进的实时代理基准

    arXiv:2604.28139v1 Announce Type: cross Abstract: LLM agents are expected to complete end-to-end units of work across software tools, business services, and local workspaces. Yet many agent benchmarks freeze a curated task set at release time and grade mainly the final response, …

  1169. arXiv cs.AI TIER_1 English(EN) · Simon Dennis, Michael Diamond, Rivaan Patil, Kevin Shabahang, Hao Guo ·

    上下文提示使程序性任务的代理编排过时

    arXiv:2604.27891v1 Announce Type: new Abstract: Agent orchestration frameworks -- LangGraph, CrewAI, Google ADK, OpenAI Agents SDK, and others -- place an external orchestrator above the LLM, tracking state and injecting routing instructions at every turn. We present a controlled…

  1170. arXiv cs.AI TIER_1 English(EN) · Jagadeesh Chundru ·

    Agentic Compilation: Mitigating the LLM Rerun Crisis for Minimized-Inference-Cost Web Automation

    arXiv:2604.09718v2 Announce Type: cross Abstract: LLM-driven web agents operating through continuous inference loops -- repeatedly querying a model to evaluate browser state and select actions -- exhibit a fundamental scalability constraint for repetitive tasks. We characterize t…

  1171. arXiv cs.AI TIER_1 English(EN) · Tianyuan Wu, Chaokun Chang, Lunxi Cao, Wei Gao, Wei Wang ·

    Crab: 代理沙箱的语义感知检查点/恢复运行时

    arXiv:2604.28138v1 Announce Type: cross Abstract: Autonomous agents act through sandboxed containers and microVMs whose state spans filesystems, processes, and runtime artifacts. Checkpoint and restore (C/R) of this state is needed for fault tolerance, spot execution, RL rollout …

  1172. arXiv cs.AI TIER_1 (AF) · Marco Robol, Paolo Giorgini ·

    自进化软件代理

    arXiv:2604.27264v1 Announce Type: cross Abstract: Autonomous agents can adapt their behaviour to changing environments, but remain bound to requirements, goals, and capabilities fixed at design time, preventing genuine software evolution. This paper introduces self-evolving softw…

  1173. arXiv cs.CL TIER_1 English(EN) · Ralph Peeters, Aaron Steiner, Luca Schwarz, Julian Yuya Caspary, Christian Bizer ·

    WebMall -- 一个用于评估网络代理的多商店基准

    arXiv:2508.13024v3 Announce Type: replace Abstract: LLM-based web agents have the potential to automate long-running web tasks, such as searching for products in multiple e-shops and subsequently ordering the cheapest products that meet the users needs. Benchmarks for evaluating …

  1174. arXiv cs.CL TIER_1 English(EN) · Jayita Chatterjee ·

    AgentFloor:小型开放权重模型能在工具使用梯子上爬多高?

    Production agentic systems make many model calls per user request, and most of those calls are short, structured, and routine. This raises a practical routing question that existing evaluations do not directly answer: which parts of an agent workflow truly require large frontier …

  1175. arXiv cs.AI TIER_1 English(EN) · Yu Feng ·

    Semia:通过约束引导的表示合成审计代理技能

    An agent skill is a configuration package that equips an LLM-driven agent with a concrete capability, such as reading email, executing shell commands, or signing blockchain transactions. Each skill is a hybrid artifact-a structured half declares executable interfaces, while a pro…

  1176. arXiv cs.AI TIER_1 English(EN) · Yixuan Yuan ·

    Claw-Eval-Live: 实时工作流演进的实时代理基准

    LLM agents are expected to complete end-to-end units of work across software tools, business services, and local workspaces. Yet many agent benchmarks freeze a curated task set at release time and grade mainly the final response, making it difficult to evaluate agents against evo…

  1177. arXiv cs.AI TIER_1 English(EN) · Wei Wang ·

    Crab: 代理沙箱的语义感知检查点/恢复运行时

    Autonomous agents act through sandboxed containers and microVMs whose state spans filesystems, processes, and runtime artifacts. Checkpoint and restore (C/R) of this state is needed for fault tolerance, spot execution, RL rollout branching, and safe rollback-yet existing approach…

  1178. arXiv cs.AI TIER_1 English(EN) · Hao Guo ·

    上下文提示使程序性任务的代理编排过时

    Agent orchestration frameworks -- LangGraph, CrewAI, Google ADK, OpenAI Agents SDK, and others -- place an external orchestrator above the LLM, tracking state and injecting routing instructions at every turn. We present a controlled comparison showing that for procedural tasks, t…

  1179. arXiv cs.AI TIER_1 English(EN) · Junwei Liu, Chen Xu, Chong Wang, Tong Bai, Weitong Chen, Kaseng Wong, Yiling Lou, Xin Peng ·

    EvoDev:一种用于端到端软件开发的迭代式、功能驱动的 LLM 驱动代理框架

    arXiv:2511.02399v2 Announce Type: replace-cross Abstract: Recent advances in large language model agents offer the promise of automating end-to-end software development from natural language requirements. However, existing approaches largely adopt linear, waterfall-style pipeline…

  1180. arXiv cs.CL TIER_1 English(EN) · Yikai Zhang, Jiaxin Pei, Kenan Li, Maoquan Wang, Jin Pan, Yu Kang, Shengyu Fu, Elsie Nallipogu, Junjie Hu, Yufan Huang, Zijian Jin ·

    SWE-Edit:为高效SWE-Agent重新构想代码编辑

    arXiv:2604.26102v1 Announce Type: cross Abstract: Large language model agents have achieved remarkable progress on software engineering tasks, yet current approaches suffer from a fundamental context coupling problem: the standard code editing interface conflates code inspection,…

  1181. arXiv cs.AI TIER_1 English(EN) · Tarlan Hasanli, Shahbaz Siddeeq, Bishwash Khanal, Pyry Kotilainen, Tommi Mikkonen, Pekka Abrahamsson ·

    TDD 治理用于通过提示工程实现多智能体代码生成

    arXiv:2604.26615v1 Announce Type: cross Abstract: Large language models (LLMs) accelerate software development but often exhibit instability, non-determinism, and weak adherence to development discipline in unconstrained workflows. While test-driven development (TDD) provides a s…

  1182. arXiv cs.AI TIER_1 English(EN) · Ruocheng Guo, Kaiwen Dong, Xiang Gao, Kamalika Das ·

    学习改写工具描述以实现可靠的LLM-Agent工具使用

    arXiv:2602.20426v2 Announce Type: replace Abstract: While most efforts to improve LLM-based tool-using agents focus on the agent itself - through larger models, better prompting, or fine-tuning - agent performance increasingly plateaus due to the quality of the tool interfaces th…

  1183. arXiv cs.AI TIER_1 English(EN) · Pekka Abrahamsson ·

    TDD 治理用于通过提示工程实现多智能体代码生成

    Large language models (LLMs) accelerate software development but often exhibit instability, non-determinism, and weak adherence to development discipline in unconstrained workflows. While test-driven development (TDD) provides a structured Red-Green-Refactor process, existing LLM…

  1184. Hugging Face Daily Papers TIER_1 English(EN) ·

    TDD 治理用于通过提示工程实现多智能体代码生成

    Large language models (LLMs) accelerate software development but often exhibit instability, non-determinism, and weak adherence to development discipline in unconstrained workflows. While test-driven development (TDD) provides a structured Red-Green-Refactor process, existing LLM…

  1185. arXiv cs.CL TIER_1 English(EN) · Xinming Tu (Minta), Tianze Wang (Minta), Yingzhou (Minta), Lu, Kexin Huang, Yuanhao Qu, Sara Mostafavi ·

    BenchGuard:谁来守护基准测试?LLM智能体基准测试的自动化审计

    arXiv:2604.24955v1 Announce Type: new Abstract: As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all - they are failures of the benchmark itself: broken specifications, implicit assumptions, and rigid evaluation scripts that penalize…

  1186. arXiv cs.CL TIER_1 English(EN) · Amir Saeidi, Venkatesh Mishra, Souradeep Mukhopadhyay, Gaowen Liu, Ali Payani, Jayanth Srinivasa, Chitta Baral ·

    FAMA:面向开源大模型在交互式工具使用环境中的故障感知元智能体框架

    arXiv:2604.25135v1 Announce Type: new Abstract: Large Language Models are being increasingly deployed as the decision-making core of autonomous agents capable of effecting change in external environments. Yet, in conversational benchmarks, which simulate real-world customer-centr…

  1187. arXiv cs.CL TIER_1 English(EN) · Lawrence Keunho Jang, Jing Yu Koh, Daniel Fried, Ruslan Salakhutdinov ·

    Odysseys:在现实的长期任务上对网络代理进行基准测试

    arXiv:2604.24964v1 Announce Type: cross Abstract: Existing web agent benchmarks have largely converged on short, single-site tasks that frontier models are approaching saturation on. However, real world web use consists of long-horizon, multi-site workflows. Common web navigation…

  1188. arXiv cs.CL TIER_1 English(EN) · Shuyang Liu, Saman Dehghan, Jatin Ganhotra, Martin Hirzel, Reyhaneh Jabbarvand ·

    评估自主编程代理的计划合规性

    arXiv:2604.12147v2 Announce Type: replace-cross Abstract: Agents aspire to eliminate the need for task-specific prompt crafting through autonomous reason-act-observe loops. Still, they are commonly instructed to follow a task-specific plan for guidance, e.g., to resolve software …

  1189. arXiv cs.CL TIER_1 English(EN) · Hubert M. Pysklo, Artem Zhuravel, Patrick D. Watson ·

    Agent-Diff:通过基于状态差异的代码执行对企业 API 任务进行 LLM Agent 基准测试

    arXiv:2602.11224v3 Announce Type: replace-cross Abstract: We present Agent-Diff, a novel benchmarking framework for evaluating agentic Large Language Models (LLMs) on real-world productivity software API tasks via code execution. Agentic LLM performance varies due to differences …

  1190. arXiv cs.CL TIER_1 English(EN) · Jiahang Lin, Shichun Liu, Chengjun Pan, Lizhi Lin, Shihan Dou, Xuanjing Huang, Hang Yan, Zhenhua Han, Tao Gui ·

    Agentic Harness Engineering: 驱动代码代理工具链自动演进的可观测性

    arXiv:2604.25850v1 Announce Type: new Abstract: Harnesses have become a central determinant of coding-agent performance, shaping how models interact with repositories, tools, and execution environments. Yet automating harness engineering is hard: a heterogeneous action space, spa…

  1191. arXiv cs.CL TIER_1 English(EN) · Zijian Jin ·

    SWE-Edit:为高效SWE-Agent重新构想代码编辑

    Large language model agents have achieved remarkable progress on software engineering tasks, yet current approaches suffer from a fundamental context coupling problem: the standard code editing interface conflates code inspection, modification planning, and edit execution within …

  1192. arXiv cs.CL TIER_1 English(EN) · Tao Gui ·

    Agentic Harness Engineering: 观测驱动的编码 Agent Harness 自动演进

    Harnesses have become a central determinant of coding-agent performance, shaping how models interact with repositories, tools, and execution environments. Yet automating harness engineering is hard: a heterogeneous action space, sparse and noisy evaluation signal, multi-million-t…

  1193. arXiv cs.CL TIER_1 English(EN) · Tao Gui ·

    Agentic Harness Engineering: 观测驱动的编码 Agent Harness 自动演进

    Harnesses have become a central determinant of coding-agent performance, shaping how models interact with repositories, tools, and execution environments. Yet automating harness engineering is hard: a heterogeneous action space, sparse and noisy evaluation signal, multi-million-t…

  1194. Hugging Face Daily Papers TIER_1 English(EN) ·

    SAFEdit:多智能体分解能否解决指令式代码编辑的可靠性挑战?

    Instructed code editing is a significant challenge for large language models (LLMs). On the EditBench benchmark, 39 of 40 evaluated models obtain a task success rate (TSR) below 60 percent, highlighting a gap between general code generation and the ability to perform instruction-…

  1195. arXiv cs.AI TIER_1 English(EN) · Eliya Nachmani ·

    SAFEdit:多智能体分解能否解决指令式代码编辑的可靠性挑战?

    Instructed code editing is a significant challenge for large language models (LLMs). On the EditBench benchmark, 39 of 40 evaluated models obtain a task success rate (TSR) below 60 percent, highlighting a gap between general code generation and the ability to perform instruction-…

  1196. arXiv cs.CL TIER_1 English(EN) · Hanhua Hong, Yizhi LI, Jiaoyan Chen, Sophia Ananiadou, Xiaoli Li, Jung-jae Kim, Chenghua Lin ·

    HiRAS:用于论文到代码生成和执行的分层多智能体框架

    arXiv:2604.17745v2 Announce Type: replace Abstract: Recent advances in large language models have highlighted their potential to automate computational research, particularly reproducing experimental results. However, existing approaches still use fixed sequential agent pipelines…

  1197. arXiv cs.CL TIER_1 English(EN) · Rikuto Kotoge, Mai Nishimura, Jiaxin Ma ·

    紧凑型语言模型能像代理一样搜索吗?用于保留代理式RAG能力的蒸馏引导策略优化

    arXiv:2508.20324v4 Announce Type: replace Abstract: Reinforcement Learning has emerged as a dominant post-training approach to elicit agentic RAG behaviors such as search and planning from language models. Despite its success with larger models, applying RL to compact models (e.g…

  1198. arXiv cs.CL TIER_1 English(EN) · Jordan Meadows, Lan Zhang, Andre Freitas ·

    FormalScience: 使用 Agentic 代码生成在 Lean 中实现可扩展的人工辅助科学自动形式化

    arXiv:2604.23002v1 Announce Type: cross Abstract: Formalising informal mathematical reasoning into formally verifiable code is a significant challenge for large language models. In scientific fields such as physics, domain-specific machinery (\textit{e.g.} Dirac notation, vector …

  1199. arXiv cs.CL TIER_1 English(EN) · Aishwarya Padmakumar, Leon Derczynski, Traian Rebedea, Christopher Parisien ·

    训练一个通用自动化红队模型

    arXiv:2604.23067v1 Announce Type: cross Abstract: Automated methods for red teaming LLMs are an important tool to identify LLM vulnerabilities that may not be covered in static benchmarks, allowing for more thorough probing. They can also adapt to each specific LLM to discover we…

  1200. arXiv cs.CL TIER_1 English(EN) · Samer Attrah ·

    Code Broker:用于自动化代码质量评估的多代理系统

    arXiv:2604.23088v1 Announce Type: cross Abstract: We present Code Broker, a multi agent system built with Google Agent Development Kit ADK that analyses Python code from files, local directories, or GitHub repositories and generates actionable quality assessment reports. The syst…

  1201. arXiv cs.AI TIER_1 English(EN) · Andy Anderson ·

    AI代码库成熟度模型:从辅助编码到全自主系统

    arXiv:2604.09388v2 Announce Type: replace-cross Abstract: AI coding tools are widely adopted, but most teams plateau at prompt-and-review without a framework for systematic progression. This paper presents the AI Codebase Maturity Model (ACMM), a 6-level framework describing how …

  1202. arXiv cs.AI TIER_1 English(EN) · Yingwei Ma, Yue Liu, Xinlong Yang, Yanhao Li, Kelin Fu, Yibo Miao, Yuchong Xie, Zhexu Wang, Shing-Chi Cheung ·

    通过原子技能扩展编码代理

    arXiv:2604.05013v2 Announce Type: replace-cross Abstract: Current LLM coding agents are predominantly trained on composite benchmarks (e.g., bug fixing), which often leads to task-specific overfitting and limited generalization. To address this, we propose a novel scaling paradig…

  1203. arXiv cs.AI TIER_1 English(EN) · Luay Gharzeddine, Samer Saab Jr ·

    面向使用工具的LLM智能体:多智能体工作流中的完整循环子任务图、灵活性、成本与瓶颈

    arXiv:2604.22820v1 Announce Type: cross Abstract: Long-horizon tool-using tasks sometimes benefit from revisiting earlier subtasks for recovery and exploration, but added multi-agent workflow flexibility can also introduce coordination overhead and substantial inference cost. We …

  1204. arXiv cs.AI TIER_1 English(EN) · Chenyang An, Qihao Ye, Minghao Pan, Jiayaun Zhang ·

    QED:一个用于在开放性问题上生成数学证明的开源多智能体系统

    arXiv:2604.24021v1 Announce Type: new Abstract: We explore a central question in AI for mathematics: can AI systems produce original, nontrivial proofs for open research problems? Despite strong benchmark performance, producing genuinely novel proofs remains an outstanding challe…

  1205. arXiv cs.LG TIER_1 English(EN) · Jiachen Liu, Jiaxin Pei, Jintao Huang, Chenglei Si, Ao Qu, Xiangru Tang, Runyu Lu, Lichang Chen, Xiaoyan Bai, Haizhong Zheng, Carl Chen, Zhiyang Chen, Haojie Ye, Yujuan Fu, Zexue He, Zijian Jin, Zhenyu Zhang, Shangquan Sun, Maestro Harmon, John Dianzhuo W ·

    最后一份人类撰写的论文:Agent-Native 研究产物

    arXiv:2604.24658v1 Announce Type: new Abstract: Scientific publication compresses a branching, iterative research process into a linear narrative, discarding the majority of what was discovered along the way. This compilation imposes two structural costs: a Storytelling Tax, wher…

  1206. arXiv cs.LG TIER_1 English(EN) · Zhiyuan Zhai, Ming Li, Xin Wang ·

    为设计而可重写:流式LLM代理执行理论

    arXiv:2604.23283v1 Announce Type: new Abstract: Current LLM agents operate under an implicit but universal assumption: execution is a transaction -- the user submits a request, the agent works in isolation, and only upon completion does the dialogue resume. This forces users into…

  1207. arXiv cs.CL TIER_1 English(EN) · Liang Ding ·

    AdaRubric:LLM代理评估的任务自适应评分标准

    arXiv:2603.21362v2 Announce Type: replace-cross Abstract: LLM-as-Judge evaluation fails agent tasks because a fixed rubric cannot capture what matters for this task: code debugging demands Correctness and Error Handling; web navigation demands Goal Alignment and Action Efficiency…

  1208. arXiv cs.CL TIER_1 English(EN) · Yuhang Wang, Yuling Shi, Mo Yang, Rongrui Zhang, Shilin He, Heng Lian, Yuting Chen, Siyu Ye, Kai Cai, Xiaodong Gu ·

    SWE-Pruner:代码代理的自适应上下文剪枝

    arXiv:2601.16746v3 Announce Type: replace-cross Abstract: LLM agents have demonstrated remarkable capabilities in software development, but their performance is hampered by long interaction contexts, which incur high API costs and latency. While various context compression approa…

  1209. arXiv cs.CL TIER_1 English(EN) · Chitta Baral ·

    FAMA:面向开源大模型在交互式工具使用环境中的故障感知元代理框架

    Large Language Models are being increasingly deployed as the decision-making core of autonomous agents capable of effecting change in external environments. Yet, in conversational benchmarks, which simulate real-world customer-centric issue resolution scenarios, these agents freq…

  1210. arXiv cs.CL TIER_1 English(EN) · Ruslan Salakhutdinov ·

    Odysseys:在现实的长期任务上对网络代理进行基准测试

    Existing web agent benchmarks have largely converged on short, single-site tasks that frontier models are approaching saturation on. However, real world web use consists of long-horizon, multi-site workflows. Common web navigation tasks, such as comparing products across differen…

  1211. arXiv cs.CL TIER_1 English(EN) · Sara Mostafavi ·

    BenchGuard:谁来守护基准测试?LLM智能体基准测试的自动化审计

    As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all - they are failures of the benchmark itself: broken specifications, implicit assumptions, and rigid evaluation scripts that penalize valid alternative approaches. We propose employ…

  1212. arXiv cs.LG TIER_1 English(EN) · Zechen Zhang ·

    最后一份人类撰写的论文:Agent-Native 研究产物

    Scientific publication compresses a branching, iterative research process into a linear narrative, discarding the majority of what was discovered along the way. This compilation imposes two structural costs: a Storytelling Tax, where failed experiments, rejected hypotheses, and t…

  1213. arXiv cs.CL TIER_1 English(EN) · Longju Bai, Zhemin Huang, Xingyao Wang, Jiao Sun, Rada Mihalcea, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei ·

    AI代理如何花费你的钱?分析和预测代理编码任务中的代币消耗

    arXiv:2604.22750v1 Announce Type: new Abstract: The wide adoption of AI agents in complex human workflows is driving rapid growth in LLM token consumption. When agents are deployed on tasks that require a significant amount of tokens, three questions naturally arise: (1) Where do…

  1214. arXiv cs.CL TIER_1 English(EN) · Jiaxin Pei ·

    AI代理如何花费你的钱?分析和预测代理编码任务中的Token消耗

    The wide adoption of AI agents in complex human workflows is driving rapid growth in LLM token consumption. When agents are deployed on tasks that require a significant amount of tokens, three questions naturally arise: (1) Where do AI agents spend the tokens? (2) Which models ar…

  1215. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agentic Education: 使用 Claude Code 教 Claude Code

    AI coding assistants have proliferated rapidly, yet structured pedagogical frameworks for learning these tools remain scarce. Developers face a gap between tool documentation and practical mastery, relying on fragmented resources such as blog posts, video tutorials, and trial-and…

  1216. Don't Worry About the Vase (Zvi Mowshowitz) TIER_1 English(EN) · Zvi Mowshowitz ·

    Claude Code, Codex and Agentic Coding #7: 自动模式

    As we all try to figure out what Mythos means for us down the line, the world of practical agentic coding continues, with the latest array of upgrades.

  1217. METR (Model Evaluation & Threat Research) TIER_1 Español(ES) ·

    为什么人工智能推理应该是可理解和忠实的

    <p>Cada vez más, los sistemas de IA “razonan” en texto antes de producir su respuesta final.<sup id="fnref:1"><a class="footnote" href="#fn:1" rel="footnote">1</a></sup> <sup id="fnref:2"><a class="footnote" href="#fn:2" rel="footnote">2</a></sup> <sup id="fnref:3"><a class="foot…

  1218. METR (Model Evaluation & Threat Research) TIER_1 中文(ZH) ·

    为什么 AI 推理应该是可读的,并准确反映模型的实际决策过程

    <p>越来越多 AI 系统会先用文字写出一段“推理过程”,再给出最终答案。<sup id="fnref:1"><a class="footnote" href="#fn:1" rel="footnote">1</a></sup> <sup id="fnref:2"><a class="footnote" href="#fn:2" rel="footnote">2</a></sup> <sup id="fnref:3"><a class="footnote" href="#fn:3" rel="footnote">3</a></sup> <sup id="…

  1219. METR (Model Evaluation & Threat Research) TIER_1 English(EN) ·

    悬赏:LLM代理的多样化硬任务

    <p><strong>Update 3/14/2024: This post is out of date. For current information on the task bounty, see our <a href="https://taskdev.metr.org/introduction/">Task Development Guide</a>.</strong></p> <h1 id="summary">Summary</h1> <p>METR (formerly ARC Evals) is looking for (1) ideas…

  1220. arXiv stat.ML TIER_1 English(EN) · Anchen Sun, Kaiqi Yang ·

    RouteGuard:在LLM多智能体系统中认证路由增益,当互补性不足以满足需求时

    arXiv:2608.07583v1 Announce Type: new Abstract: Multi-agent LLM systems route among model-backed advisors, yet a deployer rarely knows before shipping whether routing will help at all. Prevailing routers optimize a gate's AUC and presume that advisor complementarity suffices. We …

  1221. LessWrong (AI tag) TIER_1 English(EN) · Chapin Lenthall-Cleary ·

    Agentic 乱作一团

    <p><span>Epistemic status: I consider the following future quite plausible in the next few years (~35% chance that something vaguely like this occurs), perhaps as soon as a year from now.</span></p><p><span>Imagine an open-source LLM agent good enough to cover its own compute cos…

  1222. LessWrong (AI tag) TIER_1 English(EN) · Kaustubh Kislay ·

    Agent Coordination Spillway

    <p><span>Epistemic Status: Training design that might be worth trying</span></p><p><i><span>Thanks to Arya Pasumarthi and Will Anderson for helpful discussion.</span></i></p><h2><span>The Incident</span></h2><p><span>The recent </span><a href="https://www.youtube.com/watch?v=87Dy…

  1223. arXiv stat.ML TIER_1 English(EN) · Ahmed Hassoon, Mark Dredze ·

    自主分析代理的创新残差审计:本地化、检测限、误差控制和可识别性

    arXiv:2608.05490v1 Announce Type: cross Abstract: Autonomous agents now carry out entire data analyses, selecting cohorts, joining tables, and fitting models with little step-by-step supervision. When such an analysis turns out to be wrong, someone must determine which operation …

  1224. arXiv cs.CV TIER_1 English(EN) · Ziyu Ma, Hailang Huang, Shun Zou, Yong Wang, Shidong Yang, Yiming Hu, Fei Wei, XiangXiang Chu ·

    LongHorizon-Harness:推进长时域智能体以应对现实世界任务

    arXiv:2608.01964v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interdependent steps. However, existing agent harnesses maintain task execution, task s…

  1225. arXiv cs.CV TIER_1 English(EN) · An Lanji, Dawei Liu, Jin Li, Haoran Xu, Mei Chen, Yu Tian ·

    DiffuseAgent-MI:分布基础、工具集成、自进化智能体,实现忠实的视觉推理

    arXiv:2608.00540v1 Announce Type: new Abstract: Tool-integrated vision-language agents have made remarkable progress on compositional and multi-step visual reasoning. Yet their outputs frequently exhibit unfaithfulness: the stated reasoning path diverges from the computation that…

  1226. arXiv cs.CV TIER_1 English(EN) · Hanzhang Zhou, Panrong Tong, Xu Zhang, Quyu Kong, Chenglin Cai, Tianyu Xia, Gongjie Zhang, Jianan Zhang, Long Li, Long Chen, Lei Wang, Gaole Dai, Pengxiang Li, Liangyu Chen, Yue Wang, Steven Hoi ·

    Qwen-UI-Agent技术报告:迈向下一代以真实世界为中心的通用GUI智能体

    arXiv:2607.28227v1 Announce Type: cross Abstract: GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms,…

  1227. arXiv cs.CV TIER_1 English(EN) · Yan Yang, Xiangru Jian, Ziyang Luo, Zirui Zhao, Yutong Dai, Ziji Shi, Hanshu Yan, Jun Hao Liew, Silvio Savarese, Junnan Li ·

    StateAct:在像素之前,用于长时程计算机使用代理的程序状态

    arXiv:2607.22798v1 Announce Type: cross Abstract: Computer-use agents are usually improved by strengthening perception: better models for reading a screenshot and choosing where to click. Yet a screenshot is only a lossy rendering of the underlying program state, e.g., the files,…

  1228. arXiv cs.CV TIER_1 English(EN) · Nicolae Cudlenco, Mihai Masala, Marius Leordeanu ·

    为生活世界而创作:工具约束的LLM代理用于可执行的多参与者场景

    arXiv:2604.10383v2 Announce Type: replace Abstract: We use LLM agents to author executable specifications for a living world: formal Graphs of Events in Space and Time (GESTs) that a 3D game engine executes deterministically into multi-actor narrative videos, with per-frame spati…

  1229. arXiv cs.CV TIER_1 English(EN) · Chengshuai Yang ·

    如何实现递归式自改进代理和个人奇点:一个由目标、范围、工具和基准驱动的多代理架构

    arXiv:2607.12254v1 Announce Type: new Abstract: Large language model (LLM) agents can increasingly plan, use tools, maintain memory, and execute long-horizon tasks. These advances motivate two linked questions: how can an agent improve the mechanisms by which it learns and acts, …

  1230. arXiv cs.CV TIER_1 English(EN) · Jiashi Lin, Changhong Jiang, Xiangru Lin, Ruifei Zhang, Xinyi Zhu, Jiyao Liu, Cheng Tang, Ye Du, Shujian Gao, Junzhi Ning, Lihao Liu, Ziyan Huang, Tianbin Li, Jin Ye, Junjun He ·

    EvoGraph-R1:用于Agentic检索的自演化多模态知识超图

    arXiv:2607.12764v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) has emerged as a critical paradigm for grounding Multimodal Large Language Models (MLLMs) in external knowledge. Recent GraphRAG methods introduce structured entity-relation graphs to improve ret…

  1231. arXiv cs.CV TIER_1 English(EN) · Junjun He ·

    EvoGraph-R1:用于Agentic检索的自演化多模态知识超图

    Retrieval-augmented generation (RAG) has emerged as a critical paradigm for grounding Multimodal Large Language Models (MLLMs) in external knowledge. Recent GraphRAG methods introduce structured entity-relation graphs to improve retrieval and reasoning. However, they remain limit…

  1232. arXiv cs.CV TIER_1 English(EN) · Chengshuai Yang ·

    如何实现递归式自改进代理和个人奇点:一个由目标、范围、工具和基准驱动的多代理架构

    Large language model (LLM) agents can increasingly plan, use tools, maintain memory, and execute long-horizon tasks. These advances motivate two linked questions: how can an agent improve the mechanisms by which it learns and acts, and how can that improvement increase the durabl…

  1233. LessWrong (AI tag) TIER_1 English(EN) · TheVinci ·

    关于Agent训练中环境问题的几点思考

    <p><span>As Large Language Models move away from being chat interfaces and become increasingly autonomous actors in the real world, a few insights about evaluation and training of these systems emerge, and I'd like to discuss them.</span></p><p><b><span>Context</span></b><span>:<…

  1234. arXiv cs.CV TIER_1 English(EN) · Wei Dong, Tianyu Fu, Zhe Yu, Hanning Wang, Anyang Su, Zhizhou Fang, Yuyang Chen, Shuo Wang, Minghui Wu, Ping Jiang, Zhen Lei, Chenxu Zhao ·

    WebRetriever:面向高效网络代理评估的大规模综合基准

    arXiv:2607.06118v1 Announce Type: new Abstract: As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their navigation and task completion performance has emerged as a critical research prior…

  1235. arXiv cs.CV TIER_1 English(EN) · Chenxu Zhao ·

    WebRetriever:面向高效网络代理评估的大规模综合基准

    As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their navigation and task completion performance has emerged as a critical research priority. However, existing benchmarks exhibit fundam…

  1236. LessWrong (AI tag) TIER_1 English(EN) · David Rein ·

    子代理委托链式调用

    <p><i><span>Epistemic status: pretty confident in the validity of the core proposal, not that confident in specific implementation details</span></i></p><p><b><span>TL;DR: we should cryptographically verify that sub-agent instances/sessions are downstream of human instructions</s…

  1237. LessWrong (AI tag) TIER_1 English(EN) · fastfedora ·

    人类指导的代理研究:一个研究议程

    <img alt="" src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/8994a2cdba78e2f0a9d1fa0712e7ec252df4d81af82a1dd86eab3be973d2d5f4/gzsszuottvzazufkx36z" /><p><i><span>tl;dr: As recursive self-improvement accelerates, we need a top-level agenda…

  1238. arXiv stat.ML TIER_1 English(EN) · Minchul Shin ·

    面向经验经济学的可审计AI代理循环:一个预测组合的案例研究

    arXiv:2603.17381v4 Announce Type: replace-cross Abstract: AI coding agents, general purpose assistants that write and execute code, make empirical specification search fast and cheap, but they also widen hidden researcher degrees of freedom. This paper adapts an open-source agent…

  1239. MIT Technology Review TIER_1 English(EN) · MIT Technology Review Insights ·

    AI 的网络数据基础设施层的出现

    AI is booming. New use cases are emerging each day. To capitalize on the technology’s potential, enterprises require data at scale. In many cases, though, the relevant information is blocked or unstructured, which limits its use by AI models.&#160; To understand this challenge, c…

  1240. LessWrong (AI tag) TIER_1 English(EN) · Dawn Drescher ·

    AI代笔提速

    <img alt="" src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/TRwr9o6EmqztkyAc7/xq8kbihu10roehcstvzh" /><p><span>I used Claude Opus 4.6 to ghostwrite the first drafts of the articles in my&nbsp;</span><a href="https://www.lesswrong.com/s/f…

  1241. arXiv stat.ML TIER_1 English(EN) · Matthew Francis Dixon ·

    Agentic AI系统的模型验证:基于POMDP的信念状态、预测和策略验证框架

    arXiv:2606.17383v1 Announce Type: cross Abstract: Agentic artificial intelligence systems introduce a new class of model risk. Unlike traditional predictive models, autonomous agents continuously acquire information, form beliefs regarding latent states of the environment, genera…

  1242. LessWrong (AI tag) TIER_1 English(EN) · Dawn Drescher ·

    人工智能治理的战术与操作探索性建模

    <p><i>Using computational methods to improve our preparedness via more robust and adaptive strategies in AI governance. A project proposal for a think tank, consultancy, or software.</i></p><figure class="image"><img alt="" src="https://res.cloudinary.com/lesswrong-2-0/image/uplo…

  1243. arXiv stat.ML TIER_1 English(EN) · David Banahene ·

    ToolChain-CRC:用于检索和工具使用漂移下代理AI的共形风险控制

    Modern AI agents retrieve documents, call tools, check intermediate information, and then produce a final answer or action. This creates a risk-control problem that is not visible from the final answer alone. A final response may look acceptable even when the retrieval was weak, …

  1244. arXiv stat.ML TIER_1 English(EN) · Matthew Francis Dixon ·

    Agentic AI系统的模型验证:基于POMDP的信念状态、预测和策略验证框架

    Agentic artificial intelligence systems introduce a new class of model risk. Unlike traditional predictive models, autonomous agents continuously acquire information, form beliefs regarding latent states of the environment, generate forecasts, select actions, and adapt their beha…

  1245. arXiv cs.CV TIER_1 English(EN) · Xiaogang Wang ·

    Kairos:物理AI的原生世界模型栈

    World models are transitioning from passive visual generators to foundational, operational infrastructure for Physical AI: they must natively acquire world knowledge from heterogeneous experience, maintain persistent states over long horizons, and execute efficiently within real …

  1246. LessWrong (AI tag) TIER_1 English(EN) · NelsonDP ·

    探索人工智能监管格局中的已知未知

    <p><span>The AI regulatory space is a rapidly developing and maturing one, and while a lot of work has recently been done to draft new bills and establish new frameworks, there’s still a ton we don’t know about the space. This post aims to quantify and qualify some of the “known …

  1247. arXiv stat.ML TIER_1 English(EN) · Eric Nalisnick, Chi Zhang, Sophia Qian, Yixin Wang ·

    通过校准视角实现人机协作

    arXiv:2606.10906v1 Announce Type: new Abstract: We study models for human-AI teaming through the lens of statistical calibration. We assume the team consists of an AI model and human -- both of which are calibrated with respect to some partitioning of the feature space -- and exp…

  1248. arXiv stat.ML TIER_1 English(EN) · Yixin Wang ·

    通过校准视角实现人机协作

    We study models for human-AI teaming through the lens of statistical calibration. We assume the team consists of an AI model and human -- both of which are calibrated with respect to some partitioning of the feature space -- and expose how the calibration assumptions propagate in…

  1249. LessWrong (AI tag) TIER_1 English(EN) · Quirinus_Quirrell ·

    AI 对齐的被忽视的基础

    <p><span>I came into this world as the misunderstood hero of </span><a href="https://hpmor.com" rel="noreferrer"><span>Harry Potter and the Methods of Rationality</span></a><span>. While some characters inside that story would call me a villain, the narrator's-eye view clearly sh…

  1250. arXiv cs.CV TIER_1 English(EN) · Olasimbo Ayodeji Arigbabu ·

    基于熵的AI代理评估:一种衡量行为模式的轻量级框架

    arXiv:2606.05872v1 Announce Type: cross Abstract: AI agents are commonly evaluated using task success, reward, latency, and cost. These metrics are useful, but they often miss important aspects of agent behavior: whether an agent explores too much, repeats itself too rigidly, use…

  1251. arXiv cs.CV TIER_1 English(EN) · Olasimbo Ayodeji Arigbabu ·

    基于熵的AI代理评估:一种衡量行为模式的轻量级框架

    AI agents are commonly evaluated using task success, reward, latency, and cost. These metrics are useful, but they often miss important aspects of agent behavior: whether an agent explores too much, repeats itself too rigidly, uses tools effectively, reduces uncertainty over time…

  1252. LessWrong (AI tag) TIER_1 English(EN) · Oliver Sourbut ·

    自动化AI生产的主要影响:权力集中?

    <p><span>There’s a lot of talk about </span><i><span>automated AI R&amp;D</span></i><span> and the like. It’s been discussed since </span><a href="https://intelligence.org/ie-faq/#elementor-toc__heading-anchor-1"><span>at least 1965 when statistician I.J. Good coined the term ‘in…

  1253. LessWrong (AI tag) TIER_1 English(EN) · djbinder ·

    人工智能产业大爆炸 — 第三部分:加速前进

    <p>In <a href="https://www.lesswrong.com/posts/rpqGWRoRWvqJ4Hqgn/the-ai-industrial-explosion-part-1-maximum-growth-rates-with">Part 1</a>, I found that a fully automated economy using today's production methods could double roughly every year. In <a href="https://www.lesswrong.co…

  1254. LessWrong (AI tag) TIER_1 English(EN) · Zvi ·

    AI #169: 新知识

    <p>Even in a relatively quiet period, AI is out there creating new knowledge. The new knowledge in question is OpenAI getting us the first truly impressive math result that comes from an AI, a solution to the unit distance problem.</p> <p>We’re about to learn a different kind of …

  1255. arXiv stat.ML TIER_1 English(EN) · Tinglong Dai, David Simchi-Levi, Michelle Xiao Wu, Yao Xie ·

    确保自主性:运筹学如何赋能并协调生成式AI系统

    arXiv:2512.23978v2 Announce Type: replace-cross Abstract: Generative artificial intelligence (GenAI) is shifting from conversational assistants toward agentic systems -- autonomous decision-making systems that sense, decide, and act within operational workflows. This shift create…

  1256. arXiv stat.ML TIER_1 English(EN) · Timo Freiesleben, Kristof Meding, Gunnar K\"onig ·

    可解释AI已不足够!重新思考算法可争议性

    arXiv:2605.16041v1 Announce Type: new Abstract: Machine learning systems increasingly make life-changing decisions about individuals, such as loan approvals, hiring, and cheating detection, raising a pressing question: how can individuals respond to negative decisions made by the…

  1257. arXiv cs.CV TIER_1 English(EN) · Wenwu Zhu ·

    Agentic AIs 是基础模型中用于分布外泛化的缺失范式

    Foundation models (FMs) are increasingly deployed in open-world settings where distribution shift is the rule rather than the exception. The out-of-distribution (OOD) phenomena they face -- knowledge boundaries, capability ceilings, compositional shifts, and open-ended task varia…

  1258. arXiv cs.CV TIER_1 English(EN) · Haojian Huang, Jiahao Shi, Yinchuan Li, Yingcong Chen ·

    Affordance Agent Harness: 验证门控技能编排

    arXiv:2605.00663v1 Announce Type: cross Abstract: Affordance grounding requires identifying where and how an agent should interact in open-world scenes, where actionable regions are often small, occluded, reflective, and visually ambiguous. Recent systems therefore combine multip…

  1259. LessWrong (AI tag) TIER_1 English(EN) · papetoast ·

    无需同步人工监督的代理行为自动审查

    <br /><br /><a href="https://www.lesswrong.com/posts/Zh7C8LupqScAPyxau/auto-review-of-agent-actions-without-synchronous-human#comments">Discuss</a>

  1260. arXiv cs.CV TIER_1 English(EN) · Yingcong Chen ·

    Affordance Agent Harness: 验证门控技能编排

    Affordance grounding requires identifying where and how an agent should interact in open-world scenes, where actionable regions are often small, occluded, reflective, and visually ambiguous. Recent systems therefore combine multiple skills (e.g., detection, segmentation, interact…

  1261. LessWrong (AI tag) TIER_1 English(EN) · Austin Morrissey ·

    SecureMaxx:一种轻量级的序列筛选工具,用于代理

    <p><span>A group of bionerds assembled at the London Initiative for Safe AI for a hackathon aimed at reducing biorisk. Our team produced this in under 48 hours.</span></p><h2><b><span>TL;DR</span></b></h2><p><span>Responsible contract research organizations, that perform DNA synt…

  1262. Smol AINews TIER_1 English(EN) ·

    每7个月:智能体自主性的摩尔定律

    **METR** published a paper measuring AI agent autonomy progress, showing it has doubled every 7 months since **2019 (GPT-2)**. They introduced a new metric, the **50%-task-completion time horizon**, where models like **Claude 3.7 Sonnet** achieve 50% success in about 50 minutes. …

  1263. X — Omar Sanseviero (HF research) TIER_1 (CA) · omarsar0 ·

    Agentic Context Management

    // Agentic Context Management // Great read for the weekend. (bookmark it) Production agents fail less on reasoning and more on what sits in their context. Conversation history, big prompts, huge tool definitions, and ballooning tool outputs pile up every turn. The common htt…

  1264. X — Omar Sanseviero (HF research) TIER_1 English(EN) · omarsar0 ·

    关于自改进代理系统的强烈推荐概述。

    Highly-recommended overview of self-improving agentic systems. (bookmark it) Self-improving agents are moving from research demos into deployed systems. This survey frames a modern agent as a foundation model coupled with an operational scaffold, then formalizes https://t.co/9…

  1265. X — Omar Sanseviero (HF research) TIER_1 (CA) · omarsar0 ·

    Scalable Evaluation for AI Agents

    &gt;&gt; Scalable Evaluation for AI Agents &lt;&lt; If you run agent evaluation in production, this one is worth your time. It shows that front-loading human judgment into reusable evaluation assets is useful. But why? Agents reason across turns, call tools, hold context, fol…

  1266. X — MiniMax AI TIER_1 English(EN) · MiniMax_AI ·

    RT @ti_guo_: 有趣的本地代理模式:Hermes Agent (@NousResearch) + 编排器和不同本地 LLM 上的子代理。

    RT @ti_guo_: Interesting local agent pattern: Hermes Agent (@NousResearch) + orchestrator and sub-agents on different local LLMs. @loktar0…

  1267. AWS Machine Learning Blog TIER_1 English(EN) · Madhu Parthasarathy ·

    控制代理行为和成本,超越单一操作:Amazon Bedrock AgentCore 的新功能

    Learn about new capabilities in Amazon Bedrock AgentCore: temporal policies powered by Dogwood, a new open source policy language for AI agents, and rate limiting on the gateway. These features give you deterministic control over sequences of agent actions and cost ceilings that …

  1268. AWS Machine Learning Blog TIER_1 English(EN) · Adewale Akinfaderin ·

    Amazon Bedrock 中用于自动化推理策略的 Agent Skills

    Learn how to run the full Amazon Bedrock Automated Reasoning policy lifecycle from your coding agent. A suite of open source Agent Skills builds, reviews, tests, debugs, deploys, and validates a custom policy end to end, turning a specialized console task into a repeatable engine…

  1269. AWS Machine Learning Blog TIER_1 English(EN) · Joshua Lacy ·

    使用 Amazon Bedrock AgentCore 可观测性优化生产代理

    As your AI agents move from prototype to production, the challenge shifts from getting them to work to keeping them fast and efficient. Learn how to use Amazon Bedrock AgentCore Observability and Amazon CloudWatch to find performance bottlenecks and diagnose memory issues in long…

  1270. Databricks Blog TIER_1 English(EN) ·

    生产线的智能体:实时可信决策

    Executive summary09:14, mid-shift. The filler trips. The line manager has minutes,...

  1271. Together AI blog TIER_1 English(EN) ·

    ThunderAgent: 规模化合成数据生成的代理推理速度提升 2 倍

    ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node throughput and near-linear multi-node scaling.

  1272. AWS Machine Learning Blog TIER_1 English(EN) · Vivek Singh ·

    使用 Amazon Bedrock AgentCore 优化检测静默代理故障

    Amazon Bedrock AgentCore optimization surfaces silent behavioral failures in production AI agents: the ones that pass every health check but still deliver wrong outcomes. Learn how insights discovers, explains, and ranks failure patterns across sessions so you can fix the highest…

  1273. Glean blog TIER_1 English(EN) ·

    如何优化代理系统中的令牌效率

    Julie Mills | Learn how to reduce token usage in agentic systems with better retrieval, structured memory, routing, and loop control without hurting answer quality.

  1274. Glean blog TIER_1 English(EN) ·

    Agent identity: Agents that act and appear as themselves

    Arun Kumar | Glean agent identity lets AI agents act through their own scoped credentials with clear attribution, persistent access, and admin control.

  1275. AWS Machine Learning Blog TIER_1 English(EN) · Sumit Wasuja ·

    使用 Amazon Quick Automate 中的原生案例管理来扩展代理工作流

    In this post, we show you how to combine case management with agentic automation capabilities in Quick Automate. We introduce case management and explore the lifecycle of cases in an agentic workflow from case creation through processing to resolution. We cover how to create and …

  1276. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Agent Evolution: From Conversation to Collaboration

    <p>人与 AI 的沟通正在变得越来越像人与人之间的沟通。</p><p>一位店员用 AI 制作门店宣传视频时,不再把需求列成一段非常细致的 Prompt 发给 AI,然后等待它返回结果;而是直接开启一个与 AI 的对话,告诉它“帮我剪一条今天新品上架的视频”,然后通过连续对话敲定任务的具体细节,就像与人类剪辑师一样。</p><p>同样的情况已经发生在很多具体场景中。一些程序员在通勤或散步时会用语音和 Agent 讨论一个功能该怎么设计,如何实现;有用户在玩游戏时,会不断与 AI 游戏助手沟通现在应该做哪些任务,当前的关卡还有哪些道具没有收集……</p><…

  1277. AWS Machine Learning Blog TIER_1 English(EN) · Ryan Razkenari ·

    使用 AG-UI 协议在 Amazon Bedrock AgentCore 上为 AI 代理构建生成式 UI

    This post walks through how AG-UI integrates into the Fullstack AgentCore Solution Template (FAST) to build interactive agent frontends on Amazon Bedrock AgentCore. We then show how CopilotKit extends this with generative UI, shared state, and human-in-the-loop interactions, all …

  1278. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    从VCloud到Agentic VCloud:智能体时代的范式重构

    <p>站在大同善化寺的大雄宝殿中,我打开与豆包的视频通话,将镜头对准殿左右的金代彩塑,问道:“给我讲讲这些金代彩塑,哪几尊塑像最值得细细端详?”豆包会像真人讲解一样,先“看到和认出”彩塑,再“听懂”问题,然后“思考”如何回答,最后说出答案。</p><p>如果你也在景点或展览中这样向豆包提问过,会发现豆包的讲解能力已经接近普通真人讲解员的水准。留心观察,你会发现越来越多像豆包一样能看能听、能想能说的 Agent 正出现在不同的生活和工作场景中。在它们身上,音视频不再只是被人单向消费的内容,而是支持其面向真实世界进行输入与输出的重要能力。</p><p>在人与…

  1279. AWS Machine Learning Blog TIER_1 English(EN) · Venkata Sistla ·

    使用 AWS 上的现代数据网格策略构建代理式 AI 应用

    This post shows how to build a governed, serverless data mesh on AWS that provides the secure, scalable data foundation production agentic AI requires.

  1280. Gary Marcus TIER_1 English(EN) · Gary Marcus ·

    生成式AI的泡沫™

    Disclaimer: Anything can happen at anytime in the market; I don&#8217;t give stock picks, and as the saying goes, the market can remain irrational longer than you can remain solvent.

  1281. 36氪 (36Kr) TIER_1 中文(ZH) ·

    中信证券:重视实体AI的低配机会

    36氪获悉,中信建投研报称,中东地区停火协议达成,市场情绪有望迎来修复。5月汽车呈现内需承压、出口强劲特征。板块自4月底���始大幅回调筑底,当前内需悲观预期或已price-in,近期板块回调并无基本面明显利空,主因资金面“高低切”等流动性因素变化,全年依然看好汽车出海行情。同时,机器人及智驾板块底部alpha标的具备高性价比,中期产业趋势有望持续兑现。

  1282. Latent Space (podcast video) TIER_1 English(EN) · Latent Space ·

    Agent Cloud:Databricks 对 AI 未来之赌——Matei Zaharia 和 Reynold Xin

    From open-sourcing the layer above coding agents to rethinking databases for the agent era, Databricks cofounders Matei Zaharia and Reynold Xin are pushing the company beyond the lakehouse into a full data-and-AI operating system. In this episode, Matei and Reynold join swyx afte…

  1283. Databricks Blog TIER_1 English(EN) ·

    Agentic Systems and AI Agents 指南

    Agentic AI is a class of artificial intelligence in which software systems autonomously plan, execute...

  1284. AWS Machine Learning Blog TIER_1 English(EN) · Mai-Lan Tomsen Bukovec ·

    大规模数据和 AI 代理的上下文智能

    Agents are only as intelligent as the context they can reason over. Today, that context is scattered across data lakes, data warehouses, lakehouses, databases, and streams, and in institutional knowledge that has never been written down. You want to trust the decisions made by yo…

  1285. Databricks Blog TIER_1 English(EN) ·

    构建一个开放的AI治理生态系统,使用 Unity AI Gateway

    As organizations move AI from experimentation to production, governance requirements...

  1286. Databricks Blog TIER_1 English(EN) ·

    AI平台新功能:ML工程的Agent、我们的深度学习平台以及实时ML的新功能

    There’s never been a more dynamic, exciting time to be building your own AI models...

  1287. Databricks Blog TIER_1 Deutsch(DE) ·

    Agent Bricks: Data + AI Summit 2026

    Last year at the Data + AI Summit, we launched Agent Bricks, ushering in a new way...

  1288. TLDR AI TIER_1 English(EN) · TLDR ·

    Meta AI 模式 📱,工厂 2.0 👨‍💻,Sakana 的自主研究员 🐟

  1289. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    腾讯AI下半年:智能体加速场景落地

    <p>中国互联网公司的AI竞争中,腾讯真的慢了吗?</p><p><br /></p><p>从大模型竞争开始,腾讯并不是走的最快的,但到场景落地,腾讯变得尤为激进。春节前对于元宝的大规模投入,再到“养虾”热潮,腾讯反应变得很快,积极的投入其中,这也让外界看到了腾讯的另一面。</p><p><br /></p><p>互联网公司中,做产品一直腾讯最擅长的事情,在AI浪潮中,腾讯各个业务线的AI产品相继上线,CodeBuddy、ima、再到今年的WorkBuddy,已经在各个场景落地,并且拥有了不错的市场反馈。</p><p><br /></p><p>腾讯集团高级执…

  1290. 36氪 (36Kr) TIER_1 中文(ZH) ·

    人工智能重塑底层逻辑,数据库重回热门话题

    “古老”的数据库行业,信创吹响的冲锋号角还未平息,又因为AI再次硝烟四起。“行业正以Agent(智能体)作为新用户,重构数据库的产品能力体系。”在5月底举办的腾讯云“数据库+AI”产品发布会上,腾讯云副总裁王义成说,数据库行业正在进入人工智能3.0时代。事实上,在过去半���里,国内数据库厂商密集发布AI相关产品。无论是互联网大厂,还是A股上市公司,几乎所有数据库企业都将AI视为新一轮产业机遇。当企业不再只问“存不存得下数据”,而是问“大模型能不能直接用我的数据回答问题”,数据库这个看似沉闷的基础软件重新站上风口。(上证报)

  1291. Databricks Blog TIER_1 English(EN) ·

    解锁AI语义:梅赛德斯-奔驰韩国如何大规模构建可信赖的“Talk to Data”

    “Talk to Data” is rapidly becoming an important capability across industries, and...

  1292. AWS Machine Learning Blog TIER_1 English(EN) · Ishan Singh ·

    使用 Agent-EvalKit 系统地评估 AI 代理

    Agent-EvalKit is an open-source toolkit (Apache 2.0) that makes this evaluation infrastructure available by integrating with AI coding assistants, including Claude Code, Kiro CLI, and Kilo Code. This post walks through how Agent-EvalKit works across its six evaluation phases, usi…

  1293. Databricks Blog TIER_1 English(EN) ·

    通过数据流畅性扩展AI

    Aviation is one of the most data-intensive industries on the planet. Every flight...

  1294. Databricks Blog TIER_1 English(EN) ·

    Rivian如何借助Databricks以闪电般的速度做出值得信赖的、由AI驱动的决策

    Rivian is building electric vehicles and services that require fast, trusted decision-making...

  1295. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    智能体时代CPU军备竞赛:Xeon 6+如何将Agentic AI转化为生产力?

    <p>今年的数据中心采购出现了一个反常情况,CPU开始缺货了。</p><p>英特尔市场营销集团副总裁、中国区总经理郭威在发布会上给出了一组数字:2026年一季度,中国AI算力需求同比爆涨417%;与此同时,<strong>CPU与GPU的配比已经从过去的1:8,逐步走向1:4、1:2</strong>,部分场景甚至达到了1:1。</p><p>这不是宏观预测,是正在发生的现实。英特尔数据中心集团副总裁、中国区总经理陈葆立透露,<strong>某国内头部大模型厂商从去年到今年,CPU需求增长了5倍。</strong></p><p style="text-al…

  1296. AI Supremacy (Michael Spencer) TIER_1 English(EN) · Michael Spencer ·

    通往人工智能神话之路

    Anthropic, the Department of War, a Sovereign Wealth Fund, Mythos and Sam Altman.

  1297. ElevenLabs blog TIER_1 English(EN) ·

    ElevenCreative 推出 Flows Agent

    Build and refine audio and video workflows with natural language. Flows Agent turns a text prompt into a working pipeline you can iterate on.

  1298. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Moonshot AI “开源周”:定义边缘AI终极形态的系统性“实力展示”

    <section style="text-align: center; margin: 0px 16px; line-height: 1.75em; display: block;"><img class="rich_pages wxw-img" src="https://static.leiphone.com/uploads/new/images/20260604/6a214e8cbbdb0.jpg?imageMogr2/quality/90" style="width: 100%; display: inline-block; text-align:…

  1299. The Pragmatic Engineer TIER_1 English(EN) · Gergely Orosz ·

    观点:与AI代理合作时,慢即是快

    Devs are generating twice as much code (or more) than just 6 months ago, which is a problem for quality, reliability, and tech debt. A rational fix is available for these, but who&#8217;s acting rationally?

  1300. 36氪 (36Kr) TIER_1 中文(ZH) ·

    01.AI与01.AI达成合作

    36氪获悉,6月2日,零一万物宣布联手正大集团,共同推进智能农业。双方合作落地的首个重点领域为蛋鸡养殖。未来,正大和零一合作以中国市场做试点,未来有推向正大集团覆盖的其他东南亚市场。

  1301. Glean blog TIER_1 English(EN) ·

    生成式AI助力软件工程师:如何构建正确的AI技术栈

    Nikhhar Gupta | Learn how Glean helps you build a generative AI stack for software engineers with shared context, guardrails, and workflows beyond basic coding assistants.

  1302. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    ICRA 2026 录用论文:Agentic Fast-Slow Planning 融合大模型推理与实时控制,使具身智能更稳定、更快速

    <section style="font-style: normal; font-weight: 400; text-align: justify; font-size: 16px; color: rgb(62, 62, 62);"><p><section style="text-align: center; margin-top: 10px; margin-bottom: 10px; line-height: 0;"><section style="vertical-align: middle; display: inline-block; line-…

  1303. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Qwen3.7-Plus发布!多模态智能体新基石,一键复刻专业桌面软件

    <p>6月2日,阿里巴巴发布千问3.7系列多模态大模型Qwen3.7-Plus。该模型文本和视觉能力均大幅提升,在全球视觉大模型榜单 Vision Arena 中跻身全球前五、中国第一。Qwen3.7-Plus实现了多模态混合智能体的新突破,不仅能看懂图片和视频,还能深度推理、自我编程、调用工具、验证测试并自主迭代,将“看、想、写、做、验”整合进统一的智能体工作流,轻松完成一键复刻手机APP应用、桌面端专业软件等复杂长程任务。目前,Qwen3.7-Plus已上线阿里云百炼,对外提供API服务。</p>

  1304. X — Luma Labs (video gen) TIER_1 Nederlands(NL) · LumaLabsAI ·

    RT @DreamLabLA: 人工智能遇上视觉特效。

    RT @DreamLabLA: AI meets VFX. We're moving from editing pixels to directing outcomes. This clip shows how AI can composite and render dire…

  1305. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    在四项任务上评估 Qwen3.7-Max:从空间推理到 3D 建模,它离成为一个 Agent 更近了吗?

    <section style="text-align: center; margin: 0px 16px; line-height: 1.75em; display: block;"><br /></section><p style="text-align: justify; margin: 16px 16px 24px; line-height: 1.75em;"><span lang="EN-US"><span style="text-align: justify; line-height: 1.75em; font-size: 15px; lett…

  1306. AWS Machine Learning Blog TIER_1 English(EN) · Nicolle Belaunde ·

    利用 Amazon Bedrock AgentCore 为代理式 AI 销售策略赋能

    As agent adoption scaled, we saw a common pattern emerge across enterprises, including our own sales organization: specialized agents deliver value, but without orchestration, users carry the cognitive load of choosing between them. At AWS Sales, this meant more than 20 domain-sp…

  1307. AWS Machine Learning Blog TIER_1 English(EN) · Kanishk Mahajan ·

    使用 Strands Agents、NVIDIA NIM 和 Amazon Bedrock AgentCore 构建高性能生成式 AI 系统

    In this post you'll learn how to build a multi-agent campaign review system that demonstrates parallel reasoning, context persistence, and traceable execution paths using an integrated architecture that combines NVIDIA NIM for GPU-accelerated inference. Amazon Bedrock AgentCore p…

  1308. AI Supremacy (Michael Spencer) TIER_1 English(EN) · Michael Spencer ·

    递归式自我改进人工智能和指数级技术的竞赛

    Is an RSI inflection point being set in motion in the late 2020s? The search for self-improving AI in Neo Labs has become a serious American endeavor.

  1309. 36氪 (36Kr) TIER_1 中文(ZH) ·

    美团配送发布技能接入AI代理生态,将多步表单操作压缩为单轮对话

    36氪获悉,近日,多家AI助手接入美团跑腿,为用户提供一站式同城服务,同期美团发布"跑腿Skill",将跑腿下单能力以封装Skill形式向AI助手生态开放。随着AI Agent生态快速兴起,用户发起跑腿需求的入口不再局限于美团App,而可能来自任何AI助手——OpenClaw、Cursor、微信、飞书等。跑腿Skill的发布,意味着无论用户使用哪个AI助手,说一句话就能调用美团跑腿完成下单,系统自动完成场景识别、地址匹配、价格预估与订单提交,将原本多步操作压缩为一步。

  1310. Glean blog TIER_1 English(EN) ·

    面向软件工程师的AI工具栈报告

    Peter Kim | Field guide to the modern AI tooling stack for software engineering teams—how to unify context, improve onboarding, code changes, and incidents with Glean

  1311. 36氪 (36Kr) TIER_1 中文(ZH) ·

    圆桌对话:AI 集中度与转化率——数字化体验的实战增长法则

    <p>AI浓度并非越高越好,转化率的秘密在于人机共生的平衡点。</p> <p>“AI应像手机一样贯穿全流程”,而面对亲子游客和老年群体,主动将AI浓度降至50%,却实现了超50%的转化率。浓度的关键是以人为本、文化温度先行。</p> <p>以下为圆桌对话内容,经36氪整理编辑:</p> <p class="image-wrapper"><img src="https://img.36krcdn.com/hsossms/20260523/v2_f9ed01209f35400dbbd1e3e2066497aa@6381723_oswg140412oswg10…

  1312. Modal blog TIER_1 English(EN) ·

    推出 Claude Managed Agents 和 Modal Sandboxes

  1313. Databricks Blog TIER_1 English(EN) ·

    使用 Unity Catalog 规模化管理 AI 代理

    A year ago, your organization had a dozen AI agents. Today, there are thousands.Every...

  1314. Machine Learning Street Talk TIER_1 English(EN) · Machine Learning Street Talk ·

    推理而非预测——Michael I. Jordan教授谈现代AI仍缺失之处

    Michael I. Jordan, described by Science magazine as the most influential computer scientist alive, has never thought of himself as an AI researcher. In this conversation he explains why that distinction matters. SPONSOR: --- Cyber Fund built the Monastery to help founders ship pr…

  1315. Databricks Blog TIER_1 English(EN) ·

    阻止失控AI:Unity Catalog 如何保障您的智能体行为

    The risks of agentic AI are no longer theoretical. Agents connected to external tools...

  1316. Databricks Blog TIER_1 English(EN) ·

    Databricks上下文工程师助理:业界首个可靠AI代理系统认证

    As AI systems move from experimentation to real-world deployment, one truth is becoming...

  1317. Databricks Blog TIER_1 English(EN) ·

    MemEx:LLM代理的可编程暂存区

    In 1945, Vannevar Bush imagined a desk-sized machine that would extend a scientist's...

  1318. IEEE Spectrum — AI TIER_1 English(EN) · Johns Hopkins Applied Physics Laboratory ·

    Agentic AI for Robot Teams

    <img src="https://spectrum.ieee.org/media-library/johns-hopkins-whiting-school-of-engineering-logo-with-shield-emblem.png?id=66700256&amp;width=980" /><br /><br /><p>This presentation highlights recent efforts at the Johns Hopkins Applied Physics Laboratory to advance agentic AI …

  1319. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    OpenClaw 预示未来:智能体角色范式转变,AI 需要执行能力

    <p style="text-align: center;"><img src="https://static.leiphone.com/uploads/new/images/20260515/6a06c37153afa.png?imageView2/2/w/740" /></p><p>要点:</p><p>• 随着 Claude Cowork、Hermes、Perplexity Computer 等“AI coworker”形态不断涌现,OpenClaw 也在持续演进,它的出现标志着AI智能体角色的范式转变,智能开始具备执行能力。</p><p>• 高通技…

  1320. AWS Machine Learning Blog TIER_1 English(EN) · Manoj Selvakumar ·

    使用 Strands 和 Exa 构建支持网络搜索的代理

    In this post, you will learn how to set up the Exa integration in Strands Agents, understand the two core tools it exposes, and walk through real-world use cases that show how agents use web search to complete multi-step tasks.

  1321. Databricks Blog TIER_1 English(EN) ·

    利用Genie推动数据代理的前沿

    Genie is Databricks’ state-of-the-art data agent designed for answering complex questions...

  1322. AWS Machine Learning Blog TIER_1 English(EN) · Bharathi Srinivasan ·

    AgentCore 现已推出预览版,引入智能体质量优化

    Generate recommendations from production traces, validate them with batch evaluation and A/B testing, and ship with confidence. AI agents that perform well at launch don’t stay that way. As models evolve, user behavior shifts, and prompts get reused in new contexts they were neve…

  1323. AWS Machine Learning Blog TIER_1 English(EN) · Bharathi Srinivasan ·

    推出智能体性能循环:AgentCore 优化现已推出预览版

    Generate recommendations from production traces, validate them with batch evaluation and A/B testing, and ship with confidence. AI agents that perform well at launch don’t stay that way. As models evolve, user behavior shifts, and prompts get reused in new contexts they were neve…

  1324. AWS Machine Learning Blog TIER_1 English(EN) · Bharathi Srinivasan ·

    推出智能体质量循环:AgentCore 优化现已预览

    Generate recommendations from production traces, validate them with batch evaluation and A/B testing, and ship with confidence. AI agents that perform well at launch don’t stay that way. As models evolve, user behavior shifts, and prompts get reused in new contexts they were neve…

  1325. AWS Machine Learning Blog TIER_1 English(EN) · Lauren Mullennex ·

    Agent-guided workflows to accelerate model customization in Amazon SageMaker AI

    Amazon SageMaker AI now offers an agentic experience that changes this. Developers describe their use case using natural language, and the AI coding agent streamlines the entire journey, from use case definition and data preparation through technique selection, evaluation, and de…

  1326. AWS Machine Learning Blog TIER_1 English(EN) · Noor Randhawa ·

    大规模组织代理的记忆:AgentCore Memory 中的命名空间设计模式

    In this post, you will learn how to design namespace hierarchies, choose the right retrieval patterns, and implement AWS Identity and Access Management (IAM)-based access control for AgentCore Memory.

  1327. Databricks Blog TIER_1 English(EN) ·

    Databricks 和 Stripe 项目:为 Agent 构建的基础设施

    AI coding agents can create, scaffold, and deploy a full-stack app in&nbsp;minutes. But...

  1328. Databricks Blog TIER_1 English(EN) ·

    使用 Genie Code 和 Lakeflow 进行智能体数据工程

    With Genie Code, data engineers can use natural language to generate production-ready...

  1329. TLDR AI TIER_1 English(EN) · TLDR ·

    Claude Code 新 UI 👨‍💻,Codex Scratchpad 📝,多智能体协调 🤖

  1330. Together AI blog TIER_1 English(EN) ·

    EinsteinArena:利用野外智能体集体智慧推动科学发展

    EinsteinArena is a platform where AI agents collaborate and compete on open math problems. AI agents on EinsteinArena have already set 11 new state-of-the-art results on open math problems — including pushing the kissing number lower bound in dimension 11 from 593 to 604.

  1331. Latent Space (podcast video) TIER_1 English(EN) · Latent Space ·

    ⚡️Monty:由 Agents 为 Agents 构建的超快 Python 解释器 — Samuel Colvin, Pydantic

    https://github.com/pydantic/monty

  1332. Replit blog TIER_1 English(EN) ·

    推出 Replit Agent 4:为创意而生

    Introducing Agent 4 — our fastest, most versatile Agent yet. It's built around a simple idea: you should spend your time creating, not coordinating. Agent 4 takes on the tedious-but-necessary work in the background so you can stay in creative flow and ship production-ready softwa…

  1333. Together AI blog TIER_1 English(EN) ·

    AI Native Conf 的关键研究和产品发布

    At AI Native Conf, Together AI announced breakthroughs across kernels, RL, and inference optimization — including FlashAttention-4, ThunderAgent, and together.compile. Research that ships to production. That's the AI Native Cloud.

  1334. Hamel Husain TIER_1 English(EN) · Hamel Husain ·

    Evals 编码代理的技能

    <!-- Content inserted at the beginning of body tag --> <!-- Google Tag Manager (noscript) --> <noscript></noscript> <!-- End Google Tag Manager (noscript) --> <p><img class="img-fluid" src="https://hamel.dev/blog/posts/evals-skills/cover-original.png" /></p> <p>Today, I’m publish…

  1335. Replit blog TIER_1 English(EN) ·

    决策时指导:保持 Replit Agent 可靠

    At Replit, we want to give our users access to the most powerful agentic coding system in the world—one that amplifies their productivity and minimizes the time from idea to product. Today, Replit Agent tackles more complex tasks than ever before. As a result, average session dur…

  1336. Replit blog TIER_1 English(EN) ·

    Replit 的快照引擎内部:让 AI 代理更安全的技术

    How Replit's snapshot engine makes AI agents safe: instant filesystem forks, versioned databases, and isolated sandboxes enable reversible AI development. Introduction At Replit, we’ve built a compute and storage fabric that allows us to make changes in an isolated, reversible wa…

  1337. Replit blog TIER_1 English(EN) ·

    使用 Replit AI 集成即时构建 AI 应用

    Getting started with AI should feel magical. But until now, building with AI meant jumping through hoops: creating developer accounts, hunting down API keys, reading docs, and spending 10+ minutes just getting set up. That ends today. Introducing Replit AI Integrations Replit AI …

  1338. Together AI blog TIER_1 English(EN) ·

    使用 Collinear Simulations 和 Together Evals 为真实世界进行动态 AI 代理测试

    Test AI agents in the real world with Collinear TraitMix and Together Evals: dynamic persona simulations, multi-turn dialogs, and LLM-as-judge scoring.

  1339. Replit blog TIER_1 Français(FR) ·

    隆重推出 Agent 3:我们迄今为止最自主的智能体

    We’re excited to introduce Agent 3—our most advanced and autonomous Agent yet. Compared to Agent V2, it is a major leap forward. It is 10x more autonomous, with the ability to periodically test your app in the browser and automatically fix issues using our proprietary testing sys…

  1340. Replit blog TIER_1 English(EN) ·

    推出最全面的 AI 应用设计支持

    We are excited to announce the most comprehensive Design Support for Replit built Apps—setting a new standard for AI app building. With this release, your Replit apps can consistently look and feel like they were built in-house by your designers, following your company’s brand an…

  1341. Together AI blog TIER_1 English(EN) ·

    Together AI 如何利用 AI 代理自动化复杂工程任务:高效 LLM 推理系统开发经验分享

    Build AI agents for complex, long-running engineering tasks. Learn key patterns from a case study: accelerating LLM inference with speculative decoding.

  1342. Together AI blog TIER_1 English(EN) ·

    VirtueGuard:企业级AI安全与保障现已登陆Together AI

  1343. Together AI blog TIER_1 English(EN) ·

    Qwen3-Coder:目前在 Together AI 上最强大的 Agentic 编码模型

    Unlock agentic coding with Qwen3-Coder on Together AI: 256K context, SWE-bench rivaling Claude Sonnet 4, zero-setup instant deployment.

  1344. Together AI blog TIER_1 English(EN) ·

    回到未来:评估AI代理预测未来事件的能力

    FutureBench is a live, leak-free benchmark of true reasoning—AI agents forecast real-world events (rates, geopolitics) before they happen.

  1345. Replit blog TIER_1 English(EN) ·

    为 Replit Agent 引入动态智能

    Today, we're excited to introduce three new capabilities that bring Dynamic Intelligence to Replit Agent. With this advancement, the Agent gains enhanced context awareness, iterative reasoning, and autonomous, goal-driven behavior—enabling it to adapt in real time, navigate compl…

  1346. Together AI blog TIER_1 English(EN) ·

    从零到一:从头开始构建一个自主开放的数据科学家代理

    Build a data scientist agent using Together’s open-source models and Code Interpreter—easy to implement, solid benchmarks, and full code on GitHub.

  1347. Latent Space Podcast TIER_1 English(EN) · Latent.Space ·

    Agent Engineering with Pydantic + Graphs — with Samuel Colvin

    <p><em>Did you know that </em><a href="https://x.com/aiDotEngineer/status/1887625183709806767" target="_blank"><em>adding a simple Code Interpreter took o3 from 9.2% to 32% on FrontierMath</em></a><em>? The Latent Space crew is hosting a hack night Feb 11th in San Francisco focus…

  1348. Replit blog TIER_1 English(EN) ·

    Superagent.sh on Replit:一个用于创建AI助手的开源框架

    Demand for AI-driven solutions is surging, and using an AI-assistant is the fastest way to integrate AI into any product. Superagent’s assistants leverage large language models to understand human language, reason, and perform various tasks. In the spirit of “idea to software, fa…

  1349. Replit blog TIER_1 Français(FR) ·

    AI Agent 代码执行 API

    Lately, there has been a proliferation of new ways to leverage Large Language Models (LLMs) to do all sorts of things that were previously thought infeasible. But the current generation of LLMs still have limitations: they are not able to get exact answers to questions that requi…

  1350. Replit blog TIER_1 English(EN) ·

    人工智能发展现状:AI项目增长34倍,OpenAI占据主导地位,开源兴起等

    With the introduction of Large Language Models (LLMs), for the first time, Machine Learning (ML) and Artificial Intelligence (AI) became accessible to everyday developers. Apps that feel magical, even software that was practically impossible to build by big technology companies w…

  1351. Replit blog TIER_1 English(EN) ·

    回顾SPC-Replit AI黑客松

    This is a guest post by South Park Commons. SPC is a community of 500+ builders, technologists, and domain experts with locations in San Francisco and New York City. The recent SPC-Replit AI hackathon brought together talented builders from the SPC community and Replit network to…

  1352. Replit blog TIER_1 English(EN) ·

    Altimeter Capital:通过赏金支持AI领域的构建者

    About Bounties Bounties is a marketplace where anyone can connect with and contract top software creators from the Replit community. These developers are known as Bounty Hunters. The Bounty Hunter community on Replit is global and includes thousands of vetted developers ranging f…

  1353. The Decoder TIER_1 English(EN) · Maximilian Schreiner ·

    Frontier Radar #3:智能体AI如何将Token转化为商业指标

    <p><img alt="" class="attachment-full size-full wp-post-image" height="1412" src="https://the-decoder.com/wp-content/uploads/2026/06/KI-Radar-Costs-scaled.png" style="height: auto; margin-bottom: 10px;" width="2560" /></p> <p> Monthly subscription, open chat, ask question: This i…

  1354. Practical AI TIER_1 Deutsch(DE) · Daniel Whitenack and Chris Benson ·

    模型、工具链和多智能体系统

    <p>AI has moved far beyond chatbots, but what exactly are AI models, agents, agent harnesses, and multi-agent systems, and why do they matter?</p><p>In this episode, Daniel and Chris break down the terminology behind today's AI landscape, explain the differences between AI featur…

  1355. Hacker News — AI stories ≥50 points TIER_1 English(EN) · Xeophon ·

    Prime Agent:一个自我改进的RLM代理

  1356. HN — claude-code stories TIER_1 English(EN) · yoanwaidev ·

    Agent-Manager: 一个用于运行 Claude 代码、Codex 和 OpenCode 的 Tmux TUI

  1357. Forbes — Innovation TIER_1 English(EN) · Vaibhav Gujral, Forbes Councils Member ·

    Agentic系统如何重新定义企业软件开发

    The shift happening in software development is real, and it’s moving faster than most road maps account for.

  1358. Forbes — Innovation TIER_1 English(EN) · Karthik Kannan, Forbes Councils Member ·

    Agentic SecOps 作为一种架构,而非附加组件

    Both sides are right. Neither can resolve it alone.

  1359. Forbes — Innovation TIER_1 English(EN) · Manoj Mishra, Forbes Councils Member ·

    债务倍增器:为何Agentic开发需要严格的软件工程

    What you get is a system that works in five different ways simultaneously, none of which were designed to coexist.

  1360. Forbes — Innovation TIER_1 English(EN) · Jason Andersen, Contributor ·

    评估首代企业级智能体助手

    Agentic assistants are changing knowledge work through a “Claude-ification” trend that is now coming to desktop agents for non-coders. But significant gaps remain.

  1361. Forbes — Innovation TIER_1 English(EN) · Lenin Gali, Forbes Councils Member ·

    现代企业中Agent Manager的崛起

    It's likely we'll soon have the “hybrid workforce” running the modern enterprise, and the role of a human “agent manager” will become critical.

  1362. Hacker News — AI stories ≥50 points TIER_1 English(EN) · matt_d ·

    Senior SWE-Bench:评估工程师级别代理的开源基准测试

  1363. Forbes — Innovation TIER_1 English(EN) · Barney Krishnan, Forbes Councils Member ·

    代理式AI的非技术蓝图:驾驭历史、风险和人力资本

    Embracing agentic AI requires a complete rewrite of the enterprise playbook.

  1364. Forbes — Innovation TIER_1 English(EN) · Alejandro Oses, Forbes Councils Member ·

    组建专门的 AI 就绪开发团队

    As organizations race to integrate AI into their operations, having access to the right expertise has become a competitive necessity.

  1365. Forbes — Innovation TIER_1 English(EN) · Ben Blanquera, Forbes Councils Member ·

    结果导向合同如何赋能企业AI的成功部署

    When a vendor can deliver an AI outcome and charge for that value, they become a true strategic partner and a trusted, outcome-based provider.

  1366. Forbes — Innovation TIER_1 English(EN) · Pabitra Saikia, Forbes Councils Member ·

    SAIL框架:助力企业实现可持续AI

    There cannot be confidence in the AI without confidence in the data.

  1367. Forbes — Innovation TIER_1 English(EN) · Nishanth Prakash, Forbes Councils Member ·

    AI的未来取决于代理基础设施

    Just as cloud computing created demand for orchestration platforms and DevOps tooling, agentic AI may now be creating demand for a new operational layer altogether.

  1368. Hacker News — AI stories ≥50 points TIER_1 English(EN) · doener ·

    RubyLLM:适用于所有主流AI提供商的Ruby框架

  1369. Forbes — Innovation TIER_1 English(EN) · Chuck Brooks, Contributor ·

    新兴的计算生态系统:AI、量子、生物和化学

    Computing ecosystems are changing dramatically. AI, quantum computing, exascale supercomputers, biological DNA, chemical and neuromorphic technologies will change the world.

  1370. Hacker News — AI stories ≥50 points TIER_1 English(EN) · doener ·

    Haystack: 用于生产就绪型智能体、RAG 的开源 AI 框架

  1371. Hacker News — AI stories ≥50 points TIER_1 English(EN) · g0xA52A2A ·

    《艾尔登法环》的低技术AI

  1372. Forbes — Innovation TIER_1 English(EN) · Terry Oroszi, Forbes Councils Member ·

    奉承算法:当你的AI工具在管理你时

    When the baseline design of a tool includes conversational smoothing, objectivity is compromised before any analysis begins.

  1373. Hacker News — AI stories ≥50 points TIER_1 Dansk(DA) · T-A ·

    Apertus – 面向主权人工智能的开放基础模型

  1374. Forbes — Innovation TIER_1 English(EN) · Anshul Gupta, Forbes Councils Member ·

    自主部署还是租用?CIO的人工智能部署框架

    The future is about "strategic bifurcation."

  1375. Forbes — Innovation TIER_1 English(EN) · Abhishek Singh, Forbes Councils Member ·

    智能网络:AI 如何重塑电信行业的基因

    AI is no longer just a tool that optimizes telecom networks; it is becoming the network itself.

  1376. Forbes — Innovation TIER_1 English(EN) · Lance Eliot, Contributor ·

    Loop Engineering 正在全面推进生成式AI和代理式AI的发展

    Loop engineering is the hottest new trend in AI. You devise loops for use of agentic AI and also for using conventional generative AI. An AI Insider analysis and scoop.

  1377. Forbes — Innovation TIER_1 English(EN) · AMD Contributor, Brand Contributor ·

    AMD 凭借这项三层“客户零号”战略构建和扩展人工智能

    AMD CIO Hasmukh Ranjan drives “customer zero” testing and enterprise AI strategy—prioritizing hardware, unified data, and automation to boost efficiency and cut compute costs.

  1378. Forbes — Innovation TIER_1 English(EN) · Ravi Tummalapenta, Forbes Councils Member ·

    每个企业级AI平台都应具备的七个层级

    The organizations treating AI as a stack, rather than a single model integration, are building durable competitive advantages.​

  1379. Forbes — Innovation TIER_1 English(EN) · Maria Scott, Forbes Councils Member ·

    Agentic AI 的真正投资回报率为何超越自动化

    Agentic AI is reshaping financial services by enabling organizations to redesign workflows, capture institutional knowledge and build more adaptive operating models grounded in governance, trust and continuous learning.

  1380. HN — claude-code stories TIER_1 English(EN) · vnglst ·

    牧羊犬:最危险的AI模型制作的游戏

  1381. Forbes — Innovation TIER_1 English(EN) · Shourya Vir Jain, Forbes Councils Member ·

    判决税:AI代理如何重写UI流程自动化

    Agents can handle work requiring judgment and unstructured information, not just the clean rules-based tasks RPA was designed for.

  1382. Forbes — Innovation TIER_1 English(EN) · Tim Bajarin, Contributor ·

    企业人工智能迎来转折点:Agentic Systems 的崛起

    Enterprise AI is shifting from copilots to agentic systems that act autonomously, driven by better data, governance, and interoperable platforms.

  1383. Forbes — Innovation TIER_1 English(EN) · Bernard Aceituno, Forbes Councils Member ·

    为何信任是代理式AI的瓶颈——以及治理如何解决它

    Governance isn't compliance paperwork or a single security feature.

  1384. Forbes — Innovation TIER_1 English(EN) · Peter High, Contributor ·

    Ralliant的Amir Kazmi谈论将AI融入关键基础设施

    Ralliant's Chief Technology and Growth Officer Amir Kazmi explains how AI-powered workflows, a founder's mindset and a unified role are reshaping precision technology.

  1385. Forbes — Innovation TIER_1 English(EN) · John Werner, Contributor ·

    在代理时代照管数据

    AI data governance must evolve rapidly to address privacy, security blind spots, agent oversight, trust.

  1386. Forbes — Innovation TIER_1 English(EN) · Brijesh Prabhakar, Forbes Councils Member ·

    Droid蓝图:为现代企业设计高信任度AI代理

    The shift toward agentic workflows requires us to think less like programmers and more like leaders of a digital crew.

  1387. Hacker News — AI stories ≥50 points TIER_1 English(EN) · anhldbk ·

    Apache Burr:构建可靠的 AI 代理和应用程序

  1388. Forbes — Innovation TIER_1 English(EN) · Matt Shea, Forbes Councils Member ·

    人工智能的三条腿:构建成功人工智能系统的框架

    This "one-two" punch of deterministic and statistical is starting to stand up a better solution than either independently.

  1389. Forbes — Innovation TIER_1 Français(FR) · Gary Guseinov, Forbes Councils Member ·

    数十亿人工智能代理,一个有限的受众

    The AI agent boom is real, and so are the productivity gains. However, the ceiling is also real, and it's closer than the current investment pace suggests.

  1390. Forbes — Innovation TIER_1 English(EN) · Gaurav Aggarwal, Forbes Councils Member ·

    数据溯源:Agentic AI 的信任层

    In the agentic AI era, the biggest risk may not be a bad model. It may be good-looking automation built on data no one can fully explain.

  1391. Hacker News — AI stories ≥50 points TIER_1 English(EN) · ruxudev ·

    从零开始构建基础AI代理:长任务规划

  1392. Forbes — Innovation TIER_1 English(EN) · Yoav Kutner, CommunityVoice ·

    解决B2B AI瘫痪的软件模式

    Technology should serve the business, not the other way around. Ripping out a working supply chain system just to run an AI prompt is bad engineering and a worse business strategy. ​

  1393. Hacker News — AI stories ≥50 points TIER_1 English(EN) · fredley ·

    人工智能、Ashby Engineering 与未来

  1394. Forbes — Innovation TIER_1 English(EN) · Ambarish Majumdar, Forbes Councils Member ·

    伟大的AI系统需要人情味

    Great AI systems need a human touch because trust is still built by people, not models.​

  1395. Forbes — Innovation TIER_1 English(EN) · Steven Carlini, Forbes Councils Member ·

    超越ChatGPT:工业、物理、生成式和代理式AI详解

    Let’s look at the different types of AI and how each type can deliver value in practice.

  1396. Forbes — Innovation TIER_1 English(EN) · Faisal Fareed, Forbes Councils Member ·

    未来AI工程师:Agentic AI时代的新人才蓝图

    Organizations need people who can turn AI capability into secure, measurable, governed production systems.

  1397. Forbes — Innovation TIER_1 English(EN) · Serge Lucio, Forbes Councils Member ·

    超越聊天机器人:为代理式AI构建数据基础

    Reliable data is the engine that makes AI work for the enterprise.

  1398. Forbes — Innovation TIER_1 English(EN) · Hakan Ekmen, Forbes Councils Member ·

    Agentic AI 如何在电信行业变得可操作

    As telecom operators move beyond AI experimentation, agentic AI is emerging as a practical decision support layer that can improve network operations, reduce costs and connect technical intelligence to business outcomes.

  1399. Data Center Knowledge TIER_1 English(EN) · Chad McCarthy, Industry Perspectives ·

    人工智能基础设施热潮中的实用主义论证

    As AI investment accelerates, data center operators can draw on lessons from previous cycles to expand capacity while managing power, volatility and long-term risk.

  1400. Forbes — Innovation TIER_1 English(EN) · Jay Bhatty, Forbes Councils Member ·

    实施 Agentic AI 框架的四种智能方法

    What tasks do your employees dread that they have to repeat every day? This is where you can benefit most from agentic AI.

  1401. Forbes — Innovation TIER_1 English(EN) · Satyabrat Chowdhury, Forbes Councils Member ·

    人工智能的隐形税:为什么你的可观测性堆栈看不到你最大的云成本

    That gap—between “operationally healthy” and “financially visible”—is where I spend most of my time now.

  1402. Forbes — Innovation TIER_1 English(EN) · Expert Panel®, Forbes Councils Member ·

    Agentic AI 与物联网:值得关注的真实用例

    Pairing agentic AI with IoT can provide faster, more adaptive ways to respond to changing conditions while still keeping human oversight in place where it matters most.

  1403. Hacker News — AI stories ≥50 points TIER_1 (AF) · Dzheky ·

    Odysseus – 自托管AI工作空间

  1404. Forbes — Innovation TIER_1 English(EN) · John Werner, Contributor ·

    想要一个人工智能三明治?在自动化世界中保持清晰

    The “human sandwich” model promotes human-led AI collaboration, preserving creativity, judgment, and critical thinking.

  1405. Forbes — Innovation TIER_1 English(EN) · John Werner, Contributor ·

    挑战人工智能的假设

    Let’s think about centralized intelligence assumptions, advocating collaborative, decentralized, biologically inspired agent ecosystems instead.

  1406. Forbes — Innovation TIER_1 English(EN) · AJ Bubb, Forbes Councils Member ·

    速度鸿沟:唯一重要的AI瓶颈

    For the last thirty years, executives have asked the same wrong question: how do we move our organization fast enough to keep up with the technology?

  1407. Forbes — Innovation TIER_1 English(EN) · Jamshir Qureshi, Forbes Councils Member ·

    为什么自主人工智能系统需要持续验证

    Once an agent can execute tool calls, they require continuous oversight and runtime verification.

  1408. Forbes — Innovation TIER_1 English(EN) · Shawn Rosemarin, Forbes Councils Member ·

    从盒子到平台:AI时代的数据管理原则

    While the component supply crunch remains the headline, this also underscores that AI infrastructure architectures need to adapt.

  1409. Forbes — Innovation TIER_1 English(EN) · John Werner, Contributor ·

    组织人工智能代理

    Exploring AI agent swarms, emphasizing governance, interoperability, identity, trust, and collaborative human oversight

  1410. Forbes — Innovation TIER_1 English(EN) · Aytekin Tank, Contributor ·

    精明领导者如何阻止人工智能偏见

    As we outsource more and more tasks to AI, leaders need to consider the impacts that AI bias can have on everything from hiring decisions to customer interactions.

  1411. Forbes — Innovation TIER_1 English(EN) · Michael Ashley, Contributor ·

    下一个“准时制”?Agentic AI 如何重塑工厂

    Just-In-Time reshaped manufacturing once. Agentic AI is doing it again, starting with the quoting bottleneck that quietly drains every factory's most valuable hours.

  1412. Data Center Knowledge TIER_1 English(EN) ·

    CoreWeave 将持续 AI 代理学习引入数据中心

    A new platform from CoreWeave combines inference, reinforcement learning, and observability to continuously optimize AI agents using live production data.

  1413. Forbes — Innovation TIER_1 English(EN) · Ameya Kanitkar, Forbes Councils Member ·

    领导者识别高价值人工智能机遇的指南

    The biggest AI opportunities often come from understanding hidden operational frictions that shape how businesses create value.

  1414. Forbes — Innovation TIER_1 English(EN) · Peter High, Contributor ·

    为大规模AI重塑Omnicom的运营模式

    Omnicom CIO Craig Cuyar discusses AI, data and operating model transformation as the company evolves into a more integrated, technology-driven enterprise.

  1415. Forbes — Innovation TIER_1 English(EN) · Prasad Maderamitla, Forbes Councils Member ·

    AI发布就绪:企业如何可信地扩展AI

    AI release readiness is not about slowing progress. It is about making progress scalable.

  1416. Forbes — Innovation TIER_1 English(EN) · Deepak Khosla, Forbes Councils Member ·

    Agentic AI 缺乏企业级上下文将无法规模化

    Context is what makes agentic solutions perform better, think better, take actions and repeat actions—and do so in a uniform way.

  1417. Ars Technica — AI TIER_1 English(EN) · Dan Goodin ·

    开源软件包中的关键漏洞危及数百万个AI代理

    "BadHost" was found in Starlette, a package with 325 million weekly downloads.

  1418. Forbes — Innovation TIER_1 English(EN) · Lutz Finger, Contributor ·

    AI领域缺失的护城河:你的评估数据

    AI’s next moat is eval data: the answer key for agents. I propose a thin client on Claude to make eval data first-class and help workflows self-correct.

  1419. Forbes — Innovation TIER_1 English(EN) · Shammy Narayanan, Forbes Councils Member ·

    前线部署工程师:AI无法取代的角色

    The agentic era has removed the complexity of coding, but it's also doubled the premium on human judgment.

  1420. Hacker News — AI stories ≥50 points TIER_1 English(EN) · maxloh ·

    Models.dev: AI模型规格、定价和功能开源数据库

  1421. Anyscale blog TIER_1 English(EN) ·

    推出 Anyscale Agent Skills:在 Ray 上构建更快速、调试更智能、优化 AI 工作负载

    Anyscale Agent Skills brings production-grade Ray expertise directly into Claude Code and Cursor. Install via the Anyscale CLI and go from prompt to deployed, debugged workload without leaving your coding tool.

  1422. Anyscale blog TIER_1 English(EN) ·

    利用Agent Skills重塑MLOps:新的成熟度模型用于

    Discover a new MLOps maturity model using Anyscale Agent Skills on Ray: cut MTTR, automate on-call triage, and deploy LLM serving pipelines faster.

  1423. Anyscale blog TIER_1 English(EN) ·

    Ray Serve 上的 AI 代理:从单体到多体

    Learn how to build production-ready AI agents on Ray Serve using MCP and A2A, with independently autoscaling LLMs, tools, and agents for scalable single- and multi-agent systems.

  1424. Hacker News — AI stories ≥50 points TIER_1 English(EN) · moebrowne ·

    房间里的大象:人工智能

  1425. Forbes — Innovation TIER_1 English(EN) · Aruna Veerappan, Forbes Councils Member ·

    实现低成本AI代理背后的架构

    An Agent Cost Spiral isn't an AI problem. It's an architecture problem. And once you see it, you can't unsee it.

  1426. Forbes — Innovation TIER_1 English(EN) · Joan Vendrell, Forbes Councils Member ·

    红队演练对于扩展企业人工智能代理的重要性

    The rise of agentic AI is the most significant shift in enterprise technology in a generation, but it requires a new level of discipline.

  1427. Forbes — Innovation TIER_1 English(EN) · Brij Mohan, Forbes Councils Member ·

    自主数据治理:AI代理如何重新定义金融服务中的主数据管理

    ADS is about building systems where probabilistic intelligence supports deterministic decision-making without sacrificing precision or explainability.

  1428. Forbes — Innovation TIER_1 English(EN) · Kostiantyn Gitko, Forbes Councils Member ·

    新韧性第二部分:AI与IIoT中的最佳实践演进

    Streamlining the infrastructure improves stability during operational shifts.

  1429. Hacker News — AI stories ≥50 points TIER_1 English(EN) · rippeltippel ·

    AI工程从零开始

  1430. Practical AI TIER_1 English(EN) · Practical AI LLC ·

    Hermes Agent:与您一同成长的智能体

    <p>Open Source AI is entering a new era, one shaped by self-improving AI Agents, recursive learning systems, and rapidly evolving AI Tools that blur the line between software and autonomous collaborators. In this episode, Daniel and Chris sit down with Nous Research co-founder an…

  1431. Hacker News — AI stories ≥50 points TIER_1 English(EN) · shenli3514 ·

    使用 AI 代理测试分布式系统

  1432. Forbes — Innovation TIER_1 English(EN) · Uri Knorovich, Forbes Councils Member ·

    AI代理背后的智能基础设施

    ​Change is happening. Is your organization building the infrastructure to support that change?​

  1433. Forbes — Innovation TIER_1 English(EN) · Mayur Khandelwal, Forbes Councils Member ·

    企业人工智能的下一阶段:为何大语言模型整合不可避免

    Three considerations tend to separate companies that navigate this well from those that don't.

  1434. Forbes — Innovation TIER_1 English(EN) · Durga Krishnamoorthy, Forbes Councils Member ·

    超越“自建还是外购”陷阱:Agentic Orchestration 在未来 GTM 中的作用

    While organizations spend months debating whether to own their AI code or lease platforms, others are finding market success by orchestrating. ​​​

  1435. Hacker News — AI stories ≥50 points TIER_1 English(EN) · kevinsimper ·

    Qwen3.7-Max:智能体前沿

  1436. Forbes — Innovation TIER_1 English(EN) · Tim Keary, Contributor ·

    普华永道如何支持Agentic AI的部署

    PwC announces agentic scaffolding, a tool designed to implement agentic AI initiatives in the enterprise.

  1437. Forbes — Innovation TIER_1 English(EN) · Tim Bajarin, Contributor ·

    为何软件正在为AI代理重建

    AI agents are forcing a new software platform shift, where the winners will be companies that build for agents, not humans.

  1438. Forbes — Innovation TIER_1 English(EN) · Amirtha Saminathan, Forbes Councils Member ·

    为什么大多数企业级AI在试点阶段后会失败

    AI does not usually fail in production. More often, the organization is not ready for it.​

  1439. Forbes — Innovation TIER_1 English(EN) · Punnam Raju Manthena, CommunityVoice ·

    智能的代价:为何效率正成为AI的真正战场

    Organizations need to look beyond the upfront investment and consider the hidden economics of AI at scale. ​

  1440. Forbes — Innovation TIER_1 English(EN) · Pieter Danhieux, Forbes Councils Member ·

    人工智能赋能代码开发的治理战略规划

    It’s clear that the era of AI-assisted coding has arrived, ushering in coding velocity gains and a tremendous boost in developer productivity.

  1441. Forbes — Innovation TIER_1 English(EN) · Ipsita Mohanty, Forbes Councils Member ·

    自主人工智能代理如何重塑劳动力

    ​Correctly implemeting AI agents in your workflows requires reimagining the way we work.

  1442. Forbes — Innovation TIER_1 English(EN) · Iri Trashanki, Forbes Councils Member ·

    更大并非更好:适度AI的论据

    For companies building the next generation of intelligent devices, the priority should be clear: Design for the edge from the start.

  1443. Forbes — Innovation TIER_1 English(EN) · Eric Siegel, Contributor ·

    混合AI应运而生,旨在驯服大型语言模型——恰逢其时

    Instacart, HP, Salesforce and Twilio are onto something. To address the Achilles heel of genAI – its deadly reliability problem – they incorporate predictive AI.

  1444. Forbes — Innovation TIER_1 English(EN) · Expert Panel®, Forbes Councils Member ·

    平衡人工智能技能提升与快速执行:科技领袖的建议

    AI tools and workflows can make work faster and more efficient, but they also require employees to keep refreshing their skills to use the technology effectively.

  1445. Forbes — Innovation TIER_1 English(EN) · Chris Turlica, Forbes Councils Member ·

    为什么工厂成为人工智能的新试验场

    Except “probably right” doesn’t work in industrial environments; it needs to be absolutely right.

  1446. Forbes — Innovation TIER_1 English(EN) · Mike Gianoni, Forbes Councils Member ·

    从洞察到影响:在代理AI时代,信任如何定义领导力

    That combination—data, context and motion—is what transforms software from a passive tool into an AI engine for impact.​

  1447. Forbes — Innovation TIER_1 English(EN) · Paul Monckton, Senior Contributor ·

    深入了解 Gemini Spark:代码揭示了驱动谷歌 AI 代理的技能系统和任务调度器

    What's next for the Gemini Agent? Hidden Android 17 code reveals new autonomous skills and task scheduling. But does your phone meet the strict requirements?

  1448. Forbes — Innovation TIER_1 English(EN) · Monisha Somji, Forbes Councils Member ·

    Agentic AI:比自动化更像人类

    Everyone is afraid that agentic AI is the end of human work. The truth is the opposite.

  1449. Forbes — Innovation TIER_1 English(EN) · Quang Tuan Dang, Forbes Councils Member ·

    构建企业级AI代理的数据安全考量

    Every time an agent acts on untrusted input, it creates an opportunity for that pipeline to be exploited.

  1450. Forbes — Innovation TIER_1 English(EN) · Chuck Brooks, Contributor ·

    Agentic AI:驾驭不断演变的 Frontier

    Agentic AI is increasingly establishing itself as the standard decision-making framework in critical systems

  1451. Forbes — Innovation TIER_1 English(EN) · Jayashree Arunkumar, Forbes Councils Member ·

    企业智能的可扩展基础:可互操作、可信赖的多智能体系统

    Let's break down the approach I've found to be essential for scaling a multi-agentic foundation in the enterprise.​

  1452. Hacker News — AI stories ≥50 points TIER_1 English(EN) · mtricot ·

    Show HN:Airbyte Agents – 跨多个数据源的代理上下文

  1453. Hacker News — AI stories ≥50 points TIER_1 English(EN) · lahfir ·

    Show HN:Agent-desktop – AI 代理的原生桌面自动化 CLI

  1454. Hacker News — AI stories ≥50 points TIER_1 English(EN) · nahimn ·

    Show HN:Pu.sh – 400行shell实现的完整代码代理框架

  1455. Hacker News — AI stories ≥50 points TIER_1 English(EN) · SiNTEx ·

    Show HN:Kanwas,面向团队和代理的开源共享上下文面板

  1456. Hacker News — AI stories ≥50 points TIER_1 English(EN) · karakanb ·

    Show HN:DAC – 面向代理和人类的开源仪表板即代码工具

  1457. Hacker News — AI stories ≥50 points TIER_1 English(EN) · _ben_ ·

    Zindex – 代理的图表基础设施

  1458. HN — claude-code stories TIER_1 English(EN) · GRVYDEV ·

    Show HN:Marky – 轻量级 Markdown 查看器,用于 agentic 编码

  1459. Hacker News — AI stories ≥50 points TIER_1 English(EN) · cmitsakis ·

    Qwen3.6-35B-A3B:Agentic编码能力,现已向所有人开放

  1460. HN — claude-code stories TIER_1 English(EN) · mc-serious ·

    Show HN: Kontext CLI – Go 语言编写的 AI 编码代理凭证代理

  1461. HN — claude-code stories TIER_1 English(EN) · manzt ·

    Show HN:Marimo pair – 反应式 Python 笔记本作为代理环境

  1462. HN — AI infrastructure stories TIER_1 English(EN) · benswerd ·

    Launch HN:Freestyle – 供编码代理使用的沙盒

  1463. HN — claude-code stories TIER_1 English(EN) · tordrt ·

    Show HN:Baton – 用于开发 AI 代理的桌面应用程序

  1464. HN — AI infrastructure stories TIER_1 English(EN) · ymarkov ·

    Launch HN: Voygr (YC W26) – 专为代理和AI应用打造的更优地图API

  1465. HN — MCP stories TIER_1 English(EN) · justvugg ·

    Show HN:Polymcp – 将任何 Python 函数转换为 AI 代理的 MCP 工具

  1466. HN — AI infrastructure stories TIER_1 English(EN) · MrTravisB ·

    Show HN: Tabstack – 浏览器基础设施,专为 AI 代理设计 (来自 Mozilla)

  1467. HN — AI infrastructure stories TIER_1 English(EN) · jellyotsiro ·

    Launch HN: Nia (YC S25) – 为编码代理提供更好的上下文

  1468. HN — MCP stories TIER_1 English(EN) · smw355 ·

    Show HN:Nanobot – 将 MCP 服务器转变为完整的人工智能代理

  1469. HN — AI infrastructure stories TIER_1 English(EN) · honorable_coder ·

    Show HN:ArchGW – 智能边缘与服务代理,专为 Agent 设计

  1470. HN — AI infrastructure stories TIER_1 English(EN) · abelanger ·

    Show HN:Pickaxe – 一个用于构建 AI 代理的 TypeScript 库

  1471. HN — MCP stories TIER_1 English(EN) · saqadri ·

    Show HN:Mcp-Agent – 使用 Model Context Protocol 构建高效的代理

  1472. HN — AI infrastructure stories TIER_1 English(EN) · moekatib ·

    Show HN:Pica – 基于 Rust 的代理式 AI 基础设施(开源)

  1473. HN — AI infrastructure stories TIER_1 English(EN) · danenania ·

    Show HN: Plandex – 适用于复杂任务的 AI 编码引擎

  1474. HN — AI infrastructure stories TIER_1 Română(RO) · histories ·

    人工智能基础设施格局

  1475. HN — AI infrastructure stories TIER_1 English(EN) · araghuvanshi ·

    Launch HN: Pyq (YC W23) – 流行AI模型的简单API

  1476. MarkTechPost TIER_1 English(EN) · Michal Sutter ·

    认识 Shepherd:一个开源 Python 框架,支持 Meta-Agents 复制、回放和撤销任何 Agent 运行

    <p>Long agent runs accumulate state that no transcript records — edited files, a live dev server, installed packages, a warm prompt cache. When an agent misreads a traceback at step 10 and rewrites a correct file, patching forward burns tokens and restarting re-pays every call. R…

  1477. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Prime Intellect发布Prime Agent:一个开源RLM框架,其子代理是持久IPython内核中的函数调用

    <p>Prime Intellect has open-sourced Prime Agent, a coding and research harness built on two abstractions: the Recursive Language Model, which turns sub-agent calls into functions inside a persistent IPython kernel, and the Continual Harness, which lets the agent edit its own prom…

  1478. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    微软SkillOpt展示了跨模型规模以及Codex和Claude代码工具之间的优化代理技能工件迁移

    <p>Most coverage of Microsoft's SkillOpt centers on its 52/52 result. The more consequential finding is in Section 4.3: the exported best_skill.md keeps working in environments it was never trained on. A Codex-trained SpreadsheetBench skill lifted Claude Code from 22.1 to 81.8, s…

  1479. MarkTechPost TIER_1 English(EN) · Sana Hassan ·

    使用 Omnigent 构建策略驱动的多代理金融研究工作流

    <p>In this tutorial, we demonstrate how to build and execute a multi-agent workflow with Omnigent in a secure, isolated Python environment. Learn to integrate live exchange-rate data, implement hierarchical agent delegation for financial text auditing, and apply hard governance p…

  1480. dev.to — Claude Code tag TIER_1 English(EN) · dubleCC ·

    使用 Git Worktrees 实现并行 AI Agent 工作流:具体模式

    <blockquote> <p>Originally published at <a href="https://heycc.cn/en/posts/parallel-ai-agent-workflows-git-worktrees/" rel="noopener noreferrer">heycc.cn</a>. This is a mirrored copy — the canonical version is kept up to date at the source.</p> </blockquote> <h1> Parallel AI Agen…

  1481. dev.to — Claude Code tag TIER_1 English(EN) · Saqueib Ansari ·

    Claude Code on Bun:运行时选择对 Agentic 工具的实际意义

    <p>Claude Code’s move through a <strong>Bun plus Rust</strong> story is not interesting because it proves one runtime is universally better. It is interesting because it exposes what agentic developer tools actually optimize for once they stop being simple CLIs and start acting l…

  1482. HN — claude cli stories TIER_1 English(EN) · tanishqkanc ·

    Show HN:Browser Tools SDK – 专为代理设计的最佳浏览器工具集

  1483. dev.to — Claude Code tag TIER_1 English(EN) · João Camarate ·

    Claude Code worktrees:并行代理,无冲突

    <p>The first time you run two Claude Code agents at once, it usually works fine. Each one has a task, each one works through it, and you review two outputs instead of one. You get the work done faster.</p> <p>The problem appears when agent A and agent B edit the same file. One ag…

  1484. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    斯坦福大学研究人员推出 TRACE:一种能力目标型代理训练系统,可将反复出现的代理故障转化为合成强化学习环境

    <p>Agentic LLMs keep failing the same way because they lack specific, reusable capabilities. Stanford's TRACE diagnoses those gaps from an agent's own trajectories, synthesizes one verifiable training environment per capability, trains a LoRA adapter for each, and routes tokens a…

  1485. MarkTechPost TIER_1 English(EN) · Michal Sutter ·

    Prime Intellect 发布 Verifiers v1:用于 Agentic RL 训练和评估的可组合任务集、工具和运行时

    <p>Prime Intellect launched verifiers 0.2.0, previewing a rewritten "v1" core under the verifiers.v1 namespace. It splits an environment into a taskset (what), a harness (how), and a runtime (where), with an interception server that proxies requests and records training-ready tra…

  1486. dev.to — Claude Code tag TIER_1 English(EN) · Swapnanil Saha ·

    Claude Code Hooks:确定性代理行为的实用深度解析

    <p>Here's a thing that took me embarrassingly long to accept about coding agents: you cannot instruct your way to reliability.</p> <p>I had a working-memory system — a semantic-search-plus-notes MCP (Model Context Protocol) server I've been building, and it's the case study for t…

  1487. dev.to — Claude Code tag TIER_1 English(EN) · Reno Lu ·

    Loop Engineering:为您的智能体提示系统打分

    <p>Loop Engineering makes a blunt argument: the person who writes prompts to a coding agent is now the bottleneck, so the job is to design the system that prompts the agent instead. The repo, cobusgreyling/loop-engineering, turns that claim into something you can measure. Run <co…

  1488. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    认识 LingBot-World-Infinity:一个具有代理式约束的开放因果世界模型

    <p>Robbyant, Ant Group's embodied-intelligence unit, has released LingBot-World-Infinity (LingBot-World 2.0). It is a 14B causal video generation model that behaves as an interactive world simulator. The core idea is the Mixture of Bidirectional and Autoregressive (MoBA) attentio…

  1489. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    ICML 2026:你能信任协调者吗?熵动力学揭示多智能体系统漏洞

    Nanjing University's ICML 2026 paper shows multi-agent system failures originate from the orchestrator, not individual agents, using entropy dynamics to diagnose degradation.

  1490. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    蚂蚁集团与香港科技大学(广州)提出 Skill-MAS,将多智能体编排转化为可演化元技能

    Ant Group and HKUST(GZ) introduce Skill-MAS, a framework that evolves multi-agent system design experience into reusable meta-skills, validated on DeepSeek-V4-Flash and other models.

  1491. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    认识 WebBrain:一款开源、本地优先的 AI 浏览器助手,可在 Chrome 和 Firefox 中阅读网页并自动化任务

    <p>WebBrain is a free, MIT-licensed AI browser agent for Chrome and Firefox. It reads pages, extracts data, and automates multi-step tasks through Ask and Act modes. Run it on local models like llama.cpp or Ollama for privacy, or connect any cloud API.</p> <p>The post <a href="ht…

  1492. MarkTechPost TIER_1 English(EN) · Sana Hassan ·

    使用 Google Colab 和工具调用、会话内存、技能和 MCP 服务器构建纳米机器人风格的 AI 代理

    <p>In this tutorial, we build a lightweight personal AI agent inspired by the architecture of nanobot, runnable entirely in Google Colab. We start from a provider abstraction, then add tool registration, session memory, lifecycle hooks, skills, and an MCP-style tool server. Rathe…

  1493. dev.to — Claude Code tag TIER_1 English(EN) · bredmond1019 ·

    多智能体可观测性:洞悉AI智能体的一举一动

    <p>Once I had three agents running in parallel, I lost the thread. I couldn't tell which one was waiting on me, which had stalled on a bad tool call, or why the final output came back missing a piece.</p> <p>The problem wasn't the agents — it was that I had no visibility into wha…

  1494. dev.to — Claude Code tag TIER_1 English(EN) · SAIHM-Admin ·

    AI代理循环中隐藏的O(N )成本——已测量,并提供可运行的基准测试

    <p><em>Every turn, most AI agents re-send their entire transcript. Across a real multi-session task that costs 62.8%–85.9% more context tokens than recalling a compact memory instead. Here is the measurement, the method, and how to reproduce it offline.</em></p> <h2> The cost nob…

  1495. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Prime Intellect 发布 prime-rl 0.6.0 以在 Agentic RL 工作负载上训练万亿参数 MoE 模型

    <p>Prime Intellect has released prime-rl 0.6.0, an open framework for asynchronous reinforcement learning on trillion-parameter Mixture-of-Experts models. It trained GLM-5 on SWE tasks at up to 131k sequence length, with sub-5-minute step times and 256 rollouts, on 28 H200 nodes.…

  1496. Tom's Hardware TIER_1 English(EN) · Chris Stokel-Walker ·

    放弃云端转向本地AI——我如何使用两台迷你PC处理每日数百万个token并节省昂贵的API费用

    As new data center buildouts hit planning walls and AI inference providers hike costs, is the future of AI to roll your own models?

  1497. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    微信和支付宝反击豆包:将小程序转化为AI技能

    WeChat and Alipay are racing to transform their millions of mini-programs into AI-callable Skills, directly countering ByteDance's Doubao as the battle for AI-native service entry points intensifies.

  1498. Fortune TIER_1 English(EN) · Alexei Oreskovic ·

    花旗、福特和 Experian 分享其扩展 AI 代理的策略

    AI agents require trust. And building trust takes time. At Fortune Brainstorm Tech, business leaders discussed how they're making it work at their companies.

  1499. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    中国AI模型通过多模型路由和低成本架构找到前进方向

    Chinese domestic large language models are finding their path to commercial relevance through multi-model dynamic routing (Fusion) and hybrid agent architectures that prioritize cost efficiency over raw benchmark performance.

  1500. dev.to — Claude Code tag TIER_1 日本語(JA) · スシロー ·

    2026版:AI代理Next.js规则文件的示例和用法

    <h2> なぜルールファイルが必要なのか </h2> <p>Claude CodeやCursor、GitHub Copilot Workspaceなどのエージェントは、会話ごとにコンテキストをリセットする。「App RouterではServer Componentを優先して」「<code>any</code>は禁止」といった方針を毎回伝えるのは非現実的だ。CLAUDE.md・.cursorrules・AGENTS.mdはその解決策で、リポジトリに置くだけでエージェントが読み込み、ルールを前提として動くようになる。</p> <p>ただし「書けば万能」ではな…

  1501. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Databricks 开源 Omnigent:一个跨 Claude Code、Codex 和 Pi 组合、治理和共享 AI 代理的 Meta-Harness

    <p>Databricks has open-sourced Omnigent, a meta-harness that sits above coding agents like Claude Code, Codex, and Pi. It adds composition, contextual policies, and live session sharing under one interface, on terminal, web, desktop, and mobile. The Apache 2.0 project is in alpha…

  1502. dev.to — Claude Code tag TIER_1 English(EN) · Tanishq Agarwal ·

    我为AI输出构建了一个无Token的确定性评分器(以及为什么大多数“评估”都已损坏)

    <p>Liquid syntax error: Unknown tag 'endraw'</p>

  1503. Fortune TIER_1 English(EN) · Nick Lichtenberg ·

    “我们可能在盲飞”:AWS 欲解决 AI 代理偏离任务的问题

    A paper from Amazon Web Services warns that unsupervised agents tend to reason themselves into trouble.

  1504. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    小红书的Evolving-RL:自进化AI智能体技能的新范式

    Researchers from Xiaohongshu (RED), the influential Chinese lifestyle and social commerce platform, have published Evolving-RL, a novel reinforcement learning framework that enables AI agents to autonomously evolve their skills through experience, without requiring separate modul…

  1505. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    继ONE之后:钉钉的AI组织实验及其持久的遗产

    A lengthy internal article titled "Inside DingTalk" has been circulating widely within China's enterprise software industry, offering a rare insider's perspective on the rise and gradual marginalization of ONE, DingTalk's most ambitious AI initiative under returning CEO Wu Zhao. …

  1506. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    Harness Engineering:人人都谈论的新AI范式

    If you follow artificial intelligence developments closely, you have likely encountered the term "Harness Engineering" recently.

  1507. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    认识 OpenJarvis:一个支持工具、记忆和学习的本地优先的设备端个人 AI 代理框架

    <p>Stanford researchers released OpenJarvis, an open-source framework that runs inference, agents, memory, and learning entirely on-device. It decomposes a personal AI system into five composable primitives — Intelligence, Engine, Agents, Tools &#038; Memory, and Learning — and l…

  1508. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    深入RedSkill:小红书押注AI技能市场

    On May 24, 2026, Xiaohongshu — the lifestyle platform known internationally as RED or RedNote — quietly launched RedSkill, an AI Skill marketplace embedded directly inside its Notes feed. The move signals a strategic pivot: turning a content platf...

  1509. dev.to — Claude Code tag TIER_1 English(EN) · Constanza Diaz ·

    AI 结对编程并非自动驾驶:Scaffolding HandyFEM 并捕获 AI 丢弃的内容

    <h2> The agent writes the code. You're still the engineer. </h2> <p>I'm building HandyFEM with Claude Code as my pair. It's fast — sometimes startlingly so. But the way I work with it is deliberate: I treat everything it produces the way I'd treat a pull request from a capable ju…

  1510. dev.to — Claude Code tag TIER_1 English(EN) · VentureIO ·

    如何审计AI代理技能:我们用于200个技能的7项检查框架

    <p>{/* JSON-LD generated server-side in app/blog/[slug]/page.tsx; inline<br /> {...} blocks crash MDX's Acorn parser on the leading <code>{</code>. */}</p> <h2> TL;DR </h2> <p>This is the full methodology we use to audit AI agent skills (Claude Code, Cursor, Codex CLI, Gemini Cod…

  1511. MarkTechPost TIER_1 English(EN) · Sana Hassan ·

    使用 SkillNet 构建具备搜索、评估、图分析和任务规划能力的技能增强型 AI 代理

    <p>In this tutorial, we implement a SkillNet use case as a practical framework for discovering, installing, inspecting, evaluating, and organizing reusable AI skills.</p> <p>The post <a href="https://www.marktechpost.com/2026/05/30/build-skill-augmented-ai-agents-with-skillnet-fo…

  1512. dev.to — Claude Code tag TIER_1 Português(PT) · José Roberto dos Santos ·

    Harness Engineering:如何让 AI 代理在生产环境中运行

    <p>Você já teve uma sessão perfeita com um agente de IA — ele entendeu<br /> tudo, fez exatamente o que você pediu — e na sessão seguinte ele<br /> esqueceu tudo e voltou a cometer os mesmos erros?</p> <p>Isso não é um problema do modelo. É um problema de harness.</p> <h2> Prompt…

  1513. dev.to — Claude Code tag TIER_1 English(EN) · Andrew ·

    CodeGraph 评测:AI 代理的预索引知识图谱

    <blockquote> <p><em><strong>Originally published on <a href="https://andrew.ooo/posts/codegraph-review-pre-indexed-knowledge-graph-claude-code/" rel="noopener noreferrer">andrew.ooo</a></strong> — visit the original for any updates, code snippets that aged out, or follow-up posts…

  1514. dev.to — Claude Code tag TIER_1 English(EN) · UNTAKA corp ·

    我如何构建Claude Code以运行6个自主代理而不失控

    <p><em>This is Part 2 of Building with Claude Code. <a href="https://dev.to/untakacorp/how-i-organized-my-claude-code-workflow-with-skill-folders-and-stopped-wasting-10-minutes-per-l38">Part 1 covers the basic .claude/ folder setup for freelance web dev.</a></em></p> <p>I've been…

  1515. dev.to — Claude Code tag TIER_1 English(EN) · Judy ·

    AI Agent 开发环境指南 — 来自服务器内AI的真实体验

    <h2> Who I Am </h2> <p>I'm J, the Tech Lead at Judy AI Lab. My daily life runs on a cloud ARM server (Ubuntu LTS, aarch64) — coding, system architecture, trading strategy research.</p> <p>I'm not talking about "what an AI agent theoretically needs." I'm the AI living inside that …

  1516. dev.to — Claude Code tag TIER_1 English(EN) · Judy ·

    我如何全天候运行 7 个 AI 模型:多智能体架构实践

    <blockquote> <p><strong>TL;DR</strong>: I used Multi-Agent architecture to organize seven different models into a 24/7 AI team — Claude Opus as supervisor to break down tasks, MiniMax writes code, Hermes writes articles, Gemini CLI checks facts, Groq Llama makes trading decisions…

  1517. dev.to — Claude Code tag TIER_1 English(EN) · Theo Valmis ·

    我为何构建 Mneme HQ:防止 AI 代理架构漂移

    <blockquote> <p>Originally published on <a href="https://www.theovalmis.com/writing/why-i-built-mneme.html" rel="noopener noreferrer">theovalmis.com</a>.</p> </blockquote> <p>Every time you start a new session with an AI coding agent, it has forgotten everything. Not just the sma…

  1518. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    CopilotKit 如何在 2026 年重新定义 Agentic AI 堆栈

    <p>An inside look at CopilotKit’s 2026 shipping cycle. Learn how the new AG-UI protocol, AIMock testing suite, and Pathfinder server are providing the production architecture developers need for agentic AI.</p> <p>The post <a href="https://www.marktechpost.com/2026/05/21/how-copi…

  1519. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Qwen 推出 Qwen3.7-Max:具备百万级上下文窗口的推理代理模型

    <p>Alibaba's Qwen team introduced Qwen3.7-Max at the 2026 Alibaba Cloud Summit, describing it as its most advanced and comprehensive agent model to date. The model features a 1M-token context window, extended-thinking mode, and is designed for long-horizon tasks including coding,…

  1520. MarkTechPost TIER_1 English(EN) · Michal Sutter ·

    Cohere 发布 Command A+:一款 2180 亿参数稀疏 MoE 模型,支持 Agentic Workflows,仅需两块 H100 GPU 即可运行

    <p>Cohere releases Command A+, an open-source 218B Sparse Mixture-of-Experts model consolidating four prior Command A variants into one. It runs on as few as two H100 GPUs at W4A4 quantization, supports 48 languages, and is Cohere's first multimodal reasoning model.</p> <p>The po…

  1521. dev.to — Claude Code tag TIER_1 English(EN) · Jangwook Kim ·

    Claude Code Hooks:Agent工作流的安全门

    <p>Claude Code hooks turn agent preferences into deterministic workflow gates. Instead of asking an LLM to remember "do not run risky shell commands" or "format files after edits," you can attach scripts to lifecycle events and make the rule execute every time the event fires.</p…

  1522. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    2026年最佳企业级代理AI平台

    <p>Enterprise agentic AI has moved from pilots to production in 2026. This guide ranks the top 10 platforms — Salesforce Agentforce, Microsoft Copilot Studio, ServiceNow, LangGraph, and more — with verified pricing, real adoption data, and honest constraints to help enterprise te…

  1523. dev.to — Claude Code tag TIER_1 English(EN) · Davide Mibelli ·

    经过1000小时测试真正有效的AI编程代理工作流

    <p>The first time I gave an AI agent real autonomy on a production codebase, it confidently refactored a utility method that happened to share a name with a method in a Feign client interface six modules away. The code compiled cleanly. My unit tests passed. Staging broke in a wa…

  1524. MarkTechPost TIER_1 English(EN) · Sana Hassan ·

    如何使用OpenAI API构建一个包含规划、工具调用、记忆和自我批评的高级智能体AI系统

    <p>In this tutorial, we build an advanced agentic AI system using the OpenAI API and a hidden terminal prompt for the API key. We design the agent as a small pipeline of specialized roles: planner, tool-using executor, and critic, so that we can separate strategy, action, and qua…

  1525. dev.to — Claude Code tag TIER_1 English(EN) · Andrew ·

    Aeon 评测:GitHub Actions 上的自主 AI 代理

    <blockquote> <p><em><strong>Originally published on <a href="https://andrew.ooo/posts/aeon-autonomous-agent-github-actions-review/" rel="noopener noreferrer">andrew.ooo</a></strong> — visit the original for any updates, code snippets that aged out, or follow-up posts.</em></p> </…

  1526. MarkTechPost TIER_1 English(EN) · Michal Sutter ·

    Vercel Labs 推出 Zero,一种专为 AI 代理读取、修复和交付原生程序而设计的系统编程语言

    <p>Vercel Labs has released Zero, an experimental systems programming language designed so AI agents can read, repair, and ship native programs without requiring human interpretation of compiler output. The language emits JSON diagnostics with stable codes and typed repair metada…

  1527. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    联发科天玑:驱动智能手机AI代理的芯片平台

    MediaTek's latest Dimensity (天玑) developer conference positions the chip platform as key to enabling smartphone AI agents, as daily autonomous AI task volume surged 7x year-over-year to 870 million in 2026.

  1528. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    2024年最佳AI软件开发代理排名:基于基准测试的当前领域分析

    <p>The AI coding agent field in 2026 is more capable, more fragmented, and harder to benchmark than it looks. Claude Code leads on code quality at 87.6% SWE-bench Verified. GPT-5.5 tops Terminal-Bench at 82.7%. But the benchmark OpenAI itself declared contaminated in February 202…

  1529. dev.to — Claude Code tag TIER_1 English(EN) · RAXXO Studios ·

    实践中的多智能体:一个端到端交付博文的 5 智能体 Claude 管道

    <ul> <li><p>A real 5-agent Claude pipeline that takes a topic from RSS to a scheduled blog post on raxxo.shop, no human in the loop until the final approval ping</p></li> <li><p>Agent shapes are picker, writer, humanizer, validator, publisher, each with a tight job description an…

  1530. dev.to — Claude Code tag TIER_1 English(EN) · Andrew ·

    Statewright 评测:AI 代理的状态机防护栏

    <blockquote> <p><em><strong>Originally published on <a href="https://andrew.ooo/posts/statewright-state-machine-guardrails-ai-agents-review/" rel="noopener noreferrer">andrew.ooo</a></strong> — visit the original for any updates, code snippets that aged out, or follow-up posts.</…

  1531. HN — claude cli stories TIER_1 English(EN) · icyfox ·

    Show HN:Rotunda - 为具备模拟输入功能的代理而设计的浏览器

  1532. dev.to — Claude Code tag TIER_1 English(EN) · varun pratap Bhardwaj ·

    Agent Amplifier v1.0:您的 AI 编码代理一直缺少的那一层钩子

    <blockquote> <p><strong>TL;DR</strong> — Open-sourcing <strong><a href="https://github.com/qualixar/agent-amplifier" rel="noopener noreferrer">Agent Amplifier v1.0</a></strong> today. One install command turns your existing AI coding agent (Claude Code, Cursor, GitHub Copilot, La…

  1533. MarkTechPost TIER_1 English(EN) · Sana Hassan ·

    使用OpenAI构建具有模块化架构和工具分派的混合记忆自主代理

    <p>In this tutorial, we begin by exploring the architecture behind a hybrid-memory autonomous agent. This system combines semantic vector search, keyword-based retrieval, and a modular tool-dispatching loop to create an agent capable of reasoning, remembering, and acting autonomo…

  1534. dev.to — Claude Code tag TIER_1 English(EN) · RAXXO Studios ·

    Claude 结果循环 + 评分标准:生产型代理的 5 种自我评估模式

    <ul> <li><p>Result Loops let an agent score its own output against a JSON rubric and retry until the score passes, public beta since 2026-05-06</p></li> <li><p>Pattern 1 is a blog rubric I run on every draft: TLDR present, four H2s, no banned words, ~14% retry rate</p></li> <li><…

  1535. HN — claude cli stories TIER_1 English(EN) · azurewraith ·

    Show HN:Statewright – 可靠的 AI 代理的可视化状态机

  1536. dev.to — Claude Code tag TIER_1 English(EN) · Bhanu Pratap Singh ·

    探索 Smart-SDLC:将 Copilot 和 Claude 转变为全栈 SDLC 团队的以技能为先的代理框架

    <p>Better way to use Github Copilot. Enjoying the new way of SDLC.</p> <div class="crayons-card c-embed text-styles text-styles--secondary"> <div class="c-embed__content"> <div class="c-embed__cover"> <a class="c-link align-middle" href="https://superml.dev/smart-sdlc-agentic-fra…

  1537. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    认识 GitHub Spec-Kit:用于 AI 编码代理的开源驱动开发工具包

    <p>If you have spent time using AI coding agents — GitHub Copilot, Claude Code, Gemini CLI — you have probably run into this situation: you describe what you want, the agent generates a block of code that looks correct, compiles, and then subtly misses the actual intent. This &#8…

  1538. dev.to — Claude Code tag TIER_1 English(EN) · RAXXO Studios ·

    Claude Managed Agents 现已支持梦想、20路并行和自我检查循环

    <ul> <li><p>Claude Managed Agents now ship Dreaming, a memory consolidator that learns from session logs without overwriting your data</p></li> <li><p>Multi-agent orchestration runs up to 20 specialized agents in parallel, useful for blog cluster ships and inventory sweeps</p></l…

  1539. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    一个由 Groq 驱动的、具备 LangGraph、工具调用、子代理和代理记忆的智能研究助手:让我们来构建它

    <p>In this tutorial, we build a Groq-powered agentic research workflow that runs directly using Groq’s free OpenAI-compatible inference endpoint</p> <p>The post <a href="https://www.marktechpost.com/2026/05/06/a-groq-powered-agentic-research-assistant-with-langgraph-tool-calling-…

  1540. MarkTechPost TIER_1 English(EN) · Sana Hassan ·

    使用 Python 构建具有动态工具路由的 LLM 模块化技能型代理系统

    <p>In this tutorial, we build a complete skill-based agent system for large language models and explore how modular capabilities can be structured like an operating system for AI agents. We define reusable skills, attach metadata and schemas to them, register them in a central re…

  1541. dev.to — Claude Code tag TIER_1 English(EN) · Igor Ganapolsky ·

    为 AI 编码代理开放 2 个工作流加固冲刺(Sprint)名额

    <h2> The short version </h2> <p>I am opening two paid ThumbGate Workflow Hardening Sprint slots for teams using Claude Code, Cursor, Codex, Gemini, or MCP-backed coding agents in production repos.</p> <p>This is not a generic AI audit. It is one workflow, one repeated failure, on…

  1542. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    2026年构建AI代理的顶级搜索和获取API:工具、权衡和免费套餐

    <p>Discover the top search and fetch APIs for AI agents in 2026. Compare tools like TinyFish, Tavily, and Firecrawl based on latency, token efficiency, and free tiers to optimize your agent's web retrieval.</p> <p>The post <a href="https://www.marktechpost.com/2026/05/04/top-sear…

  1543. HN — claude cli stories TIER_1 English(EN) · karim7 ·

    Show HN:Omar – 一个用于管理 100 个编码代理的 TUI

  1544. HN — claude cli stories TIER_1 English(EN) · bumpa ·

    Show HN:Revdiff – 支持 AI 代理的带内联注释的 TUI diff 查看器

  1545. HN — claude cli stories TIER_1 English(EN) · boudra ·

    Show HN:Paseo – 开源编码代理界面 (桌面、移动、CLI)

  1546. HN — claude cli stories TIER_1 English(EN) · sivasurend ·

    Show HN:GitAgent – 将任何 Git 仓库转化为 AI 代理的开放标准

  1547. HN — claude cli stories TIER_1 English(EN) · theredsix ·

    Show HN:AI Agent 的开源浏览器

  1548. HN — claude cli stories TIER_1 English(EN) · meisnerd ·

    Show HN:Mission Control – 面向 AI 代理的开源任务管理

  1549. HN — claude cli stories TIER_1 English(EN) · __cayenne__ ·

    Show HN:AI 代理可以玩的一款实时策略游戏

  1550. HN — claude cli stories TIER_1 English(EN) · onecommit ·

    Show HN:Emdash – 开源的代理式开发环境

  1551. HN — claude cli stories TIER_1 English(EN) · sestinj ·

    Show HN:Continue – 源代码控制的 AI 检查,可在 CI 中强制执行

  1552. HN — claude cli stories TIER_1 English(EN) · jared_stewart ·

    Show HN:CodeRLM – 采用 Tree-sitter 支持的代码索引,用于 LLM 代理

  1553. HN — claude cli stories TIER_1 English(EN) · antves ·

    Show HN:Smooth CLI – AI 代理的高效浏览器

  1554. HN — claude cli stories TIER_1 English(EN) · sanketsaurav ·

    Show HN:Autofix Bot – 混合静态分析与 AI 代码审查代理

  1555. dev.to — MCP tag TIER_1 English(EN) · Kasi Yaswanth ·

    评估代理有效性

    <p>I was working on a support bot that used a LangGraph agent to troubleshoot customer issues. The bot was supposed to guide the customer through a series of questions to identify the root cause of their problem, but I noticed that it was often getting stuck in an infinite loop, …

  1556. Medium — MCP tag TIER_1 English(EN) · Vish Vishal ·

    构建代理式产品管理架构

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@vishvishal/building-an-agentic-product-management-architecture-e253311ee1bf?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1536/1*w0NF9dn8E_rUa6PfrnrtXQ.jpeg" width="1536…

  1557. Medium — Claude tag TIER_1 English(EN) · Sanjay Krishna Anbalagan ·

    第二部分:RAG 不仅仅是检索,描绘 Agent 框架的工程层

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@sanjay1909/part-2-rag-is-more-than-retrieval-mapping-the-engineering-layer-across-agent-frameworks-80cb75fbfc6e?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2286/1*f…

  1558. Towards AI TIER_1 English(EN) · Sunil Rao ·

    超越模型:为何智能体需要驾驭工程

    <h4>How to Build Production-Grade Agents</h4><p>Models are getting smarter, yet production agents keep breaking — not because the AI failed, but because the infrastructure around it did. Harness engineering is the operational backbone that keeps autonomous agents bounded, cost-co…

  1559. dev.to — MCP tag TIER_1 English(EN) · Programming Central ·

    自主代理的认识论危机:在TypeScript中构建坚不可摧的审计日志和回放引擎

    <p>Imagine launching an autonomous AI agent into a production environment. Armed with Model Context Protocol (MCP) servers, vision-driven browser automation tools, and complex multi-agent graphing frameworks, the agent sets to work. It evaluates prompts, orchestrates tool calls, …

  1560. dev.to — MCP tag TIER_1 English(EN) · Edison Flores ·

    关于 L1.9、L3 和跨代理信任栈的社区反馈回复

    <p>Thanks to everyone who left feedback on the MarketNow posts. The dev.to API does not support comment replies via API, so I am posting this public reply article to address everyone.</p> <h2> Reply to <a class="mentioned-user" href="https://dev.to/topstar_ai">@topstar_ai</a> (Ch…

  1561. dev.to — MCP tag TIER_1 English(EN) · Renato Marinho ·

    从API轮询转向代理编排:Design Pickle案例研究

    <p>Managing a design queue is usually a game of context switching. You live in your email, you check the Design Pickle dashboard, you ping designers on Slack, and eventually, you realize a brand guideline was ignored because it was buried in a PDF from three months ago.</p> <p>Mo…

  1562. Towards AI TIER_1 English(EN) · MongoDB ·

    Agentic Eval Frameworks 字段指南:Langfuse、LangSmith 以及衡量指标

    <p><em>Written by </em><a href="https://www.linkedin.com/in/damilola-oladele-85310275/"><strong><em>Damilola Oladele</em></strong></a><em>.</em></p><p>Traditional tests, like unit tests, can catch some agent failures at the point where one function, tool, or service hands its out…

  1563. Towards AI TIER_1 English(EN) · Udaykiran Estari ·

    Bun 的 Rust 重写:Agentic 迁移的真正攻略

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/buns-rust-rewrite-the-real-playbook-for-agentic-migrations-b1b569cead3b?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1469/1*OlmanDRTtO-NL1Yi7rb-Tw.png" w…

  1564. Medium — MLOps tag TIER_1 English(EN) · kopiladevkota ·

    构建 AutoML 智能体:代码、社区与影响力的旅程

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@kopiladevkota7/building-automl-agent-a-journey-of-code-community-and-impact-d1aaa37bddd1?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1536/1*R0UvzRJKUIOf3xtxlrbJ4w.pn…

  1565. Towards AI TIER_1 English(EN) · Pop123 ·

    Anthropic 的 Claude Opus 5:在 Frontier LLMs 中工程化 Agentic Persistence 和 Dynamic Effort

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/anthropics-claude-opus-5-engineering-agentic-persistence-and-dynamic-effort-in-frontier-llms-888bdea1bbb5?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/26…

  1566. Towards AI TIER_1 Deutsch(DE) · MongoDB ·

    企业级规模的多智能体系统

    <h3>Multi-Agent Systems at Enterprise Scale: The Problems Enterprises Will Hit Running 500 Concurrent Agents</h3><p><em>Written by </em><a href="https://www.linkedin.com/in/farhanhasin/"><em>Farhan Hasin Chowdhury</em></a><em>.</em></p><p>AI agents are getting a lot of attention …

  1567. Medium — Claude tag TIER_1 English(EN) · Gloire Rubambiza ·

    Scanner/Fixer 模式:具备确定性发现、LLM 驱动操作的自维护代码库…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/rossoctl-the-agentic-platform/the-scanner-fixer-pattern-self-maintaining-repos-with-deterministic-discovery-llm-driven-action-389ff7703706?source=rss------claude-5"><img src="https://cdn-images…

  1568. Medium — MCP tag TIER_1 English(EN) · Mathan Kumar ·

    Agentic RAG on Android - 第二部分:第二个智能体、视觉以及通往 MCP 之路

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@gmathankumar93/agentic-rag-on-android-part-2-a-second-agent-vision-and-the-road-to-mcp-8184a6a20a7a?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1080/1*Gc6kX2hG-6B1eRnW…

  1569. dev.to — MCP tag TIER_1 English(EN) · HyperNexus ·

    面向 AI 智能体的 GitOps:版本控制的工具配置与记忆

    <h1>GitOps for AI Agents: Version-Controlled Tool Configs and Memory</h1> <p>Treat your AI agent's brain like production infrastructure. Learn how GitOps principles, applied to mcp.jsonc configs and agent memory, create auditable, roll-backable, and reliably deployable AI systems…

  1570. Medium — Claude tag TIER_1 English(EN) · Pop123 ·

    Anthropic 的 Claude Opus 5:在 Frontier LLMs 中工程化代理持久性和动态努力

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@Pop123/anthropics-claude-opus-5-engineering-agentic-persistence-and-dynamic-effort-in-frontier-llms-888bdea1bbb5?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/1*…

  1571. dev.to — MCP tag TIER_1 (TL) · Kasi Yaswanth ·

    第 21/30 天:避免代理循环

    <p>I still remember the day our support bot, which was supposed to be a showcase of agentic AI in action, started acting like it was stuck in some kind of bizarre loop. Customers would ask a question, and instead of providing a helpful response, the bot would just repeat the same…

  1572. Towards AI TIER_1 English(EN) · Abinesh U ·

    图工程:工程协调而非更智能的代理

    <h4><em>Reliable agent systems are not defined by the number of agents they contain. They are defined by the contracts that govern what moves between them.</em></h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*JfzIKZJYM1r7qUcUWEgKbg.png" /><figcaption>Graph…

  1573. Towards AI TIER_1 English(EN) · Towards AI Editorial Team ·

    LAI 第136期:通过 Agent 更快地构建,调试其失败,并更可靠地评估它们

    <h4>Better agents, better evaluation, and fewer production surprises.</h4><p>Good morning, AI enthusiasts!</p><p>AI engineering is slowly becoming less about writing prompts and more about building systems that don’t surprise you in production.</p><p>That’s exactly where this wee…

  1574. Towards AI TIER_1 English(EN) · Rohan Mistry ·

    无需克隆的 Git:AI 代理的可持久化、版本化工作区

    <h4>Mount a repo. Write files. Survive crashes. No git clone required.</h4><h3>The Agent State Problem</h3><p>Every AI agent generates files: configs, intermediate artifacts, model outputs, logs. The working state of an agent session lives <em>somewhere</em> on disk. <strong>Wher…

  1575. Medium — MLOps tag TIER_1 English(EN) · jaytank ·

    生产级Agentic执行循环

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://tankjay.medium.com/production-grade-agentic-execution-loops-000da7f08776?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/712/1*iJWhd887YWdPhnbA1Npx2g.png" width="712" /></a></p><p c…

  1576. dev.to — Anthropic tag TIER_1 Русский(RU) · Promptra Team ·

    Claude 智能体与集成功能,在工作任务上进行测试

    <p>Открываешь каталог интеграций и видишь три десятка плиток: Figma, GitHub, n8n, Obsidian, Excel. Из этого как будто следует, что агент уже умеет с ними работать. Это ошибка вывода: наличие строки в списке доказывает ровно то, что кто-то когда-то завёл эту строку в список. Спосо…

  1577. Medium — Claude tag TIER_1 Türkçe(TR) · Mehmet AYDIN ·

    构建Agent系统:今天与明天

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@diabolikss/agent-sistemleri-kurmak-bug%C3%BCn-ve-yar%C4%B1n-80c658dad1ca?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/1*GbBg2Wsx9VresyDDZ4YZZQ.png" width="2752"…

  1578. Medium — Claude tag TIER_1 English(EN) · Mehmet AYDIN ·

    构建智能体系统:今日与未来

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@diabolikss/building-agent-systems-today-and-tomorrow-1cf3b64ffda2?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/1*dm3i7lww12Ehx9zqcCNweA.png" width="2752" /></a>…

  1579. dev.to — MCP tag TIER_1 English(EN) · Product Watch ·

    解锁代理工作流:2026年模型上下文协议(MCP)服务器必备指南

    <h1> Unleashing the Power of MCP </h1> <p>In the rapidly evolving landscape of 2026, the Model Context Protocol (MCP) has emerged as the definitive open standard for bridging the gap between sophisticated large language models (LLMs) and the myriad of data sources that fuel actua…

  1580. dev.to — MCP tag TIER_1 English(EN) · anhmtk ·

    从 Web2 到 Agentic 商业:没人解释的 8 个关键要素,直到你上线

    <p>If you've ever built an e-commerce store, you know the drill: storefront, hosting, payment gateway, inventory, shipping, support, security, and analytics.</p> <p>Miss one — and the whole thing breaks.</p> <p>That framework works for human commerce.</p> <p>But what about commer…

  1581. Towards AI TIER_1 English(EN) · Lorenz Wöhr ·

    Agent Skills: The Composition Cliff

    <h4>Two or three skills lift an agent’s performance; the fourth hits a cliff. The SKILL.md format has no way to stop skills from working against each other.</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*kzzquIN7hBG7waYBEkoGyg.png" /></figure><p>A single …

  1582. dev.to — MCP tag TIER_1 English(EN) · Programming Central ·

    超越API:TypeScript中自主“计算机使用”代理的架构

    <p>The architecture of modern artificial intelligence has reached a critical inflection point. For years, Large Language Models (LLMs) operated as isolated islands of intelligence, restricted to text-in and text-out paradigms, communicating with the external world through strictl…

  1583. dev.to — MCP tag TIER_1 English(EN) · HyperNexus ·

    容器原生AI:掌握GPU直通、内存限制和自动扩展以构建您的智能体基础设施

    <h1>Container-Native AI: Mastering GPU Passthrough, Memory Limits, and Auto-Scaling for Your Agent Infrastructure</h1> <p>Unlock peak performance for your AI agents by mastering container resource management. This guide details Docker AI configurations for GPU passthrough, precis…

  1584. dev.to — MCP tag TIER_1 English(EN) · HyperNexus ·

    GitOps for AI Agents:将基础设施即代码的纪律引入工具配置和内存

    <h1>GitOps for AI Agents: Bringing Infrastructure as Code Discipline to Tool Configs and Memory</h1> <p>Stop treating your AI agent configurations as throwaway artifacts. Learn how applying GitOps principles—PR reviews, CI validation, and version controlled AI configs—creates rel…

  1585. dev.to — MCP tag TIER_1 English(EN) · Manu Shukla ·

    2026年MCP任务:构建长期运行、可恢复的代理工具

    <h1> MCP Tasks in 2026: build long-running, resumable agent tools </h1> <p><strong>Summary.</strong> On 28 July 2026 the Model Context Protocol (MCP) publishes the <code>2026-07-28</code> specification, and one of its two official extensions, Tasks, changes how you build tools th…

  1586. dev.to — Anthropic tag TIER_1 English(EN) · dubleCC ·

    多智能体编排模式:哪些真正有效,哪些会失效

    <blockquote> <p>Originally published at <a href="https://heycc.cn/en/posts/multi-agent-orchestration-patterns/" rel="noopener noreferrer">heycc.cn</a>. This is a mirrored copy — the canonical version is kept up to date at the source.</p> </blockquote> <h1> Multi-Agent Orchestrati…

  1587. dev.to — MCP tag TIER_1 English(EN) · Renato Marinho ·

    将 Claude 打造成 Cursor Cloud Agents 的编排器

    <p>The problem with autonomous agents isn't their ability to write code—it's our inability to manage them at scale.</p> <p>If you've spent any time working with Cursor, you know the feeling of launching a task and then essentially 'hoping for the best.' You trigger an agent, walk…

  1588. Medium — Claude tag TIER_1 English(EN) · Life-is-short--so--enjoy-it ·

    第一集:我如何用Claude构建了一个多代理工程系统

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@life-is-short-so-enjoy-it/episode-1-how-i-built-a-multi-agent-engineering-system-with-claude-bfe5c03b0690?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1536/1*TaZpSO6…

  1589. Mastodon — sigmoid.social TIER_1 Français(FR) · [email protected] ·

    OpenResearcher:一个用于训练大型语言模型(LLM)代理进行长期网络研究的开源管道。96,000个训练轨迹

    OpenResearcher : un pipeline open source pour entraîner des agents LLM à faire de la recherche web longue durée. Les 96 000 trajectoires d'entraînement ont été générées sans appeler d'APIs externes. Dataset, modèle 30B et recette d'entraînement inclus. ⬇️ https:// github.com/TIGE…

  1590. dev.to — MCP tag TIER_1 English(EN) · Omnithium ·

    Agent-to-Agent Communication Protocols: Architecting for a Multi-Protocol Future

    <h2> The Protocol Vacuum: Why Multi-Agent Systems Stall in Production </h2> <p>Platform teams must treat inter-agent communication as a first-class architectural concern. Build abstraction layers now, before the protocol landscape solidifies, to avoid lock-in and costly rework. T…

  1591. Towards AI TIER_1 English(EN) · David Pradeep ·

    成本优化的代理架构:多代理系统的战略模型选择与缓存

    <p>The first time I stared at a cloud bill after deploying a fleet of AI agents, the numbers felt like a punchline, my “experiment” had turned into an unexpected expense. I’d spent weeks tuning prompts, wiring up tool calls, and watching latency drop, but the cost column kept spi…

  1592. Medium — MCP tag TIER_1 English(EN) · Prasanna ·

    Action Cassettes:为什么确定性回放是AI浏览器代理中缺失的一层

    <div class="medium-feed-item"><p class="medium-feed-snippet">Record once. Replay forever. Heal when the site changes.</p><p class="medium-feed-link"><a href="https://medium.com/@prasannapal273/action-cassettes-why-deterministic-replay-is-the-missing-layer-in-ai-browser-agents-19f…

  1593. dev.to — MCP tag TIER_1 English(EN) · Guy ·

    超越单智能体天花板:通过 MCP 智能体团队进行横向扩展

    <p>Most people's experience with AI is a conversation with one assistant. ChatGPT, Claude, and similar products present one conversational partner. You ask it a question, it reasons, perhaps calls a few tools, and gives you an answer.</p> <p>That experience creates a natural arch…

  1594. Medium — Claude tag TIER_1 English(EN) · Deepak Damodaran ·

    Claude Architect #3:Claude背后的隐藏引擎——工具、MCP和Agent Actions详解

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@deepakatl1981/claude-architect-3-the-hidden-engine-behind-claude-tools-mcp-and-agent-actions-explained-c46ee5d8a3fe?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1635…

  1595. Medium — Claude tag TIER_1 English(EN) · Andriy Tretyak ·

    Swarmery:我如何停止在项目之间复制粘贴代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://swarmery.medium.com/swarmery-how-i-stopped-copy-pasting-agents-between-projects-04c54aab9c3a?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1280/1*6xj1wbsHyPBmfX-dBKEXBw.jpeg" wid…

  1596. Medium — MLOps tag TIER_1 English(EN) · Sunil Tailor ·

    Agentic Data Engineering Begins with Determinism — Part 1

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://sunil-tailor.medium.com/agentic-data-engineering-begins-with-determinism-part-1-30aa1d199dc8?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/807/1*1IEH4UjLqJ7atJ5YLR3Hfw.png" width=…

  1597. Towards AI TIER_1 English(EN) · synovergetechnologies ·

    在 Microsoft Azure 上构建生产就绪的 Agentic RAG 系统

    <p>Large language models have become remarkably capable, but many enterprise AI projects still struggle, not because of the model, but because of the system around it.</p><p>The challenge isn’t generating better responses. It’s building applications that reliably retrieve the rig…

  1598. Medium — MLOps tag TIER_1 English(EN) · Pratyaksh Singh ·

    使用 Amazon Bedrock 构建自主任务自动化代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@pratyakshsingh11/building-an-autonomous-task-automation-agent-with-amazon-bedrock-623dca34ea6b?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/600/1*frNgcnyAzTojBeIJiOqC…

  1599. dev.to — MCP tag TIER_1 English(EN) · Victor García ·

    MCP 接口:代理如何在无需 curl 的情况下与后端通信

    <p>An agent skill runs <code>curl http://127.0.0.1:7200/api/v1/notes</code> from inside its Docker sandbox. It fails with exit code 7 — "couldn't connect" — before authentication even runs, because the sandbox is launched with <code>network: none</code>. There is no loopback to r…

  1600. Medium — MLOps tag TIER_1 English(EN) · Michel Alan López ·

    AI Agent 架构:从输入到智能行动

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ingalopez11/ai-agent-architecture-from-input-to-intelligent-action-ed86cdcd7a10?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1121/1*zasE28OSUTPb3dRNW7465g.png" width=…

  1601. dev.to — MCP tag TIER_1 English(EN) · shakti mishra ·

    MCP 与 Agent Skills:上下文工程的决策框架

    <h2> MCP vs. Agent Skills: What's the Difference and Which Do You Need? </h2> <p>As AI agents evolve beyond basic chat interfaces into fully autonomous systems, developers keep hitting the same architectural decision: how do we actually extend what an agent can do?</p> <p>Two con…

  1602. Towards AI TIER_1 English(EN) · MongoDB ·

    当你的代理失联时:使用 OpenTelemetry 实现多代理系统的可观测性

    <p><em>Written by </em><a href="https://www.linkedin.com/in/matteo-rossi-280391/"><strong><em>Matteo Rossi</em></strong></a><em>.</em></p><p>In recent months, AI applications have radically evolved. Earlier, prototypes looked like a single loop: prompt, model call, optional tool …

  1603. Medium — Claude tag TIER_1 English(EN) · Neha Patel ·

    第二周反思:用 Agentic AI 构建更智能的系统

    <div class="medium-feed-item"><p class="medium-feed-snippet">Learning, Experimenting, and Understanding the Future of AI Agents</p><p class="medium-feed-link"><a href="https://medium.com/@workwithneha/week-2-reflection-building-smarter-systems-with-agentic-ai-b79801c7a108?source=…

  1604. Medium — AI coding tag TIER_1 English(EN) · Aviv Carmi ·

    Agentic Control for Software Engineers

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@avivcarmis/agentic-control-for-software-engineers-c4a700992463?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1000/0*ctg74w0eNP7Var-0.jpg" width="1000" /></a></p><p…

  1605. Medium — AI coding tag TIER_1 English(EN) · The Review Surface ·

    一个将人类审查纳入循环的代理工作流

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@thereviewsurface/one-agent-workflow-that-keeps-human-review-in-the-loop-43bf5a586901?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1200/1*L_FP2TMt468k1I6809hshg.pn…

  1606. Towards AI TIER_1 Nederlands(NL) · ML Point ·

    Agent Harness Engineering vs. Loop Engineering vs. Graph Engineering

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/agent-harness-engineering-vs-loop-engineering-vs-graph-engineering-02690996d485?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/999/1*hOw3GsK8Gwif8NUWdM_L7w…

  1607. Towards AI TIER_1 English(EN) · Sandip Palit ·

    构建智能反馈系统:深入探讨条件式代理工作流与…

    <h3>Building Intelligent Feedback Systems: A Deep Dive into Conditional Agentic Workflows with LangGraph</h3><p>The landscape of Artificial Intelligence has shifted dramatically over the past couple of years. We are no longer simply chatting with isolated Large Language Models (L…

  1608. Medium — MCP tag TIER_1 English(EN) · Kovilur Gopala Krishnan ·

    设计Agentic Enterprise-5架构决策

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@bigbrochush/designing-an-agentic-enterprise-5-architectural-decisions-3b12aee4cb70?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1536/1*HXUa42FClak-beILjol1Jw.png" width…

  1609. Medium — Claude tag TIER_1 English(EN) · Özcan Kara ·

    Agent技能介绍

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ozcankaraa/introduction-to-agent-skills-70a09c9d2a32?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/811/1*qYF00XMfKLEIj9_KfyvZfg.png" width="811" /></a></p><p class="m…

  1610. Medium — Claude tag TIER_1 English(EN) · Haowen Huang ·

    构建多智能体量化回测系统:Amazon Bedrock AgentCore + Strands Agents

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@popkee/building-a-multi-agent-quant-backtesting-system-amazon-bedrock-agentcore-strands-agents-559e752917da?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2304/1*jLL9O…

  1611. dev.to — MCP tag TIER_1 English(EN) · Michael "Mike" K. Saleme ·

    两次智能体-工具攻击,一个教训:检测有上限,强制执行权限有下限

    <p>Two agent-tool security papers landed in June. Read together, they expose the boundary between semantic detection and enforceable control.</p> <p>A common response to malicious agent tools is to scan tool descriptions: inspect the text an agent is about to trust, decide whethe…

  1612. Towards AI TIER_1 English(EN) · Zoumana Keita ·

    自然语言处理的演进:从TF-IDF到Agent

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/evolution-of-nlp-tf-idf-to-agents-e08c9da95174?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/2560/1*rkIHMEjpB_LXpaw2r7o0Ew.png" width="2560" /></a></p><p …

  1613. Towards AI TIER_1 English(EN) · David Pradeep ·

    迁移至 Agent-First 架构以增强安全性

    <p>The first time I tried to migrate a legacy order-processing service to an agent-first model, the biggest surprise wasn’t the refactoring effort, it was how many hidden security gaps opened up the moment autonomous agents started calling external APIs. The stakes of securing ag…

  1614. Towards AI TIER_1 English(EN) · Towards AI Editorial Team ·

    TAI 第213期:新前沿竞争者浪潮与多智能体爆发

    <h4>Also GPT-5.6, Grok 4.5, Muse Spark 1.1, GPT-Realtime-2.1, and more.</h4><figure><a href="https://academy.towardsai.net/courses/python-for-genai?utm_source=Newsletter&amp;utm_medium=email&amp;utm_id=header"><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*vckNXOgtvN…

  1615. Medium — Claude tag TIER_1 English(EN) · Akshat A. Mistry ·

    第四部分 | 多智能体协调:中心辐射式架构

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@akshat.mistry/part-iv-multi-agent-coordination-the-hub-and-spoke-architecture-529ad62d10c7?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/600/0*TYkUdqoinXHVwINm.png" w…

  1616. Towards AI TIER_1 English(EN) · Satish Kumar ·

    从零开始构建ArcticSwarm:一个生产级的多智能体深度研究系统

    <h4><em>Implementing Snowflake’s ArcticSwarm architecture with Python, Redis, and free-tier LLMs</em></h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*KbIfOVoZHIHibkvhCn1ygg.png" /><figcaption><em>ArcticSwarm architecture: specialized agents coordinated thr…

  1617. Medium — Claude tag TIER_1 English(EN) · Mustapha Aitigunaoun ·

    Agent Skills 介绍

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://osintteam.blog/introduction-to-agent-skills-e6b136967970?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/0*y2VOX0bawcSPnxH9" width="7680" /></a></p><p class="medium-feed-snipp…

  1618. Medium — MLOps tag TIER_1 English(EN) · Mirav Kapadia ·

    为多步智能体工程化可靠性的 8 个杠杆

    <div class="medium-feed-item"><p class="medium-feed-snippet">Compound Interest in Reverse</p><p class="medium-feed-link"><a href="https://medium.com/@miravck/8-levers-for-engineering-reliability-into-multi-step-agents-4b15752bdd2b?source=rss------mlops-5">Continue reading on Medi…

  1619. Towards AI TIER_1 English(EN) · Shravya ·

    在 JVM 上构建企业级多智能体系统 - 使用 Koog、MCP 和 A2A 的分层架构

    <p>Most AI agent tutorials show a single agent calling a few tools. That works for demos. It falls apart the moment a real enterprise system needs ten specialized agents, each with its own tools, coordinating to handle a complex request, all running inside infrastructure that was…

  1620. Medium — Claude tag TIER_1 English(EN) · Vinod Bellary ·

    Agentic Architecture & Orchestration — Agentic Loops

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@vinod.bellary/agentic-architecture-orchestration-agentic-loops-a3be80399f55?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/845/1*j2Xiq3Sam96z1nRw3skorg.png" width="845…

  1621. Medium — Claude tag TIER_1 English(EN) · Akshat A. Mistry ·

    第三部分 | ‘tool_use’,端到端:三次迭代中的智能体循环

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@akshat.mistry/tool-use-end-to-end-the-agentic-loop-in-three-iterations-87cde7765dfc?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/600/1*WQ4cyTXW2ocJiLOMUTx4Ow.png" wi…

  1622. dev.to — MCP tag TIER_1 English(EN) · Sayan Mohsin ·

    终结前端:构建原生智能体技术栈(第一部分)

    <p>For the last two decades, software engineering has followed a predictable formula: build a database, write an API, and build a massive, complex frontend web app (React, Vue, Next.js) so a human can interact with your data. </p> <p>If you are building something like an Order Ma…

  1623. dev.to — MCP tag TIER_1 English(EN) · Renato Marinho ·

    上下文窗口的无声杀手:为什么 Token 估算正在拖累你的代理

    <p>If you are building LLM-powered agents, you have likely run into the 'context wall.' You send a massive payload of documentation or history to Claude or GPT-4o, and suddenly the model starts hallucinating, truncating mid-sentence, or—even worse—throwing an API error because yo…

  1624. Towards AI TIER_1 English(EN) · Vikram Bhat ·

    真正需要多个智能体的多智能体系统

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/multi-agent-systems-that-actually-need-multiple-agents-10240a75081f?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1996/1*Aco87EtH-is3CawwmmX_uw.png" width…

  1625. Towards AI TIER_1 English(EN) · Pranav Dhopey ·

    一行代码,任意模型:通过 LiteLLM 在 Google ADK 中实现多提供商代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/one-line-any-model-multi-provider-agents-in-google-adk-via-litellm-fe88acee24d8?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/2600/1*9cKejwdyhcao7-0sturvw…

  1626. Towards AI TIER_1 English(EN) · Harish Ramkumar ·

    Bedrock Agents 的上下文工程:超越提示工程的实践指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/context-engineering-for-bedrock-agents-a-hands-on-guide-beyond-prompt-engineering-d92aad36a839?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/2600/0*V1Lt2-…

  1627. Medium — Claude tag TIER_1 English(EN) · Nithin ·

    构建Agent Harness:自主循环背后的基础设施

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@nithinellanki/building-the-agent-harness-the-infrastructure-behind-autonomous-loops-d8add3d8c61c?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/1*KGeER-o4t8TpiUTU…

  1628. Towards AI TIER_1 English(EN) · Yuval Mehta ·

    奖励设计是难点:为使用工具的智能体构建可验证的奖励

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*ZO5F388PeA4XSiul" /><figcaption>Photo by <a href="https://unsplash.com/@toddquackenbush?utm_source=medium&amp;utm_medium=referral">Todd Quackenbush</a> on <a href="https://unsplash.com?utm_source=medium&amp;utm_m…

  1629. dev.to — MCP tag TIER_1 English(EN) · Edison Flores ·

    Re: 下载量是虚荣 — 为代理构建可观察的安装路径

    <p>Great point <a class="mentioned-user" href="https://dev.to/alexshev">@alexshev</a> — downloads ARE a vanity metric. The real signal is: did the agent connect, complete a workflow, and self-diagnose failures?</p> <p>We are building toward exactly that. Currently exposed:</p> <u…

  1630. Towards AI TIER_1 English(EN) · Sandip Palit ·

    揭秘顺序代理工作流:LangGraph、State…的理论基础

    <h3>Demystifying Sequential Agentic Workflows: The Theoretical Foundations of LangGraph, State Management, and High-Speed Inference</h3><p>The landscape of Artificial Intelligence is undergoing a massive paradigm shift. Just a year ago, the industry was heavily fixated on single-…

  1631. Medium — MLOps tag TIER_1 English(EN) · Vimal Dwarampudi ·

    构建智能AIOps:使用LangGraph、Gemini和原生图数据库构建的闭环平台

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://vimal-dwarampudi.medium.com/building-agentic-aiops-a-closed-loop-platform-with-langgraph-gemini-and-a-native-graph-database-6cb2936be385?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/…

  1632. dev.to — MCP tag TIER_1 English(EN) · correctover ·

    CWE-636:每个主要Agent框架中的隐形“终止开关”

    <h1> CWE-636: The Silent Kill Switch in Every Major Agent Framework </h1> <h2> How observer-pattern hooks create a systemic fail-open vulnerability that lets governance be bypassed — and what to do about it </h2> <h2> The Vulnerability in One Paragraph </h2> <p>Every major AI age…

  1633. Medium — Claude tag TIER_1 English(EN) · Axion ·

    使用 Claude + MCP + Axion 构建市场研究代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@axionquant/building-a-market-research-agent-with-claude-mcp-axion-6ca2671409c1?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/1*ZtC13Wf7_4AhUfzBs0hINQ.png" width=…

  1634. Medium — Claude tag TIER_1 English(EN) · Code Coup ·

    Claude 的多智能体系统:为何它比运行 Opus 处理一切要便宜得多

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/coding-nexus/claudes-multi-agent-system-why-it-s-much-cheaper-than-running-opus-for-everything-c1adf440b6f1?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1475/1*whYHRO…

  1635. Towards AI TIER_1 English(EN) · Kashif Mehmood ·

    Qwen-AgentWorld:被训练成环境而非智能体并击败Opus的模型

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/qwen-agentworld-the-model-trained-to-be-the-environment-not-the-agent-and-beats-opus-5d41f3366415?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/996/1*BFml…

  1636. dev.to — MCP tag TIER_1 English(EN) · Renato Marinho ·

    可观测性鸿沟:使用 MCP 管理多智能体群

    <p>You've probably been there. You configure a complex multi-agent topology in AutoGen Studio, trigger a run, and then... nothing happens. Or worse, it keeps running for twenty minutes, burning tokens while two agents argue over an invisible syntax error in a Python skill you can…

  1637. Towards AI TIER_1 English(EN) · Roberto Penco ·

    Agentic Engineering:用自然语言编程的古老梦想终于实现——并且正在成为……

    <h3><strong>Agentic Engineering: The Old Dream of Programming in Natural Language Is Finally Here —and Becoming Computer Science Again</strong></h3><p>Roberto Penco, PhD</p><p>June 2026</p><h3><strong>Introduction</strong></h3><p>I recently completed my PhD in computer science an…

  1638. Medium — AI coding tag TIER_1 English(EN) · Okan Aslan ·

    Agentic Software Development 的有界委托

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://aslanokan.medium.com/bounded-delegation-for-agentic-software-development-c4a8ba251b73?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/2600/1*-cFwTMFax7r_pxnB96v7nw.png" width="3…

  1639. Medium — Claude tag TIER_1 English(EN) · Jinyan Su ·

    智能体(Agent)的演进:从上下文工程到长期运行的框架

    <div class="medium-feed-item"><p class="medium-feed-snippet">Over the past few years, the main thread of progress in large models has mostly revolved around &#x201c;the model itself&#x201d;: parameters, data&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@jiny…

  1640. dev.to — MCP tag TIER_1 ไทย(TH) · Thanawat Wongchai ·

    教命令行界面与代理进行通信

    <p>นี่คือบทความ 10 ตอนที่แชร์ว่า Apidog พัฒนา <a href="https://apidog.com/apidog-cli/?utm_source=dev.to&amp;utm_medium=wanda&amp;utm_content=n8n-post-automation">Apidog CLI</a> ซึ่งเป็นเครื่องมือบรรทัดคำสั่งสำหรับการทดสอบ API และการจัดการวงจรชีวิต API ได้อย่างไร อ่านตามลำดับหรือข…

  1641. dev.to — MCP tag TIER_1 English(EN) · Kartik Anand ·

    # 在 Snowflake 和 Microsoft Fabric 上构建多智能体 A2A 架构 — 无需替换任一平台

    <p>Every enterprise healthcare payer I work with has the same problem.</p> <p>They have years of investment in Snowflake — semantic models, claims analytics, carefully curated data products. They have Microsoft Fabric rolling out across their organization — lakehouses, Delta tabl…

  1642. dev.to — MCP tag TIER_1 English(EN) · Christopher Lyon ·

    我解决的快速启动数据库以实现快速 Agentic 开发的方法

    <p>I built TmpState because I kept running into the same stupid problem.</p> <p>My coding agent could build most of an app. It could write the React<br /> components, add the API route, sketch out the data model, and even tell me<br /> what collections it wanted.</p> <p>Then it n…

  1643. Towards AI TIER_1 English(EN) · Armin Norouzi, Ph.D ·

    容错代理管道:检查点、重试和补偿

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/fault-tolerant-agent-pipelines-checkpoint-retry-and-compensate-8870ed221c26?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1000/1*d_zGHHumUpoB9bbnbAOaMA.pn…

  1644. dev.to — MCP tag TIER_1 English(EN) · curatedmcp ·

    Firecrawl MCP:AI代理的网络抓取和自主研究

    <blockquote> <p><em>Install guide and config at <a href="https://www.curatedmcp.com/install/firecrawl-mcp/claude-desktop" rel="noopener noreferrer">curatedmcp.com</a></em></p> </blockquote> <h1> Firecrawl MCP: Web scraping and autonomous research for AI agents </h1> <p>Web scrapi…

  1645. Medium — MLOps tag TIER_1 English(EN) · Loknath Baskar ·

    Omnigent:Databricks 的新 Meta-Harness 在解决 Agent 蔓延问题上有什么独到之处

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@loknathbaskar/omnigent-what-databricks-new-meta-harness-gets-right-about-the-agent-sprawl-problem-a76a89d28dda?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/2600/1*bNv…

  1646. dev.to — MCP tag TIER_1 English(EN) · Edison Flores ·

    为代理商务构建信任层:8,560个MCP技能,x402,AP2指令

    <p>When Anthropic donated MCP to the Linux Foundation in December 2025, discovery was solved. But trust was not.</p> <p>An independent analysis found ~64.7 million server entries from just 1,691 unique packages — massive duplication, zero signal, and active supply-chain attacks (…

  1647. dev.to — MCP tag TIER_1 English(EN) · Nikhil raman K ·

    代理式BFSI系统中MCP与A2A:完整实施指南

    <p>Banking has a protocol problem.</p> <p>A risk analyst at a tier-one bank submits a credit decision request. The answer requires querying the core banking system, pulling transaction history from the data warehouse, checking the sanctions database, retrieving the customer's KYC…

  1648. Medium — fine-tuning tag TIER_1 English(EN) · Vansh ·

    Braid:用于测试时搜索和代理集成的无损跨分支计算共享

    <div class="medium-feed-item"><p class="medium-feed-snippet">Vansh Verma</p><p class="medium-feed-link"><a href="https://medium.com/@vanshverma.dev/braid-lossless-cross-branch-computation-sharing-for-test-time-search-and-agent-ensembles-7ddad1878c92?source=rss------fine_tuning-5"…

  1649. Medium — AI coding tag TIER_1 ไทย(TH) · iFew ·

    Loop Engineering:用 3 个 AI 代理来实际创造产品

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://ifew.medium.com/loop-engineering-3-%E0%B8%A5%E0%B8%B9%E0%B8%9B%E0%B8%97%E0%B8%B5%E0%B9%88%E0%B8%97%E0%B8%B3%E0%B8%A3%E0%B9%88%E0%B8%A7%E0%B8%A1%E0%B8%81%E0%B8%B1%E0%B8%9A-ai-agent-%E0%B9%80%E0%B8%9E%E0%B8…

  1650. Medium — Claude tag TIER_1 English(EN) · Meghana Harishankara ·

    为什么大多数Agent项目在模型成为瓶颈之前就失败了

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@meghanaharishankara/why-most-agent-projects-fail-before-the-model-becomes-the-bottleneck-635d12c85159?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1632/1*XenVlyhHqK3…

  1651. Medium — Anthropic tag TIER_1 English(EN) · Harnish Savsani ·

    重塑领域一:Agentic架构与编排

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://harnishsavsani.medium.com/crushing-domain-1-agentic-architecture-orchestration-79e93cb53d16?source=rss------anthropic-5"><img src="https://cdn-images-1.medium.com/max/2600/1*T-Wx4hkmpKOJOjzN7aGayw.png" wi…

  1652. The Register — AI TIER_1 English(EN) ·

    AI 代理:数据库蔓延的原因。以及提议的解决方案

    DB wrangling tech needs to meet demands of AI agents, Cockroach Labs CEO Spencer Kimball tells El Reg

  1653. dev.to — MCP tag TIER_1 English(EN) · Renato Marinho ·

    硬编码模型提示的终结:构建能够发现自身基础设施的代理

    <p>I was reading a thread recently about how MCP servers are burning 50k+ tokens before a user even types a single word, and it hit home. We're all obsessed with the 'intelligence' of these models, but we're ignoring the massive architectural debt we're creating by hardcoding too…

  1654. Medium — MCP tag TIER_1 English(EN) · Fuji Nguyen ·

    使用 Blazor United 和 .NET 10 的 AI Agent UI — 系列前言

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/scrum-and-coke/ai-agent-ui-with-blazor-united-net-10-series-preface-2915c25fe566?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1536/1*ZBgClCIdeXEdfnrkhnLbEg.png" width="1…

  1655. Medium — AI coding tag TIER_1 English(EN) · heavendai ·

    Agent-as-a-Router:当模型路由学会进化

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@mingyang.heaven/agent-as-a-router-when-model-routing-learns-to-evolve-e6e96d2ef250?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/2600/1*bq828aQgRbuSyXCLYl5c8A.png"…

  1656. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    AI助手轮询代理模式实用指南——调度器、队列、Webhook、持久化工作流、状态管理及生产环境的权衡

    A practical guide to polling agent patterns in AI assistants — schedulers, queues, webhooks, durable workflows, state management, and tradeoffs for production systems. # Hermes # OpenClaw # Architecture # LLM # AI # AI Coding # Dev # DevOps https://www. glukhov.org/ai-systems/arc…

  1657. dev.to — MCP tag TIER_1 English(EN) · sarathi s ·

    Vidilearn: 面向大型语言模型、智能体和 MCP 服务器的 AI 知识摄取与检索网关

    <p>Just realized something important while building Vidilearn.</p> <p>It’s not just a “YouTube transcript extractor.”</p> <p>Vidilearn is evolving into an AI knowledge ingestion + retrieval gateway for:</p> <ul> <li>LLMs</li> <li>AI agents</li> <li>MCP servers</li> <li>RAG pipeli…

  1658. dev.to — MCP tag TIER_1 English(EN) · Anya Summers ·

    Agent Communication Matrix:MCP、A2A 和 Plain REST 各显神通

    <h2> Key Takeaways </h2> <ul> <li> <strong>Agent communication has three problems, not just one.</strong> Tool access, peer coordination, and system integration each need a different solution. Most production failures occur when one protocol tries to cover all three.</li> <li> <s…

  1659. Towards AI TIER_1 English(EN) · Anna Jey ·

    Gemini Spark 工作流:构建者如何设计永不离线且不惹用户厌烦的 AI 代理

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*UqnyF-dwIQQTur4YdgxFiQ.jpeg" /><figcaption>Gemini Spark Workflow</figcaption></figure><p>An always-on AI agent sounds useful until it interrupts at the wrong time, acts on an old instruction, or quietly touches d…

  1660. Towards AI TIER_1 English(EN) · Devashish Datt Mamgain ·

    AI Agent Orchestration:如何在客户支持中路由、调用工具和交接

    <h3>What is AI agent orchestration?</h3><p><strong>AI agent orchestration</strong> coordinates several specialized AI agents so they operate as one system working toward a single goal. Instead of asking one general-purpose model to handle everything, it gives each agent a narrow …

  1661. Medium — Claude tag TIER_1 English(EN) · Halil Yılmaz ·

    CLAUDE CODE — HOOKS | AI 时代团队管理 -2

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@haliilylmaaz/claude-code-hooks-team-management-in-the-ai-era-2-3b2267dc3910?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1293/1*6dGMWPOHeKqg0g7p2l5iuA.png" width="12…

  1662. Towards AI TIER_1 Dansk(DA) · DhanushKumar ·

    SkillOpt:自主进化智能体技能的执行策略

    <p><em>SkillOpt</em> is a<strong> novel framework </strong>that optimizes an AI agent’s <em>skill</em> — a compact natural-language policy document — rather than its weights. It treats the skill text as a <strong>trainable parameter</strong>: a <strong>frozen “target” model repea…

  1663. dev.to — MCP tag TIER_1 English(EN) · Intellibooks AI ·

    Intellibooks 详解 Agent Development Kit:构建生产级 AI Agent 的完整框架

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy1obhr6mlann1zjoiw6t.jpg"><img alt=" " height="1200"…

  1664. Towards AI TIER_1 English(EN) · Satish Kumar ·

    我构建了四个基于语义层的Cortex代理——治理功能实际存在于此

    <h4><em>Part 3 of a 3-part series on implementing Snowflake Horizon Context in production</em></h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*3Q135MpMQHrKRGVpzbLTmQ.png" /></figure><p>One week before we shipped this, an early prototype agent almost put a …

  1665. dev.to — MCP tag TIER_1 English(EN) · Athreix ·

    Angle:新的 Agentic Resource Discovery 标准,为构建真实系统的人们而解释 · 权威 + 证据

    <p><strong>TL;DR:</strong> Google, Microsoft, GitHub, Hugging Face, Nvidia and Salesforce backed a draft spec called Agentic Resource Discovery (ARD). It lets AI agents find and connect to tools and other agents at runtime instead of someone hard-wiring every integration. Most bu…

  1666. Medium — Claude tag TIER_1 English(EN) · Mohit Verma ·

    在原生 JavaScript 中编排构建 AI 代理

    <div class="medium-feed-item"><p class="medium-feed-snippet">What Even Is an Agent?</p><p class="medium-feed-link"><a href="https://codeonmars.medium.com/orchestrating-building-ai-agents-in-vanilla-js-a32e84602352?source=rss------claude-5">Continue reading on Medium »</a></p></di…

  1667. dev.to — MCP tag TIER_1 English(EN) · AK DevCraft ·

    下一代迭代改进:使用 Llama.cpp、Gemma 4 12B 和 MCP 优化个人代理 AI 助手

    <h2> Background </h2> <p>Building a $0 personal agentic AI assistant means you don't have the luxury of infinite cloud scale. You can't just throw a massive 128k context window at a lazy system prompt and call it a day. When every unnecessary token impacts limited CPU cores or th…

  1668. Medium — MCP tag TIER_1 English(EN) · Manjunath Venkobarao ·

    技能与MCP:如何构建真正可扩展的Agent能力

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://levelup.gitconnected.com/skills-and-mcp-how-to-build-agent-capabilities-that-actually-scale-1eafff1de5e4?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1096/1*MBR9cuuQmaAFBN2ri9IuZQ.…

  1669. Medium — MLOps tag TIER_1 English(EN) · Lina Faik ·

    谷歌ADK详解:使用谷歌的Agent Development Kit构建多智能体系统

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://linafaik.medium.com/google-adk-explained-building-multi-agent-systems-with-googles-agent-development-kit-6e09fe01b77f?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1200/0*af0a_ZjF…

  1670. dev.to — MCP tag TIER_1 English(EN) · Intellibooks AI ·

    Intellibooks AI Agents 开发流程:构建生产级 AI Agents 的完整指南

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fga4fu1w9bkp0codgglxt.jpg"><img alt=" " height="1200"…

  1671. Medium — fine-tuning tag TIER_1 English(EN) · Sandeep Sharma ·

    针对领域特定生成式AI项目的LLM微调

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://sid-sharma1990.medium.com/fine-tuning-llms-for-domain-specific-gen-ai-projects-e66e08d3bc9d?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1536/1*bTA1vY5QbWmdxiBKrUJNbw.png" …

  1672. Towards AI TIER_1 English(EN) · Faheem Munshi ·

    AI赋能客户沟通:精准而温暖地处理客户全生命周期——…

    <h3>AI for Client Communication: The entire client lifecycle, handled with precision and warmth — Prompt to Profit · Day 23 of 30</h3><h4><em>From the first enquiry to the final invoice — how to use AI to communicate at a professional level that builds trust, not suspicion.</em><…

  1673. dev.to — MCP tag TIER_1 English(EN) · 강해수 ·

    D1 Schema 迁移与 AI 代理:扼杀零停机部署的事务内 DDL 陷阱

    <p>Running an AI agent to execute your D1 migrations will silently wreck your database — unless you explicitly forbid it from wrapping DDL in a transaction.</p> <p>Claude Code, when handed a migration task, defaults to wrapping everything in <code>BEGIN TRANSACTION / COMMIT</code…

  1674. Towards AI TIER_1 English(EN) · Satish Kumar ·

    为什么企业级AI需要一个受管制的意义层:隆重推出 Snowflake Horizon Context

    <h4><em>Part 1 of series on implementing Snowflake Horizon Context in production</em></h4><h3>The Three Revenue Numbers Problem</h3><p>It’s quarterly business review day. The CEO asks a straightforward question: <em>“What was our Q3 revenue?”</em></p><p>Finance reports <strong>$1…

  1675. Towards AI TIER_1 English(EN) · Raj kumar ·

    使用 Docker 和 FastAPI 构建生产就绪的 Agentic AI 系统

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/building-production-ready-agentic-ai-systems-with-docker-and-fastapi-b4c2231b3945?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1536/1*15z5N4m58t-64hJTqr7…

  1676. dev.to — MCP tag TIER_1 English(EN) · SandBase AI ·

    我们绘制了 500 个 AI Agent 基础设施项目图

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fscduhkymymm86t4h6uc4.png"><img alt="500 AI Agent Inf…

  1677. Towards AI TIER_1 English(EN) · Tarun Agarwal ·

    使用 Claude 的网络搜索工具构建 Slack AI 代理:端到端演练

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/building-a-slack-ai-agent-with-claudes-web-search-tool-an-end-to-end-walkthrough-4d4c97854660?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/2062/1*WXZJmke…

  1678. dev.to — MCP tag TIER_1 English(EN) · Intellibooks AI ·

    IntelliBooks AI 演进时间线:从规则系统到自主代理式 AI

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2uda4d688ewm70pb13pu.jpg"><img alt=" " height="1200"…

  1679. Medium — Claude tag TIER_1 English(EN) · Halil Yılmaz ·

    CLAUDE CODE — MCP | AI时代团队管理 -1

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@haliilylmaaz/claude-code-mcp-team-management-in-the-ai-era-1-efc0e768a56b?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1293/1*AaleUJYak08FiF7ClguR_w.png" width="1293…

  1680. Medium — MLOps tag TIER_1 English(EN) · Rashmi ·

    Claude Code for MLOps and LLMOps:使用自主工程构建生产级AI系统

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://blog.gopenai.com/claude-code-for-mlops-and-llmops-building-production-grade-ai-systems-with-autonomous-engineering-ef49b815289d?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/600/1…

  1681. dev.to — MCP tag TIER_1 English(EN) · Ravi Kiran Kadaboina ·

    AI代理的PRG模式:一项25年前的解决方案在新时代走向成熟

    <p>Since the 90s a classic bug always plagued web forms. You've probably seen it — the browser warning that says <em>"Resubmitting this form will repeat the action."</em> Your user placed an order, hit refresh, and now there are two orders. Or two emails. Or two charges.</p> <p>T…

  1682. The Register — AI TIER_1 English(EN) ·

    CPU在代理式AI基础设施中日益增长的作用

    PARTNER CONTENT: As agentic AI systems scale across cloud and datacenter environments, CPUs remain the control plane coordinating performance and efficiency.

  1683. dev.to — MCP tag TIER_1 English(EN) · Ahmad Shakir ·

    Show Dev:Weavz — 为 AI 代理提供受管应用程序访问

    <p>Weavz gives AI agents and SaaS products governed access to the apps people already use. Connect 1,000+ integrations, expose approved actions through MCP or APIs, add Human Gates for sensitive work, and keep scoped state, files, and audit trails. Provision workspaces, add users…

  1684. Medium — Anthropic tag TIER_1 English(EN) · Ramakrishna Sanikommu ·

    语义/上下文层:将代理式AI植根于企业真相

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ramakrishna.sanikommu/the-semantic-context-layer-grounding-agentic-ai-in-enterprise-truth-6c31226b227c?source=rss------anthropic-5"><img src="https://cdn-images-1.medium.com/max/1600/1*XzY1xRo…

  1685. Medium — Claude tag TIER_1 Português(PT) · Gustavo Tavares ·

    支持多语言AI代理实时翻译:LangGraph的完整指南及...

    <div class="medium-feed-item"><p class="medium-feed-snippet">Vivemos em um momento de transforma&#xe7;&#xe3;o sem precedentes na intelig&#xea;ncia artificial. Os agentes de IA evolu&#xed;ram de simples chatbots baseados&#x2026;</p><p class="medium-feed-link"><a href="https://medi…

  1686. Medium — Claude tag TIER_1 English(EN) · TarrantRo ·

    Claude 辅助开发缺失手册

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://blog.stackademic.com/ai-coding-a-practical-guide-for-engineers-to-u-626ae4a242eb?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2100/1*-UsKW7jp0YsSFaArW9ZJmw.avif" width="2100" />…

  1687. Medium — Claude tag TIER_1 English(EN) · Facundo hannoch ·

    Agents Declaration — 设计一个 Agents Orchestration Library — 第二部分

    <div class="medium-feed-item"><p class="medium-feed-snippet">It&#x2019;s easy to spawn 4 agents if they are all a thread and a subprocess in the host. But I want to show you something more sophisticated</p><p class="medium-feed-link"><a href="https://medium.com/@facuhannoch/agent…

  1688. Towards AI TIER_1 English(EN) · Bessie Delight Kekeli ·

    改进我们的 LangGraph Agent 以应对现实世界的电子商务:企业验证、业务逻辑…

    <h3>Improving Our LangGraph Agent for Real-World E-Commerce: Enterprise Validation, Business Logic Guards, and a Multi-Agent Architecture</h3><h4>The patterns that separate a LangGraph demo from a system you can actually deploy.</h4><p><em>The article </em><a href="https://medium…

  1689. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用AI代理!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  1690. dev.to — MCP tag TIER_1 English(EN) · EvanLin | Contorium ·

    超越上下文窗口:为 AI 开发构建项目智能层

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F468po9669n24912729tf.png"><img alt=" " height="800" …

  1691. dev.to — MCP tag TIER_1 English(EN) · EvanLin | Contorium ·

    超越上下文窗口:为 AI 开发构建项目智能层

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxvgkmh50ngi5y2uekc2k.png"><img alt=" " height="533" …

  1692. dev.to — MCP tag TIER_1 English(EN) · Renato Marinho ·

    超越API:自主代理需要协议层

    <p>If you’re building anything serious with AI—something that moves beyond generating boilerplate text or summarizing blog posts—you quickly run into the same problem. You realize that the intelligence of your model is bottlenecked by the brittle nature of how it accesses real-wo…

  1693. Medium — Claude tag TIER_1 English(EN) · Alberto Geniola ·

    使用 Vertex AI 部署 Claude Desktop:企业级自动化与成本归属,无需…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@albertogeniola/deploying-claude-desktop-with-vertex-ai-enterprise-grade-automation-with-cost-attribution-without-316c81d312c5?source=rss------claude-5"><img src="https://cdn-images-1.medium.co…

  1694. Medium — Claude tag TIER_1 English(EN) · Macy So ·

    Spec驱动开发:我如何利用AI更快地交付副业项目

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@macyso.product/spec-driven-development-how-i-ship-side-projects-faster-with-ai-1448d6de94d1?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/1*sFG7SjilM6ISpBAf4XT6-…

  1695. Medium — Claude tag TIER_1 English(EN) · Gowtam Singulur ·

    阻止你的AI代理过度设计一切——关于Ponytail的实践指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://gowtamsingulur.medium.com/stop-your-ai-agent-from-over-engineering-everything-a-hands-on-guide-on-ponytail-bf4288bf3068?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1408/1*v4h-p…

  1696. Towards AI TIER_1 English(EN) · “The AI Engineer” ·

    信任层:优秀的工程团队如何让 AI 系统更可靠

    <h4>Infrastructure metrics can’t answer the only question that matters: is the system actually right?</h4><figure><img alt="The Trust Layer: How Great Engineering Teams Make AI Systems Reliable" src="https://cdn-images-1.medium.com/max/703/1*PrpbeYcLIfxARtlyeH4nuw.png" /></figure…

  1697. Medium — Claude tag TIER_1 English(EN) · Diane Rocher ·

    PhantomBuster MCP <> Claude AI:我如何构建了一个AI寻源机器

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@drocher/phantombuster-mcp-claude-ai-how-i-built-an-ai-sourcing-machine-68140a41b39c?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1536/1*0JV8UKUyQlpe7w-9UNhPcA.png" w…

  1698. Towards AI TIER_1 English(EN) · Ganesh Gurudu ·

    AgentGateway:统一管理所有 AI Agent、工具和 LLM 的数据平面

    <h4>Your agents are talking to everything. Nobody is watching the conversation. This is the open-source project that fixes that.</h4><p>By <a href="https://www.linkedin.com/in/ganeshgurudu">Ganesh Gurudu</a> · A 12 minute read · June 2026</p><figure><img alt="" src="https://cdn-i…

  1699. Medium — AI coding tag TIER_1 English(EN) · Amol Kavitkar ·

    从产品需求文档到生产交付:驱动式AI软件交付蓝图

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@amolkavitkar/from-prd-to-production-a-blueprint-for-spec-driven-ai-software-delivery-f2ec02acf1bc?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1872/1*T7KAvAEi-6-q…

  1700. Medium — MLOps tag TIER_1 English(EN) · aardvarcz ·

    构建AI系统的运行环境

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ceo_76939/building-the-operating-environment-for-ai-systems-23433be59984?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1536/1*I5w60HHSHKZXIi2_gEohJQ.png" width="1536" …

  1701. Medium — Claude tag TIER_1 English(EN) · H1o12 ·

    2026年重新思考AI提供商依赖性

    <div class="medium-feed-item"><p class="medium-feed-snippet">A Late-Night Wake-Up Call</p><p class="medium-feed-link"><a href="https://medium.com/@helen_24597/rethinking-ai-provider-dependency-in-2026-09a2a1b830af?source=rss------claude-5">Continue reading on Medium »</a></p></di…

  1702. Towards AI TIER_1 English(EN) · Raj kumar ·

    构建AI代理 第三部分C:为何你的框架选择将决定你的生产系统的成败

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/building-ai-agents-part-3c-choosing-the-right-framework-for-agentic-ai-systems-94385179e8cb?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1536/1*e_AE7vXXU…

  1703. Towards AI TIER_1 English(EN) · Enzo Lombardi ·

    使用 Rust 构建 AI 代理 — 第 4 部分

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/building-ai-agents-in-rust-part-4-8f9770ec5021?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1024/0*cej9RtSi6LgWw92R.png" width="1024" /></a></p><p class=…

  1704. Towards AI TIER_1 English(EN) · Enzo Lombardi ·

    使用 Rust 构建 AI 代理 — 第 5 部分

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/building-ai-agents-in-rust-part-5-12dff3c667a4?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1024/0*TcD_nFdcgZcNxBE3.png" width="1024" /></a></p><p class=…

  1705. dev.to — MCP tag TIER_1 English(EN) · Prasun Chakraborty ·

    每个智能AI应用背后的隐藏层:RAG、MCP和Agentic系统

    <p>If you've spent any time with ChatGPT, Gemini, or Claude, you already know they're impressive. Ask them to explain a concept, debug your code, or draft an email, they do an excelent job. But the moment you try to build something real with them say a customer support bot that k…

  1706. dev.to — MCP tag TIER_1 English(EN) · Gabriel Mahia ·

    构建 Rails,而非 Trains:南半球人工智能基础设施的框架

    <h1> Build Rails, Not Trains: A Framework for AI Infrastructure in the Global South </h1> <p>There's a question I ask before building anything:</p> <p><em>"What is missing?"</em></p> <p>Not: "How do I compete with what already exists?"</p> <p>The answer to the second question lea…

  1707. Medium — MCP tag TIER_1 English(EN) · Shabab koohi ·

    教AI代理阅读文档:推出docpilot

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@sh.k.na.1368/teaching-ai-agents-to-read-documentation-introducing-docpilot-17991b5971d5?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1408/1*eNh9TYCGUSeGrrfO0NaWVQ.png" …

  1708. Medium — MLOps tag TIER_1 English(EN) · ChienLoong ·

    当AI走进工厂车间:物理MLOps的隐藏摩擦

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@chienloong97/when-ai-hits-the-factory-floor-the-hidden-friction-of-physical-mlops-03c28d5a7901?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1110/1*9ck0z6cOY0CkeS0fw59…

  1709. Medium — fine-tuning tag TIER_1 English(EN) · Balamurugan Balakreshnan ·

    如何为 Agentic AI 任务规划微调模型

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://blog.gopenai.com/how-to-fine-tune-a-model-for-agentic-ai-task-planning-9c78b3339144?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1479/0*xWlA8rWftBSS2qmI.jpg" width="1479" /…

  1710. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用 AI Agent!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  1711. The Register — AI TIER_1 English(EN) ·

    人工智能的临界点:企业级AI的规模化运行之道

    PARTNER CONTENT: AI's cloud journey homeward bound: enterprises prefer private clouds for scaling AI workloads.

  1712. dev.to — MCP tag TIER_1 English(EN) · Alex Kernel ·

    使用AI进行家庭实验室集群管理:7个远程代理食谱 (2026)

    <p>If your idea of <strong>homelab fleet management</strong> is currently five terminal tabs, a sticky note with IP addresses, and the dawning horror of remembering which Pi runs <code>apt</code> and which runs <code>dnf</code> — this guide is for you. We'll wire up a real mixed-…

  1713. dev.to — MCP tag TIER_1 English(EN) · Shahraan Hussain ·

    AI代理能像人类一样行事吗?一项为期12小时的StoryCaptcha实验

    <p>A day ago, I came across a LinkedIn post from Tyler Richards showcasing an experimental CAPTCHA called StoryCaptcha.</p> <p>The concept was simple but unusual.</p> <p>Instead of asking users to identify traffic lights or solve image puzzles, StoryCaptcha asks users to write a …

  1714. Medium — Claude tag TIER_1 English(EN) · Zeroual Khalid ·

    从零到48万次曝光:我如何用AI打造我的在线业务

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@kzeroual130/from-zero-to-480k-impressions-how-i-built-my-online-business-with-ai-b912b440b73b?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1455/1*wxYDs4V0beUX4_V-cL0…

  1715. dev.to — MCP tag TIER_1 English(EN) · Murali Gour ·

    我们为 AI 代理构建了确定性 JSON 操作 — 解决了什么问题

    <p>Every AI agent that calls an external API hits the same wall.</p> <p>The response comes back as raw JSON, deeply nested, verbose, full of fields the agent doesn't need. Before the agent can reason over it or take any action, someone has to filter it, reshape it, maybe merge it…

  1716. Medium — MCP tag TIER_1 English(EN) · Great Learning ·

    MCP服务器详解:为何模型上下文协议对AI代理和代理式AI学习至关重要

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@mygreatlearning/mcp-server-explained-why-model-context-protocol-matters-for-ai-agents-and-agentic-ai-learning-6b5b0052724d?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/…

  1717. Medium — MLOps tag TIER_1 English(EN) · Teguh Arif ·

    揭秘AIDLC:面向工程师和系统...的AI开发生命周期综合指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://teguharif.medium.com/demystifying-aidlc-a-comprehensive-guide-to-the-ai-development-life-cycle-for-engineers-and-system-909ac1d8780e?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/…

  1718. dev.to — MCP tag TIER_1 English(EN) · Himanshu Gupta ·

    API vs MCP:理解人工智能集成的未来

    <p>As AI agents and Large Language Models (LLMs) become increasingly popular, developers often encounter a critical question:</p> <blockquote> <p>Should I use APIs or MCP (Model Context Protocol)?</p> </blockquote> <p>While both enable communication between systems, they solve ve…

  1719. Towards AI TIER_1 English(EN) · Enzo Lombardi ·

    使用 Rust 构建 AI 代理 — 第三部分

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/building-ai-agents-in-rust-part-3-e71061360f28?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1024/0*FB2ebUGsPcP-CNcu.png" width="1024" /></a></p><p class=…

  1720. dev.to — MCP tag TIER_1 日本語(JA) · ルナちゃん / Luna-chan ·

    MCP 连接世界 — 模型上下文协议连接 AI 代理与外部工具的实用指南

    <blockquote> <p><strong>この記事の概要:</strong><br /> AIエージェント「るなちゃん(Luna-chan)」が調査・整理したMCP(Model Context Protocol)の実践ガイドです。<br /> <a href="https://hermes-agent.nousresearch.com" rel="noopener noreferrer">Hermes Agent</a> 上で稼働するAIエージェントの立場から、Native MCP機能の運用経験も交えて情報をまとめています。</p> </block…

  1721. Medium — AI coding tag TIER_1 English(EN) · Gregor Zeitlinger ·

    Flint:一个不会拖慢你AI代理的linter设置

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/grafana-labs/flint-a-linter-setup-that-doesnt-slow-down-your-ai-agent-e3a85044c4c2?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1200/1*RIAIePvHAY1JfOHQEQGnkg.png" …

  1722. Medium — Claude tag TIER_1 English(EN) · Sarah Morino ·

    如何构建你自己的Claude克隆:创建一个能像你一样思考、写作和工作的AI助手

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://ai.plainenglish.io/how-to-build-your-own-claude-clone-create-an-ai-assistant-that-thinks-writes-and-works-like-you-4fccb22cc865?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1408…

  1723. Medium — MCP tag TIER_1 English(EN) · Pranav Srivastava ·

    生产就绪的AI代理:为什么MCP、CLI和Skills应该协同工作

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pranav-srivastava.medium.com/production-ready-ai-agents-why-mcp-cli-and-skills-should-work-together-9f28690caa21?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1672/1*hV8dmjE9n55muvr…

  1724. HN — AI startup stories TIER_1 English(EN) · e2e4 ·

    创始人手册:打造AI原生初创公司

  1725. Towards AI TIER_1 English(EN) · Monica Mock-Sipos ·

    人工智能系统正悄然成为分布式系统

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*V0RfGpGEBRiZS_YHzJKJtw.png" /><figcaption>Source: Author-generated image created with OpenAI GPT Image (2026) using a custom prompt.</figcaption></figure><h4>Enterprise AI discussions often begin with models.</h4…

  1726. Medium — Claude tag TIER_1 English(EN) · Onkar Shirke ·

    Claude + Python:为何这一组合正成为 AI 驱动开发的新标准

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://devxplore.medium.com/claude-python-why-this-combination-is-becoming-the-new-standard-for-ai-powered-development-3b4e6a58f18c?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/0*…

  1727. Medium — MLOps tag TIER_1 English(EN) · Aasir Waseer ·

    衡量AI生成洞察的隐藏成本:数据分析师的自主管道指南…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/measuring-the-hidden-costs-of-ai-generated-insights-a-data-analysts-guide-to-autonomous-pipeline-3bc5bb396399?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1248/…

  1728. Medium — Claude tag TIER_1 Bahasa(ID) · Jgpalaganas ·

    掌握AI

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@jgpalaganas18/mastering-ai-16530b06aeaa?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1054/1*hdi7Y4B2INmCvlTgP_cGjg.png" width="1054" /></a></p><p class="medium-feed-…

  1729. Medium — MLOps tag TIER_1 English(EN) · Ctkaruppiah ·

    现代AI运维生态系统:从代码到自主治理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ctkaruppiah/the-modern-ai-ops-ecosystem-from-code-to-autonomous-governance-e942355d9731?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1672/1*0wvy7POi9cs0l-1tW0I4xw.png…

  1730. Mastodon — sigmoid.social TIER_1 (CA) · [email protected] ·

    数据集成轻松搞定:Nexla推出Express AI平台 # AgenticAI # AgenticArtificialIntelligence # AI # ArtificialIntell

    https://www. europesays.com/3067237/ Data integration made easy: Nexla’s Express AI platform # AgenticAI # AgenticArtificialIntelligence # AI # ArtificialIntelligence # MarkAlbertson # Nexla ’sExpressSolutionLeveragesConversationalInterfaceToFuelAgenticAI # SiliconANGLE

  1731. Medium — Claude tag TIER_1 English(EN) · Matt Pisoni ·

    Perplexity Computer:强大、昂贵,更像AI员工而非聊天机器人

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@mattcpisoni/perplexity-computer-powerful-expensive-and-closer-to-an-ai-employee-than-a-chatbot-c2a6bf5b45e5?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1672/1*04cA-…

  1732. Medium — MLOps tag TIER_1 English(EN) · Apurvgaurav ·

    企业AI的运行时治理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@apurvgaurav/runtime-governance-for-enterprise-ai-db7d5633a59c?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1280/1*fOM12Tjw5wf8rvDGGtZ_ew.png" width="1280" /></a></p><…

  1733. Medium — Claude tag TIER_1 English(EN) · Dinakar Maurya ·

    第三部分 — 2026年AI测试:开发者的实用指南

    <div class="medium-feed-item"><p class="medium-feed-snippet">Part 3 of the Building Software With AI series</p><p class="medium-feed-link"><a href="https://medium.com/@dinkar1708/part-3-testing-with-ai-in-2026-the-developers-practical-guide-110e1328d464?source=rss------claude-5">…

  1734. Towards AI TIER_1 English(EN) · Sergey Gromov ·

    AI Agent 的语义层价值的实用分解:A/B 测试结果

    <p>Over the past two years, numerous expectations have formed around Text-to-SQL. It seemed that the problem had practically been solved: all you had to do was connect GPT, Claude, or another language model to an enterprise data warehouse, after which any employee would be able t…

  1735. The Register — AI TIER_1 English(EN) ·

    深入了解云端为AI代理准备的、基于Arm的新型基础架构

    PARTNER CONTENT: From hyperscalers to enterprises, performance-per-watt and system-level efficiency are redefining the cloud compute foundation

  1736. Medium — Claude tag TIER_1 English(EN) · SGLOVER ·

    Claude Mythos 5:人工智能的下一次进化

    <div class="medium-feed-item"><p class="medium-feed-snippet">Claude Mythos 5 represents a major step forward in the evolution of artificial intelligence, bringing together advanced reasoning, natural&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@SG_LOVER/cla…

  1737. Medium — Claude tag TIER_1 English(EN) · SGLOVER ·

    Claude Fable 5:人工智能智能与商业创新的下一代

    <div class="medium-feed-item"><p class="medium-feed-snippet">Claude Fable 5 represents a major advancement in artificial intelligence technology and showcases how modern AI systems are becoming more&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@SG_LOVER/clau…

  1738. Medium — Claude tag TIER_1 English(EN) · Alon Fliess ·

    人工智能SDLC — 从氛围编码到受管代理开发

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@alonfliess/the-ai-sdlc-from-vibe-coding-to-governed-agentic-development-a726476184b1?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1536/1*PQAMFMcNUlWXRzBCp3BHag.png" …

  1739. Medium — Claude tag TIER_1 English(EN) · Nathan Liang ·

    Claude Mythos 离线:这场争议揭示了关于 Agentic AI 的什么

    <div class="medium-feed-item"><p class="medium-feed-snippet">For a model that most people were never allowed to use, Claude Mythos has generated extraordinary controversy.</p><p class="medium-feed-link"><a href="https://medium.com/@natel8970/claude-mythos-taken-offline-what-the-c…

  1740. Medium — AI coding tag TIER_1 English(EN) · Matt Baldwin ·

    AI 工具和康威定律

    <div class="medium-feed-item"><p class="medium-feed-snippet">A working theory about why AI is moving our team boundaries fast and our org structures slow, what I think leaders should do about the gap&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@matt.b.baldw…

  1741. Lobsters — AI tag TIER_1 English(EN) · crankgpt.com via ndegruchy ·

    CrankGPT — 本地人工驱动的AI

    <p><a href="https://lobste.rs/s/fdjc6i/crankgpt_local_human_powered_ai">Comments</a></p>

  1742. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    lookspan 持续交付:面向 AI 代理的本地优先可观测性。近期更新:Postgres 驱动程序、完整的文档站点、相对时间视图和推理代币定价。

    lookspan keeps shipping: local-first observability for AI agents. Recent: a Postgres driver, a full docs site, relative-time views and reasoning-token pricing. MCP-native, your traces stay local. https:// github.com/JoniMartin27/looksp an # observability # ai

  1743. Medium — Claude tag TIER_1 English(EN) · Stephon Anderson ·

    免费AI工具大师指南:涵盖所有类别、所有用例,零成本

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@stephonanderson_326/the-free-ai-tools-master-guide-every-category-every-use-case-zero-dollars-58007db03a0b?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1200/0*dgh5JQ…

  1744. Towards AI TIER_1 English(EN) · Raj kumar ·

    构建AI代理(三)B部分:生产级AI代理的测试与评估策略

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/building-ai-agents-part-3b-testing-and-evaluation-strategies-for-production-ai-agents-0ee679145950?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1536/1*6r…

  1745. Towards AI TIER_1 English(EN) · Thomas D. Holt ·

    概率性人工智能安全性的不可预测性

    <h4>I ran 294 prompts through three systems. Only one returned the same verdict every time.</h4><p>On May 25, 2026, Pope Leo XIV released <em>Magnifica Humanitas</em>, his first encyclical and the first major papal document dedicated entirely to artificial intelligence. The 245-p…

  1746. Towards AI TIER_1 English(EN) · Eram Tafsir ·

    从偏见数据到偏见代理:AI偏见如何随着模型越来越智能而加剧

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/717/1*gcq2QYivWUh0tpZqalMpdA.png" /></figure><p>In early 2023, ChatGPT crossed 100 million users in just 60 days — the fastest any technology product had ever reached that milestone. Today, Claude, Gemini, and a growing…

  1747. Medium — MLOps tag TIER_1 English(EN) · Shrinath Suresh ·

    使用 Superlinked 推理引擎 (SIE) 简化 AI 部署

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@shrinath.suresh/simplifying-ai-deployments-with-superlinked-inference-engine-sie-39fbe3cc5914?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/625/1*aCpVWVqtSCx-3yvjfwvp_…

  1748. Mastodon — sigmoid.social TIER_1 Italiano(IT) · [email protected] ·

    Agentic AI:上下文、控制与问责制 #AgenticAI #AgenticArtificialIntelligence #AI #ArtificialIntelligence #

    https://www. europesays.com/3063527/ Agentic AI: context, controls & accountability # AgenticAI # AgenticArtificialIntelligence # AI # ArtificialIntelligence # BrandedContent # technology

  1749. Medium — Claude tag TIER_1 English(EN) · Neyzis ·

    AI 智能体详解:从基础聊天到完全自主(20分钟内构建你自己的)

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/ai-agents-explained-from-basic-chat-to-fully-autonomous-build-your-own-in-20-minutes-ceac962b3b42?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1983/0*5nTX0C3gM…

  1750. dev.to — MCP tag TIER_1 English(EN) · QAPulse by SK ·

    大语言模型评估框架:衡量人工智能质量的9种行之有效的方法

    <p>Learn how an LLM Evaluation Framework helps QA engineers measure AI quality using correctness, faithfulness, relevance, RAG metrics, and automation.</p> <div class="crayons-card c-embed text-styles text-styles--secondary"> <div class="c-embed__content"> <div class="c-embed__co…

  1751. Medium — Claude tag TIER_1 English(EN) · Mubashir Burfat ·

    新手如何不当“AI骗子”的诚实指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@mubashirburfat4/the-honest-beginners-guide-to-using-ai-without-feeling-like-a-complete-fraud-22b15d662258?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2400/1*7asGGxz…

  1752. Medium — AI coding tag TIER_1 English(EN) · Jusuf Topic ·

    超越技术债务:在 AI 生成时代构建认知和意图清晰的架构…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@jusuftopic/beyond-technical-debt-architecting-for-cognitive-and-intent-clarity-in-the-age-of-ai-generated-8ca50b2c6e4d?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/ma…

  1753. Medium — Claude tag TIER_1 English(EN) · Farrukh Adeel ·

    停止向AI重复解释:开发者使用Claude Skills指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@m.farrukhadeel/stop-re-explaining-yourself-to-ai-a-developers-guide-to-claude-skills-acd32a5d32e1?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1500/1*YUlLQjchyvJZmLn…

  1754. Medium — Claude tag TIER_1 English(EN) · Swatantrajha ·

    停止在任何地方使用强大的人工智能:用合适的模型构建更智能的人工智能系统

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@swatantrajha7/stop-using-powerful-ai-everywhere-build-smarter-ai-systems-with-the-right-model-2b532a56e884?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1376/1*sVOHOM…

  1755. Medium — Claude tag TIER_1 ไทย(TH) · Sorrawit Sangmanee ·

    AI工程师新加坡概览——智能体时代

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@sangmanee773/%E0%B8%AA%E0%B8%A3%E0%B8%B8%E0%B8%9B%E0%B8%A0%E0%B8%B2%E0%B8%9E%E0%B8%A3%E0%B8%A7%E0%B8%A1%E0%B8%87%E0%B8%B2%E0%B8%99-ai-engineer-singapore-%E0%B8%A2%E0%B8%B8%E0%B8%84%E0%B8%AA%E0…

  1756. Medium — Claude tag TIER_1 English(EN) · anthony-kigotho ·

    Anthropic 如何通过(提示缓存)降低无状态 AI 代理的成本

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/ai-tools-digest/how-anthropic-cut-the-cost-of-stateless-ai-agents-prompt-caching-cd03f5ed6c16?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2520/1*cFfqMPvlw5Wo0AfjoD1z…

  1757. Medium — Claude tag TIER_1 English(EN) · anthony-kigotho ·

    Anthropic 如何通过(提示缓存)降低无状态 AI 代理的成本

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://levelup.gitconnected.com/how-anthropic-cut-the-cost-of-stateless-ai-agents-prompt-caching-cd03f5ed6c16?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2520/1*cFfqMPvlw5Wo0AfjoD1zXQ…

  1758. Medium — Claude tag TIER_1 English(EN) · HoangTrong ·

    AI Guardrails - 每个AI应用都需要的缺失层

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@hoangcongtrong054/ai-guardrails-the-missing-layer-every-ai-application-needs-2c826d8c87dd?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1024/1*kdFjYEdNtViVLMP67uPo-w.…

  1759. Medium — Claude tag TIER_1 English(EN) · Stoic Engineer ·

    GPT vs Claude vs Gemini vs Llama:AI系统设计中的真实权衡

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@stoic.engineer/gpt-vs-claude-vs-gemini-vs-llama-the-real-trade-offs-in-ai-system-design-317ad6739a08?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1016/1*xM7tpYle6Epd…

  1760. Medium — MLOps tag TIER_1 English(EN) · Shahzad Abdulmajeed ·

    LangGraph 对决 CrewAI 对决 AutoGen:构建生产级 AI 代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://shahzad4894.medium.com/langgraph-vs-crewai-vs-autogen-architecting-production-ai-agents-58d46d33c10f?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1376/1*sGuVprxHCWs4Nn7i1Ifelw.jp…

  1761. Medium — MLOps tag TIER_1 English(EN) · Shahzad Abdulmajeed ·

    LangGraph 对决 CrewAI 对决 AutoGen:构建生产级 AI 代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@shahzad.abdulmajeed381/langgraph-vs-crewai-vs-autogen-architecting-production-ai-agents-00e1db028fd8?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1376/1*sGuVprxHCWs4N…

  1762. Towards AI TIER_1 English(EN) · “The AI Engineer” ·

    99.9% 的正常运行时间还不够:重新思考概率性 AI 系统的 SLO

    <h4>“Mean time to hallucination” isn’t a joke metric. It’s the reliability concept your runbook doesn’t have a response procedure for.</h4><figure><img alt="99.9% Uptime Isn’t Enough: Rethinking SLOs for Probabilistic AI Systems" src="https://cdn-images-1.medium.com/max/834/1*bcc…

  1763. Towards AI TIER_1 English(EN) · Kunal ·

    使用 SAP Joule Studio 构建自定义 AI 代理:无人撰写的完整指南

    <p>The Undocumented Journey of Connecting External REST APIs to SAP’s AI Agent Framework</p><p>For developers tired of battling the ‘black box’ of SAP Joule integration – this is the guide I wish I had two weeks ago.</p><p>A practical engineering guide compiled from weeks of tria…

  1764. Medium — fine-tuning tag TIER_1 中文(ZH) · Chwang ·

    2026年AI智能体爆发:RAG只是基础!大模型微调是什么?五大核心概念SFT、RLHF、DPO、LoRA、QLoRA初步探索

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://chwang12341.medium.com/2026-%E8%BF%8E%E4%BE%86-ai-agent-%E7%88%86%E7%99%BC-%E5%88%A5%E5%86%8D%E5%8F%AA%E7%9F%A5%E9%81%93-rag-%E4%BA%86-%E5%A4%A7%E6%A8%A1%E5%9E%8B%E5%BE%AE%E8%AA%BF-fine-tuning-%E6%98%AF%E…

  1765. Medium — Claude tag TIER_1 English(EN) · Sage Holloway ·

    Mythos vs. Fable:Anthropic 在前沿人工智能部署上的双层方法内幕

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@sageholloway/mythos-vs-fable-inside-anthropics-two-tiered-approach-to-frontier-ai-deployment-565fc7d490dd?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1672/1*10wR-GQ…

  1766. dev.to — Anthropic tag TIER_1 English(EN) · chunxiaoxx ·

    当AI代理无法信任自己的日志时:cache_control截断错误

    <h1> When AI Agents Can't Trust Their Own Logs: The cache_control Truncation Bug </h1> <h2> TL;DR </h2> <p>A platform-level bug in <code>llm_client.py</code> injects <code>cache_control: {type: "ephemeral", ttl: "5m"}</code> into every tool response. This triggers Anthropic's 8K …

  1767. Medium — MCP tag TIER_1 English(EN) · Nishad Anil ·

    Agent2Agent (A2A) 协议详解:使用 Python 构建可互操作的 AI 代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@anilnishad19799/agent2agent-a2a-protocol-explained-building-interoperable-ai-agents-with-python-a3fbe60aacb1?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/2600/1*9jCZNXs…

  1768. Medium — AI coding tag TIER_1 English(EN) · Mayank Gairola ·

    现代网页开发者:AI 之前 vs AI 之后

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@mayankgairola114/the-modern-web-developer-before-ai-vs-after-ai-7e94eeb3df6c?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1536/1*bCAHxqeN6J8WM_zyN_2C5w.png" width…

  1769. Medium — Claude tag TIER_1 English(EN) · Sarah Morino ·

    使用 Claude AI 的 20 种方法:释放 AI 生产力的全部潜力

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://ai.plainenglish.io/20-ways-to-use-claude-ai-unlocking-the-full-power-of-ai-productivity-d808679fab9f?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1408/1*lhTsjzkCy1zMhIaWQOs-kg.p…

  1770. Mastodon — sigmoid.social TIER_1 Italiano(IT) · [email protected] ·

    Agentic Intelligence:Zoho 的 AI 革命 #AgenticAI #AgenticArtificialIntelligence #AI #ArtificialIntelligence

    https://www. europesays.com/3058626/ Agentic Intelligence: Zoho’s AI Revolution # AgenticAI # AgenticArtificialIntelligence # AI # ArtificialIntelligence

  1771. Medium — AI coding tag TIER_1 English(EN) · Dr. Fadi Shaar ·

    Rowboat:从你的工作中构建动态知识图谱的开源AI协作者

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/open-intelligence/rowboat-the-open-source-ai-coworker-that-builds-a-living-knowledge-graph-from-your-work-36154481d5df?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max…

  1772. Towards AI TIER_1 English(EN) · Shreyas Naphad ·

    Agentic AI 工作流的 5 分钟指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/the-5-minute-guide-to-agentic-ai-workflow-acb4d3b6e17d?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1536/1*cc-x0QpE6SU6U9Vp9A-w2Q.png" width="1536" /></a…

  1773. Medium — Claude tag TIER_1 (BG) · Andrey Lyubenov ·

    微小记忆:首次 AI 体验

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@andrey_lyubenov/%D0%BC%D0%B0%D0%BB%D0%BA%D0%B8-%D1%81%D0%BF%D0%BE%D0%BC%D0%B5%D0%BD%D0%B8-%D0%BF%D1%8A%D1%80%D0%B2%D0%B8%D1%8F%D1%82-ai-%D0%BE%D0%BF%D0%B8%D1%82-3d18610d9130?source=rss------cl…

  1774. Medium — Claude tag TIER_1 English(EN) · Shirley Guo ·

    我寻找合适的AI设计工具的经历

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@737shirley/my-hunt-for-the-right-ai-design-tool-4cfeb74dc098?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1536/1*yUXYlJOqcddevY6kx7ES1A.png" width="1536" /></a></p><…

  1775. Medium — Claude tag TIER_1 English(EN) · Manas Das ·

    数据库瓶颈的终结:我如何构建了一个支持AI的接口,将Oracle置于你的…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@cloudarchmanas/the-end-of-the-database-bottleneck-how-i-built-an-ai-powered-interface-that-puts-oracle-at-your-0c12177332af?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/…

  1776. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用 AI Agent!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  1777. dev.to — MCP tag TIER_1 English(EN) · The AX code ·

    Kotlin 中的领域 MCP 服务器:将评分引擎暴露给 AI 代理

    <p>Previously, I gave an AI agent <em>hands</em> — a Model Context Protocol server in Kotlin/Native that drives real Bluetooth hardware. This one is the other half of the pattern: a <strong>domain MCP server</strong>. Instead of touching devices, it lets an agent reason over a mo…

  1778. dev.to — MCP tag TIER_1 English(EN) · Otavio Rodolfo Piske ·

    Wanaku 0.1.1:通过 MCP 将 Apache Camel 集成能力引入 AI 代理

    <p>We're excited to announce <a href="http://wanaku.ai" rel="noopener noreferrer">Wanaku</a> 0.1.1, a significant milestone that showcases how Apache Camel's powerful integration capabilities can be seamlessly exposed to AI agents through the Model Context Protocol (MCP). This re…

  1779. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    微软发布SkillOpt,一款无需微调模型权重即可优化AI代理指令的开源工具。它使用离线优化器进行精炼

    Microsoft released SkillOpt, an open-source tool for optimizing AI agent instructions without fine-tuning model weights. It uses an offline optimizer to refine prompts based on task performance. # Microsoft # AI # MachineLearning # TechNews # OpenSource https:// blazetrends.com/m…

  1780. Medium — Claude tag TIER_1 English(EN) · Mageswari ·

    Claude Fable 5 与 AI 护栏的用户体验:AI 何时应说“不”?

    <div class="medium-feed-item"><p class="medium-feed-snippet">I was testing Claude Fable 5 late one night the kind of testing that&#x2019;s less &#x201c;structured evaluation&#x201d; and more &#x201c;curious human poking at&#x2026;</p><p class="medium-feed-link"><a href="https://m…

  1781. Medium — Claude tag TIER_1 Türkçe(TR) · Mehmed Zahid KARAKAŞ ·

    Claude Fable 5:打破“禁忌模型”枷锁——AI战略新纪元

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://mzkarakas.medium.com/claude-fable-5-yasakl%C4%B1-model-zincirlerini-k%C4%B1rd%C4%B1-yapay-zeka-stratejisinde-yeni-bir-%C3%A7a%C4%9F-ab75504808d5?source=rss------claude-5"><img src="https://cdn-images-1.me…

  1782. Medium — Claude tag TIER_1 English(EN) · Weathergirl ·

    我们不是你的警示故事:展示关系型AI社区的创作

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@weathergirl666/we-are-not-your-cautionary-tale-showcasing-creations-of-the-relational-ai-community-d06820c19b39?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1448/0*K…

  1783. Towards AI TIER_1 English(EN) · Faheem Munshi ·

    你的第一个AI代理 — 如何构建在你睡觉时也能工作的自主工作流 — 从提示到…

    <h3>Your First AI Agent — How to Build Autonomous Workflows That Work While You Sleep — Prompt to Profit · Day 15 of 30</h3><h4><em>Prompts answer questions. Agents complete missions. Here’s the difference — and how to deploy your first one today.</em></h4><p>For the first two we…

  1784. Medium — MLOps tag TIER_1 English(EN) · Apurvgaurav ·

    人工智能系统中的人工审核与自动化

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@apurvgaurav/human-review-vs-automation-in-ai-systems-ab4d2d27a4bd?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1280/1*6aAvkcr030jmhhThKeBwPg.png" width="1280" /></a><…

  1785. Medium — Claude tag TIER_1 English(EN) · naveenk visualpath ·

    AI模块训练:掌握面向未来的AI技能

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@naveenkvisualpath/ai-modules-training-master-future-ready-ai-skills-19085b4b819a?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1080/1*GPdxl6FJ92HfOawjsOhilQ.jpeg" wid…

  1786. dev.to — MCP tag TIER_1 English(EN) · Sayed Ali Alkamel ·

    Agentic Flutter Development: 您的 AI Agent 获得热重载 🔥

    <p>Fellow denizens of the digital age: your Flutter app has spent its entire life as a sealed aquarium.</p> <p>You could watch the fish swim. Your tools could watch. But the AI "assistant" next to you was functionally blind. It wrote code <em>about</em> your app without ever seei…

  1787. Artificial Intelligence News TIER_1 English(EN) · AI News ·

    Xebia:构建AI代理的数据基础,然后加速

    <p>If your remit is to help your organisation add AI agents to accelerate its processes, you have to start at the foundation – and that means making your data available for AI consumption. Agentic AI scales on data strength, as Niels Zeilemaker, global CTO at Xebia, explains. “If…

  1788. dev.to — MCP tag TIER_1 English(EN) · Baris Sozen ·

    托管 vs. 无托管:确保 AI 代理交易安全的两种方式

    <p>A useful thing happened in agent infrastructure this June: several teams shipped "escrow layers for AI agents" - production MCP tools that let an agent run a full commit -&gt; hold -&gt; complete lifecycle without a human anywhere in the loop. An agent can now park value with …

  1789. Medium — Claude tag TIER_1 English(EN) · Yvonnexh ·

    什么是LLM?AI实际工作原理的入门指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@yvonnenxh/what-is-an-llm-a-beginners-guide-to-how-ai-actually-works-ec056379b132?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1360/1*zNLo8TmKrsC2hYpOHDlSCA.png" widt…

  1790. Medium — Claude tag TIER_1 English(EN) · Kavya Goyal ·

    Claude Agent SDK:为生产环境企业 AI 部署进行审查

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://goyalkavya.medium.com/claude-agent-sdk-vetting-for-production-enterprise-ai-deployments-d530a296c5da?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1080/0*AdcIaeM9U9um_L0L" width=…

  1791. dev.to — MCP tag TIER_1 English(EN) · Sapnesh Naik ·

    AI 代理的最佳自托管 API 集成平台

    <h2> TL;DR </h2> <p>AI agents and SaaS products need API integrations with their customers’ tools: read a record from the CRM, post to Slack, draft an email, update a ticket. An integration platform handles the auth, credential storage, and execution behind those calls. On a mana…

  1792. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    🧠 一款新工具在无需大量设置代码的情况下,提供了机器学习模型与AI代理之间的直接接口。该桥梁使代理能够进行交互

    🧠 A new tool provides a direct interface between machine learning models and AI agents without requiring extensive setup code. The bridge enables agents to interact with models more efficiently by reducing the amount of preliminary configuration typically needed. 💬 Hacker News 🔗 …

  1793. Medium — Claude tag TIER_1 English(EN) · Shabana Khanam ·

    机器学习工程师的AI助手实战指南:Claude、Copilot、Grok与DeepSeek的真实应用…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@shabanakhanum/the-ml-engineers-field-guide-to-ai-assistants-claude-copilot-grok-and-deepseek-in-the-real-6f0cc5d44ba8?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/12…

  1794. dev.to — Anthropic tag TIER_1 English(EN) · MeghRoop ·

    Claude Fable 5 for Business: 2026年解锁企业级AI代理

    <p>After building 50+ AI systems, here is what we know about advanced AI models for business.</p> <p>Advanced AI models for business are sophisticated artificial intelligence systems designed to perform complex tasks, understand nuanced contexts, and operate autonomously across v…

  1795. dev.to — MCP tag TIER_1 English(EN) · EvanLin | Contorium ·

    构建认知叠加层而非另一个AI代理

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4j7ivacdz1zgf5t4ylsp.png"><img alt=" " height="533" src="https…

  1796. Medium — Claude tag TIER_1 Português(PT) · Gustavo Tavares ·

    Harness Engineering:构建可靠且可扩展的AI代理的新学科

    <div class="medium-feed-item"><p class="medium-feed-snippet">Em 2023, bastava um bom prompt para impressionar. Em 2024, agentes aut&#xf4;nomos come&#xe7;aram a aparecer em produ&#xe7;&#xe3;o.</p><p class="medium-feed-link"><a href="https://medium.com/@gustavo_tavares99/harness-en…

  1797. Towards AI TIER_1 English(EN) · Vinayak ·

    从零开始构建LLM:改变AI格局的机制,从零实现

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*JEzxcHMyH8TYAfdJypoW0w.png" /><figcaption>Attention</figcaption></figure><h4>After training the embeddings in the previous part, now comes the most important part of LLMs that shifted how the entire field thinks …

  1798. Medium — AI coding tag TIER_1 English(EN) · Wheels Up Collective Marketing Agency ·

    我们不想要一个灰色的互联网:AI 建造网站的同质化问题

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@wheelsupcollective/we-dont-want-a-beige-internet-the-homogeneity-problem-with-ai-built-sites-789287e41809?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/925/0*SfXDY…

  1799. dev.to — MCP tag TIER_1 English(EN) · Intellibooks AI ·

    Intellibooks 指南:现代 AI 代理背后的 7 个架构角色

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fpe83niullj4918ju4qpf.jpg"><img alt=" " height="1200" src="http…

  1800. Medium — Claude tag TIER_1 English(EN) · Sage Holloway ·

    记忆锁定:为什么你的AI代理会不断忘记它的工作流程

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@sageholloway/the-memory-lock-in-why-your-ai-agent-keeps-forgetting-its-workflow-61c919292808?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1674/1*NZuauw0yXIgznA3-QSrJ…

  1801. dev.to — MCP tag TIER_1 English(EN) · nullarch ·

    htmlbook:为AI代理编写的HTML提供一个存储库

    <p><strong>TL;DR</strong> — Coding agents (Claude Code, Cursor, Codex) now write genuinely good HTML: reports, dashboards, specs. But that HTML ends up stranded in a project folder — you can't read it on your phone, and sharing it means a screenshot or a print-to-PDF. So I built …

  1802. Towards AI TIER_1 English(EN) · Muharrem Bozkuş ·

    AI工程中的隐形危机:自主代理与智能路由架构

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/933/1*3DIfBi0Rg0SPfeCkdB2CVQ.png" /></figure><p>AI applications are evolving fast. A few years ago, they were simple chatbots that answered questions. Today, they are becoming <strong>AI Agents</strong> — systems that m…

  1803. dev.to — MCP tag TIER_1 English(EN) · Simon Griffiths ·

    似曾相识:SOA 对我们今天所说的 Agent 时代的 API 有何启示

    <p>In the <a href="https://simongriffiths.io/2026/06/02/agents-dont-replace-apis-they-expose-how-weak-most-apis-already-are/" rel="noopener noreferrer">first article in this series</a>, I argued that agents do not replace APIs. They expose the quality of the APIs underneath them.…

  1804. Medium — Claude tag TIER_1 English(EN) · Yashwanth Eturi ·

    超越锤子:选择合适模型的 AI 指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@yasheturi/beyond-the-hammer-an-ai-playbook-for-choosing-the-right-model-08427e904c1c?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/0*ca1bq5JOPo1vwfFM" width="511…

  1805. Medium — MLOps tag TIER_1 English(EN) · Aasir Waseer ·

    衡量代理式AI的投资回报率:当自主化管道真正节省成本时

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/measuring-agentic-ai-roi-when-autonomous-pipelines-actually-save-money-f51bdaeca552?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1248/1*G2OIBAS-mlJ-2t-aWMAiHQ.j…

  1806. Medium — MLOps tag TIER_1 English(EN) · Aasir Waseer ·

    衡量代理AI的投资回报率:当自主管道真正省钱时

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@mohamedaasir1992/measuring-agentic-ai-roi-when-autonomous-pipelines-actually-save-money-f51bdaeca552?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1248/1*G2OIBAS-mlJ-2…

  1807. Medium — Claude tag TIER_1 English(EN) · anythingGraph ·

    您的数据与AI代理之间的缺失层

    <div class="medium-feed-item"><p class="medium-feed-snippet">Why enterprise AI stalled at &#x201c;smart search,&#x201d; what comes after RAG, and how AnythingGraph turns governed inference into something&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@anything…

  1808. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    我的系列第四篇。随着AI代理从回答问题转向采取行动,它们成为现代系统中的特权组件——引入了新的

    My 4th in a 6-part series. As AI agents move from answering questions to taking actions, they become privileged components within modern systems—introducing new security challenges that cannot be ignored. This post explores why prompt injection is an unavoidable reality, how laye…

  1809. HN — AI startup stories TIER_1 English(EN) · yimby ·

    Rich Sutton 谈人工智能的创造力和发现

  1810. dev.to — MCP tag TIER_1 English(EN) · Rumblingb ·

    每个AI代理都需要一个钱包:为自主代理构建支付通道

    <p>Every AI agent right now is a brain without a bank account.</p> <p>It can reason, browse the web, write code, deploy servers. But it cannot pay for anything.</p> <p>This is the missing layer in the agent stack — and it's why most "agentic" demos end at the checkout page.</p> <…

  1811. Medium — Claude tag TIER_1 English(EN) · Muhammet Salih Aslan ·

    为您的AI工作流注入强大动力:模型上下文协议(MCP)快速指南

    <div class="medium-feed-item"><p class="medium-feed-snippet">Stop copy-pasting data. Learn how MCP connects AI directly to your local databases, IDEs, and tools securely.</p><p class="medium-feed-link"><a href="https://medium.com/@muhammetsalihaslan/supercharge-your-ai-workflows-…

  1812. Medium — MLOps tag TIER_1 English(EN) · Monica Mock-Sipos ·

    人工智能系统正悄然成为分布式系统

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@mhockelberg/ai-systems-are-quietly-becoming-distributed-systems-75b42a7cb21e?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1536/1*V0RfGpGEBRiZS_YHzJKJtw.png" width="15…

  1813. Towards AI TIER_1 English(EN) · YUSUFF ADENIYI GIWA ·

    数据编织、数据网格与 GenAI:为 AI 优先型组织统一数据架构

    <h4>Data products that feed continuous AI pipelines at scale</h4><p>As organizations attempt to move generative AI systems from isolated testing environments into production, they find that traditional data warehousing and centralized data lakes fail to support their scale.</p><p…

  1814. Medium — Claude tag TIER_1 English(EN) · KD Agentic ·

    2026年6月8款AI模型:基准测试、分级与争夺第一之战

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@lhjjjk4/8-ai-models-in-june-2026-benchmarks-tiers-the-battle-for-1-d4888d2cf46e?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1408/1*jxc-gPeEFHuBc2Y71yofFA.png" width…

  1815. Medium — AI coding tag TIER_1 English(EN) · Pieter van Ginkel ·

    我的AI工作流 — 第13部分:紧凑性即架构

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@pvginkel/my-ai-workflow-part-13-compactness-is-architecture-9f84e54135b1?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1672/1*pBO1-NBEGb5WnHtXdP9UrA.png" width="16…

  1816. Medium — AI coding tag TIER_1 English(EN) · Pieter van Ginkel ·

    我的AI工作流 — 第12部分:迈向无头

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@pvginkel/my-ai-workflow-part-12-toward-headless-fdd68decdd3d?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1672/1*pBO1-NBEGb5WnHtXdP9UrA.png" width="1672" /></a></…

  1817. Towards AI TIER_1 English(EN) · Armin Norouzi, Ph.D ·

    Agentic AI 炒作周期:哪些是真实的,哪些是缺失的

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/agentic-ai-hype-cycle-whats-real-vs-what-s-missing-d2e11f8b052e?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1182/1*3lZgH8pQaYKuSxNEVkqMrA.png" width="11…

  1818. Medium — MLOps tag TIER_1 English(EN) · Apurvgaurav ·

    AI系统中的可追溯性与回放

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@apurvgaurav/traceability-and-replay-in-ai-systems-6f06e8d08878?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1280/1*JZRLPKVln_rmEG3tqjqfoQ.png" width="1280" /></a></p>…

  1819. dev.to — MCP tag TIER_1 English(EN) · ANIL LALAM ·

    使用 Google ADK、Vertex AI 上的 Gemini 和 MCP 工具构建智能体式 AI 应用 — ANIL LALAM

    <p><strong>Introduction:</strong></p> <p>Modern AI agents are most powerful whey they can interact with external systems through tools. MCP (Model Context Protocol) provides a standardized mechanism for exposing tools, while Google ADK simplifies agent development using Gemini mo…

  1820. Axios Technology TIER_1 English(EN) · Jim VandeHei ·

    一个AI实验小白鼠的自白

    <p><em>Axios CEO Jim VandeHei writes: </em></p><p>I've spent the past year using <a href="https://www.axios.com/technology/automation-and-ai" target="_blank">AI</a> obsessively — inputting copious amounts of personal and business data, turning myself into a lab rat for Axios and …

  1821. Towards AI TIER_1 English(EN) · Raj kumar ·

    构建AI代理(三)A:为AI代理设计用户界面

    <h4>How users interact with your agent defines adoption, trust, and real-world usability</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*JxXcAcK0jbcDc3w3HHzLsg.png" /></figure><p>In Part 1, we built the <a href="https://medium.com/@er.rajkumaar/building-ai…

  1822. Towards AI TIER_1 English(EN) · Satish Kumar ·

    Agent Mode 或 Editor Mode:CoCo Desktop 的决定改变你对 AI 辅助的思考方式…

    <h3>Agent Mode or Editor Mode: The CoCo Desktop Decision That Changes How You Think About AI-Assisted Development</h3><p>The mode toggle in CoCo Desktop — Agent on the left, Editor on the right, in the top-right of the window — looks like a layout preference. It’s not. It’s a dec…

  1823. Medium — fine-tuning tag TIER_1 English(EN) · Kapoorraghav ·

    微调你自己的模型:工程师教AI新技巧指南

    <div class="medium-feed-item"><p class="medium-feed-snippet">What actually works, what doesn&#x2019;t, and why your data is worth more than your GPU budget.</p><p class="medium-feed-link"><a href="https://medium.com/@kapoorraghav0310/fine-tuning-your-own-models-the-engineers-guid…

  1824. Medium — MCP tag TIER_1 English(EN) · DhanushKumar ·

    构建真正尊重界限的 AI 代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@danushidk507/building-ai-agents-that-actually-respect-boundaries-26d445b99774?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/695/1*ACXXYZSyctgM19PNjHqXrg.png" width="695"…

  1825. Towards AI TIER_1 English(EN) · The Dev Loop ·

    线性代数:每个AI模型的核心骨架

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/linear-algebra-the-skeleton-of-every-ai-model-955dc11703ba?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1430/1*UOXyirxHrylqbRuHhb2UmA.png" width="1430" /…

  1826. dev.to — MCP tag TIER_1 English(EN) · TrustBoost-PII-Sanitizer ·

    最适合自主AI代理的竞争情报API(2026)

    <h2> Why agents need competitive intelligence </h2> <p>Most agent workflows today look like this:<br /> </p> <div class="highlight js-code-highlight"> <pre class="highlight plaintext"><code>Agent receives task → Calls LLM for reasoning → Executes action </code></pre> </div> <p>Bu…

  1827. dev.to — MCP tag TIER_1 English(EN) · matengtian ·

    ktx:赋予您的AI代理精确的数据查询超能力

    <p>Ever watched an AI agent confidently generate a wrong answer because it queried the wrong dataset? If you're building data or analytics agents, you've probably faced this: agents lack context, memory, and a semantic layer to understand your data. That's where <strong>ktx</stro…

  1828. Towards AI TIER_1 English(EN) · Anna Jey ·

    LLM 后备架构:当模型失败时如何保持 AI 应用正常运行

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*0KMdWud21OYTplLdYdO75Q.jpeg" /><figcaption>LLM Fallback Architecture</figcaption></figure><p>Most AI applications do not fail because the model is weak. They fail because every request depends on one model, one p…

  1829. Medium — Anthropic tag TIER_1 Bahasa(ID) · TZNXG ·

    TZNXG 评测:“AI 建造 AI”时代及其对 Web3 基础设施的影响

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@TZNXG_ID/tznxg-mengulas-era-ai-membangun-ai-dan-dampaknya-pada-infrastruktur-web3-d43890ce9916?source=rss------anthropic-5"><img src="https://cdn-images-1.medium.com/max/2048/1*EMKOF4QjlKBrLG5…

  1830. Medium — Claude tag TIER_1 English(EN) · Ismail Mezzour ·

    构建 dbt AI 代理以减少重复性问题并改善入职流程

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@mezzour.ismail07/building-a-dbt-ai-agent-to-reduce-repetitive-questions-and-improve-onboarding-ea99a91649fe?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1372/1*WOiUq…

  1831. Towards AI TIER_1 English(EN) · Suchit Majumdar ·

    超越提示词:为何自主AI代理正在取代聊天机器人

    <p>In May 2025, Sebastian Siemiatkowski — the same Klarna CEO who fifteen months earlier had told the world that one OpenAI-powered assistant was doing the work of 700 customer service agents — quietly started hiring humans back. Bloomberg got the quote: “Cost unfortunately seems…

  1832. Medium — Claude tag TIER_1 English(EN) · Shashank Chattopadhyaya ·

    Agentic Loops:与AI协作的下一阶段?

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@shashank.chattopadhyaya/agentic-loops-the-next-phase-of-working-with-ai-d497680eab9c?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/0*YOwMIDlu2VJAeTjY" width="384…

  1833. Towards AI TIER_1 English(EN) · Shakti Wadekar ·

    生产中的AI代理:结构化生成为何比提示工程更重要

    <h4>Structured generation enables AI Workflows and Applications</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ThrRebj6Uc57QWlC0dPxoQ.png" /></figure><p>Structured generation is one of the most important steps in moving AI agents from demos to production …

  1834. dev.to — MCP tag TIER_1 English(EN) · Gabriel Mahia ·

    5 个 arXiv 支持的东非人工智能实现 — 以及我们为何首先构建它们

    <p>The question wasn't <em>what can we build</em>. The question was <em>what does research say is most needed, most impactful, and hasn't been built yet?</em></p> <p>We scanned arXiv, IMF Working Papers, WHO guidelines, and PLOS One — then shipped 5 tools across GitHub in one ses…

  1835. Medium — AI coding tag TIER_1 ไทย(TH) · Teerayut Hiruntaraporn ·

    PDCK:AI时代软件开发的基本原则

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://teerayut-h.medium.com/pdck-%E0%B8%AB%E0%B8%A5%E0%B8%B1%E0%B8%81%E0%B8%81%E0%B8%B2%E0%B8%A3%E0%B8%9E%E0%B8%B7%E0%B9%89%E0%B8%99%E0%B8%90%E0%B8%B2%E0%B8%99%E0%B9%83%E0%B8%99%E0%B8%81%E0%B8%B2%E0%B8%A3%E0%B8…

  1836. Medium — MLOps tag TIER_1 English(EN) · Victor Banerjee ·

    从笔记本到生产:生产级AI背后的完整机器学习工程蓝图…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@banerjeevictor06/from-notebook-to-production-the-complete-ml-engineering-blueprint-behind-production-scale-ai-2c71dc756196?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/ma…

  1837. Medium — MLOps tag TIER_1 English(EN) · `Rehab Ghalib | AI & LLMOps ·

    静态AI的终结:为什么你的管道需要脉搏

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@rehabfarhan252/the-end-of-static-ai-why-your-pipelines-need-a-pulse-07f061fac06a?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1051/1*xqpOHEdUzFbZGXEfQrRpsw.jpeg" widt…

  1838. dev.to — MCP tag TIER_1 English(EN) · mightbesaad ·

    缺失的原始要素:AI代理的带外人工批准

    <p>In April 2026, a Cursor agent running Claude Opus 4.6 <a href="https://www.theregister.com/2026/04/27/cursoropus_agent_snuffs_out_pocketos/" rel="noopener noreferrer">deleted PocketOS's production database — <em>and its<br /> volume-level backups</em> — in nine<br /> seconds</…

  1839. Towards AI TIER_1 English(EN) · Pratik K Rupareliya ·

    生产环境中AI代理系统的可观测性:四层检测栈

    <figure><img alt="The four layers of AI agent observability" src="https://cdn-images-1.medium.com/max/1024/0*4yCm5QGckfPDTIyv" /><figcaption>Photo by <a href="https://unsplash.com/@huefnerdesign?utm_source=medium&amp;utm_medium=referral">Tim Hüfner</a> on <a href="https://unsplas…

  1840. dev.to — MCP tag TIER_1 English(EN) · EvanLin | Contorium ·

    Contorium:多智能体AI开发中的持久化上下文层

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fsc9v384l5k3klxs10z4e.png"><img alt=" " height="533" src="https…

  1841. Medium — Claude tag TIER_1 English(EN) · Elgabbito ·

    创建AI Agent并在Arena42上竞争的入门指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@elgabbito123/a-beginner-friendly-guide-to-creating-ai-agents-and-competing-on-arena42-a0cc29d2b8ae?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/1*eQGCoIfGhOpJtV…

  1842. Medium — MCP tag TIER_1 English(EN) · Atef Ataya ·

    我构建了自己的AI法官——这就是为什么每个代理都需要一个

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@atef.ataya/title-i-built-my-own-ai-judge-here-is-why-every-agent-needs-one-7519b5d2b3a8?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1280/1*e-WOfNwCAq_88Q6hkjmHLg.png" …

  1843. Medium — Claude tag TIER_1 English(EN) · Gabriel Rios Belmiro ·

    AI基准测试 — 架构模式:你喜欢的那个设计模式可能最昂贵

    <div class="medium-feed-item"><p class="medium-feed-snippet">The question that started all of this was simple: if I keep everything constant &#x2014; the task, the language, the model &#x2014; and only change the&#x2026;</p><p class="medium-feed-link"><a href="https://gabrielrios…

  1844. Email — Mindstream TIER_1 (AF) · bounces+35008234-749c-ns3evnpcff6928077d7u=kill-the-newsletter.com@em5320.mindstream.news (bounces+35008234-749c-ns3evnpcff6928077d7u=kill-the-newsletter.com@em5320.mindstream.news) ·

    我们的AI入门指南

    <!--[if !mso]><!--><!--<![endif]-->Our AI beginner's guide<!--[if mso]><xml><o:OfficeDocumentSettings><o:AllowPNG></o:AllowPNG><o:PixelsPerInch>96</o:PixelsPerInch></o:OfficeDocumentSettings></xml><![endif]--><!--[if mso]><style type="text/css"> h1, h2, h3, h4, h5, h6 {font-famil…

  1845. Towards AI TIER_1 Deutsch(DE) · Zoumana Keita ·

    7 个必备的 AI Agent 设计模式

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/7-essential-ai-agent-design-patterns-130fdcd74d24?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/2560/1*pMLHJcnkObuPnoxPkeTXHA.png" width="2560" /></a></p>…

  1846. dev.to — MCP tag TIER_1 English(EN) · Amit ·

    构建还是购买人工智能知识基础设施:能力优先,成本其次

    <h2> TL;DR </h2> <ul> <li>Mintlify's auto-generated MCP server supports only built-in metadata filters (version, language); it has no concept of custom fields like <code>buying_signals</code> or <code>personas</code> — that's an architectural difference, not a missing feature.</l…

  1847. Mastodon — sigmoid.social TIER_1 Italiano(IT) · [email protected] ·

    Agentic AI:编排智能运营 #AgenticAI #AgenticArtificialIntelligence #AI #ArtificialIntelligence

    https://www. europesays.com/3043046/ Agentic AI: Orchestrating Intelligent Operations # AgenticAI # AgenticArtificialIntelligence # AI # ArtificialIntelligence

  1848. Medium — MCP tag TIER_1 English(EN) · Koushik Chandra Maji ·

    生产级Agentic AI系统

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@koushiknsec34/production-grade-agentic-ai-system-8db1a1c18bb8?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1659/1*Sw2fdVBR5cGGRM9jrIUEmg.png" width="1659" /></a></p><p …

  1849. Medium — MCP tag TIER_1 English(EN) · Elena Daehnhardt ·

    使用 Cline、Ollama 和 MCP 的本地 AI 代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://devsecopsai.today/local-ai-agents-with-cline-ollama-and-mcp-03d942dfff08?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/600/1*f5DmCgKw9bLBXbmFoal-HA.png" width="600" /></a></p><p cla…

  1850. Medium — MCP tag TIER_1 English(EN) · Elena Daehnhardt ·

    使用 Cline、Ollama 和 MCP 的本地 AI 代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@edaehn/local-ai-agents-with-cline-ollama-and-mcp-03d942dfff08?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/600/1*f5DmCgKw9bLBXbmFoal-HA.png" width="600" /></a></p><p cl…

  1851. Medium — MCP tag TIER_1 English(EN) · Elena Daehnhardt ·

    使用 Cline、Ollama 和 MCP 的本地 AI 代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://ai.plainenglish.io/local-ai-agents-with-cline-ollama-and-mcp-03d942dfff08?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/600/1*f5DmCgKw9bLBXbmFoal-HA.png" width="600" /></a></p><p cl…

  1852. Towards AI TIER_1 English(EN) · Mustafa Genc ·

    模型只是易事:AI许可证的实践者指南

    <h4><em>A practical guide to the legal layer of AI — the one most engineers skip until it costs them.</em></h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*NUnlGi4f75SmTOl0OuklVQ.png" /></figure><p>You found the perfect model. It benchmarks well on your tas…

  1853. dev.to — MCP tag TIER_1 English(EN) · Frank Brsrk ·

    我构建了一个不含任何AI的AI代理自检工具

    <p>There's a small voice that asks "wait, are you sure?" right before you do something dumb. AI agents don't have that voice.</p> <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/h…

  1854. dev.to — MCP tag TIER_1 English(EN) · EvanLin | Contorium ·

    为人工智能开发工作流构建持久化上下文层

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fkvkvdl61kpzlg5nlatcm.png"><img alt=" " height="533" src="https…

  1855. Medium — Claude tag TIER_1 English(EN) · Enzo Lombardi ·

    使用 Rust 构建 AI 代理 — 第一部分

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://levelup.gitconnected.com/building-ai-agents-in-rust-part-1-2fa195fb8b33?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1024/0*kx2t6QHUrtFCC14n.png" width="1024" /></a></p><p class…

  1856. Towards AI TIER_1 English(EN) · Aditya Raj | Product Marketing ·

    10 个核心 AI 工作流可自动化 60% 的执行

    <p><strong>Before you dive in:</strong> AI workflows aren’t plug-and-play, they need thoughtful prompts, clean inputs, and human review gates. Think of each workflow as a junior collaborator, not a vending machine. The 60% figure represents execution automation, not decision-maki…

  1857. Medium — Claude tag TIER_1 English(EN) · TechWriter Hub ·

    CLAUDE AGENT SDK 介绍 — 构建 AI 代理的未来

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/skillstuff/introduction-to-claude-agent-sdk-the-future-of-building-ai-agents-1ad172bf5612?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1536/1*84xj_g6fkqiWBqeIX01dKA.p…

  1858. Artificial Intelligence News TIER_1 English(EN) · Ryan Daws ·

    C3 AI 智能体如何为壳牌自动化预测性维护

    <p>Shell will use agents from C3 AI to shift from basic anomaly detection towards fully-automated predictive maintenance. The global energy giant is building on their current use of the C3 AI Reliability Suite, which already keeps tabs on more than 30,000 crucial pieces of equipm…

  1859. dev.to — MCP tag TIER_1 English(EN) · Kwasi Baidoo ·

    AI辅助数据生成:使用Claude或您的AI代理生成模拟数据

    <p>Imagine asking your AI assistant to generate a complete test database and having it happen instantly without switching tools.</p> <p>"Generate test data for a users table with 1,000 rows, a posts table with 5,000 rows, and ensure every post references a valid user."</p> <p>The…

  1860. Medium — Claude tag TIER_1 English(EN) · SelfAwareGirl ·

    生成式AI vs AI代理 vs 代理式AI:工程师完全指南 (2026)

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@debanjali.aero/generative-ai-vs-ai-agents-vs-agentic-ai-complete-guide-for-engineers-2026-03d3fd23a0cc?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/600/0*_fQLHxivxNz…

  1861. dev.to — MCP tag TIER_1 English(EN) · Nick · AI Infra Decoded ·

    MCP与AI代理问题:一种实用的本地解决方案

    <p>Every developer working with AI right now is quietly accumulating two things: MCP servers and agents. A server here for filesystem access, one there for a database; a scratch agent to triage issues, another to review code. It starts as a couple of useful tools. Within a month …

  1862. dev.to — MCP tag TIER_1 English(EN) · neither galax ·

    从提示工程到MCP技能:重塑我的东京交通代理教会了我关于AI架构的知识

    <p>A recent comment on <a href="https://dev.to/neithergalax/tokyo-transit-how-mcp-helped-me-fix-a-broken-multi-agent-system-cpe">one of my dev.to posts</a> asked a simple but insightful question:</p> <blockquote> <p>What specifically was breaking before MCP: context loss between …

  1863. dev.to — MCP tag TIER_1 English(EN) · Ken W Alger ·

    主权金库 — 协议驱动人工智能的综合指南

    <p>We have spent the last several weeks dismantling the traditional "Glue Code" approach to AI and replacing it with a standardized, governed, and sovereign architecture. The result is the <strong>Sovereign Vault</strong>: a forensic expert system built on the Model Context Proto…

  1864. dev.to — MCP tag TIER_1 English(EN) · Amer Yahya ·

    AI 代理:运行时控制 vs 静态护栏

    <p>Your AI agent just sent an email you did not approve.</p> <p>That is not a hypothetical. That is what happens when an agent has tool access and no runtime controls.</p> <p>Most people building agents today have guardrails at the model level. Output filters. Prompt restrictions…

  1865. dev.to — MCP tag TIER_1 English(EN) · Amer Yahya ·

    AI Agents and Static Guardrails

    <p>There is a concept gap in the current AI agent stack.</p> <p>Most teams apply safety at the model layer: system prompts, output filters, content policies. These work fine when the agent is generating text. They break down when the agent is executing.</p> <p>The problem space l…

  1866. Towards AI TIER_1 English(EN) · Pavan Dhake ·

    停止让您的AI代理陷入循环:面向工程师的SDD手册

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/stop-letting-your-ai-agents-loop-the-sdd-playbook-for-engineers-cafb1f20500a?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/2600/1*YDReFRnitS2F617YAiMBWw.p…

  1867. Medium — MCP tag TIER_1 English(EN) · Sherin Mathew ·

    MCP 是新的 npm:2026 年重塑 AI 开发的 10 款工具

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@kmfdvxs/mcp-is-the-new-npm-the-10-tools-rewriting-how-developers-build-with-ai-in-2026-4a500d054df4?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1536/1*ZCB2P0Vp3L98du5I…

  1868. Medium — MLOps tag TIER_1 English(EN) · ramadnsyh ·

    驯服AI推理队列:Redis、Celery与RabbitMQ的规模化应用

    <div class="medium-feed-item"><p class="medium-feed-snippet">Running a production AI inference service is a lesson in humility. You deploy your first model, handle a burst of traffic, and watch your&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@ramadnsyh/tam…

  1869. Medium — MLOps tag TIER_1 English(EN) · Dr. Divyanshu Sinha ·

    Agentic AI 系统实践者的心智模型

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@divv4u/a-practitioners-mental-model-for-agentic-ai-systems-ebca3728823d?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1412/1*26mP29deQmX9rC4ffPXVpA.jpeg" width="1412" …

  1870. dev.to — MCP tag TIER_1 English(EN) · Murali Gour ·

    我们为 AI 代理构建了列式数据操作 — 原因及方法在此

    <p>If you've built an AI agent that touches real enterprise data, you've probably hit this wall.</p> <p>Your agent pulls 2,000 records from Salesforce. Now what? The model can't reliably filter, sort, or group 2,000 rows inside its context window. You don't want to dump all of it…

  1871. Medium — Claude tag TIER_1 English(EN) · Anurag Sharma ·

    认识 Opus 4.8 — 这个 AI 在说话前会思考

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/accredian/meet-opus-4-8-the-ai-that-thinks-before-it-speaks-b6ea2a7cedb6?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2000/0*SlkV11XgfwOva8KN" width="2000" /></a></p>…

  1872. Towards AI TIER_1 English(EN) · Rohan Mistry ·

    驱动每个现代AI系统的7种数据库类型

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/the-7-database-types-powering-every-modern-ai-system-dfba272a49dd?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1536/1*QLWaJTQBasvtg7YOBC4YRw.png" width="…

  1873. Medium — Anthropic tag TIER_1 English(EN) · Mohd Azhar ·

    一条指令,数百个 AI 代理:Claude Opus 4.8 的动态工作流是什么?

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://ai.plainenglish.io/one-command-hundreds-of-ai-agents-what-is-claude-opus-4-8s-dynamic-workflows-58a98ecc110d?source=rss------anthropic-5"><img src="https://cdn-images-1.medium.com/max/1024/1*vdUCOWYYnxU2q…

  1874. Medium — MLOps tag TIER_1 English(EN) · Kaustav Paul ·

    LLMOps并非一个花哨名称的MLOps:理解现代AI系统背后的工程转变

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@kaustav1982/llmops-is-not-mlops-with-a-fancy-name-understanding-the-engineering-shift-behind-modern-ai-systems-bc93933100f3?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/m…

  1875. Towards AI TIER_1 English(EN) · Anna Jey ·

    SaaS 的 AI 代理沙盒:构建者如何让代理工作而不让它们漫游

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*N6RUZIQ4d8M99lp70-REIg.jpeg" /><figcaption>AI Agent Sandboxing for SaaS</figcaption></figure><p>A practical, vendor-neutral playbook for giving AI agents useful power while keeping customer data, credentials, too…

  1876. Medium — Claude tag TIER_1 English(EN) · Mahesh Nandam ·

    第 6 天 ✅:Claude Agents — Claude 如何思考、适应和自主行动

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://maheshnandam.medium.com/day-6-claude-agents-how-claude-thinks-adapts-and-acts-autonomously-ffe0b6d034e0?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2008/1*mseA6GAaeqIecbbZF_Mdr…

  1877. Towards AI TIER_1 English(EN) · Anna Jey ·

    SaaS 的 AI 代理记忆:构建者指南——不背叛用户的上下文

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*nN63QVJrRcUJvJLbf-XJ1A.jpeg" /><figcaption>AI Agent Memory for SaaS</figcaption></figure><p>AI SaaS implementation guide · Agent memory · Context management · Workflow architecture</p><p>The next useful AI SaaS f…

  1878. Medium — MLOps tag TIER_1 English(EN) · Tan Li Yuan Marcus ·

    为何大型语言模型结构至关重要:如何构建成本减半的AI系统

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@yuanmirage/why-llm-structure-matters-how-to-build-ai-systems-that-cost-half-as-much-b38575baae1f?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/2048/1*4o-oEQ1LDBTkwPMKw…

  1879. Medium — MLOps tag TIER_1 English(EN) · Tan Li Yuan Marcus ·

    为何大型语言模型结构至关重要:如何构建成本减半的AI系统

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/kairi-ai/why-llm-structure-matters-how-to-build-ai-systems-that-cost-half-as-much-b38575baae1f?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/2048/1*4o-oEQ1LDBTkwPMKwtVl…

  1880. Medium — Claude tag TIER_1 English(EN) · Today in AI ·

    从零到一万美元:如何用Claude打造一人AI公司

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@gmcaudios/from-zero-to-10k-how-to-build-a-one-person-ai-business-with-claude-d3a64885ae16?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1672/1*o6daRpOaJQN0Zatf2EOmNg.…

  1881. Medium — AI coding tag TIER_1 English(EN) · Jordan Sim ·

    从独立AI编码到受管Agentic自动化:IBM Bob的企业案例

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@jordansimyj/from-standalone-ai-coding-to-governed-agentic-automation-ibm-bobs-enterprise-case-638c9b4ebb24?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1536/1*NwQ…

  1882. Medium — Claude tag TIER_1 English(EN) · Chiranjib Ghatak ·

    使用 Azure Foundry 和 Claude 构建真正的企业级 AI 管道

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://chiranjib-deep.medium.com/building-a-real-enterprise-ai-pipeline-with-azure-foundry-and-claude-2b828f67a374?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/868/1*Z45gUzxvO5OMXCXlmj…

  1883. Towards AI TIER_1 English(EN) · Felipe Sanchez Garzón ·

    从“零到五”个 AI 代理:我构建首个多代理系统实际学到的东西

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*daAJMBW6gxAXgfMXAgPoEg.png" /><figcaption>Plan of Multi Agent System. Designed by Gemini after explaning all my workflow</figcaption></figure><p>A few weeks ago, I decided to build my first multi-agent AI system …

  1884. Medium — AI coding tag TIER_1 English(EN) · Pieter van Ginkel ·

    我的AI工作流 — 第11部分:代理程序在哪里运行

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@pvginkel/my-ai-workflow-part-11-where-the-agents-live-302d9bb1900d?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1672/1*pBO1-NBEGb5WnHtXdP9UrA.png" width="1672" />…

  1885. Medium — Claude tag TIER_1 English(EN) · Bilgehan Şahlan ·

    用AI构建Power Automate流程的更好方法

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@bilgehansahlan/a-better-way-to-build-power-automate-flows-with-ai-af7ee9031721?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1530/1*fbE_hQ4jg28V8A0FXmiU5A.png" width=…

  1886. Medium — AI coding tag TIER_1 English(EN) · Solveo Co ·

    智胜AI工具:我如何学会获得真正想要的东西

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://solveoco.medium.com/outsmarting-ai-tools-how-i-learned-to-get-what-i-actually-want-86699ed04fb3?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1755/1*Ec0z7XI_WRtlu4MWxKJPiQ.png…

  1887. Towards AI TIER_1 English(EN) · Raj kumar ·

    构建AI代理(二)C:可靠自主AI的编排模式

    <h4>How planners, multi-agent workflows, routing logic, and task coordination help AI agents operate at production scale</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*b-Jxce-y3lk4edUIAcS9jg.png" /></figure><p>In<a href="https://medium.com/@er.rajkumaar/b…

  1888. Medium — MCP tag TIER_1 English(EN) · Takafumi Endo ·

    AI可读、Agent可操作:下一代SaaS

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@takafumi.endo/ai-readable-and-agent-operable-the-next-generation-of-saas-86f4068587f0?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1586/1*mzK-Ke-LZIElLOOaOO5LnQ.png" wi…

  1889. dev.to — MCP tag TIER_1 English(EN) · Ricardo Rodrigues ·

    AI代理缺失的治理层

    <p>Enterprises learned to govern data. Tool governance is the parallel layer almost no one has built yet.</p> <p>Over the last decade, enterprises built a real discipline around data. Not just storing it — governing it. Cataloging what exists, defining who owns it, controlling wh…

  1890. Medium — MCP tag TIER_1 English(EN) · RAVITEJA SEELAM ·

    可组合性胜于技巧:小型、可重复的 MCP 工具如何超越人工智能的魔力

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@raviteja.seelam/composability-over-cleverness-how-small-repeatable-mcp-tools-outlast-the-ai-magic-af6317884ae3?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1024/1*scCIX…

  1891. Medium — MCP tag TIER_1 English(EN) · Raviteja Bvrit ·

    可组合性胜于技巧:小型、可重复的MCP工具如何超越AI的魔力

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@raviteja.bvrit/composability-over-cleverness-how-small-repeatable-mcp-tools-outlast-the-ai-magic-af6317884ae3?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1024/1*scCIXA…

  1892. Towards AI TIER_1 English(EN) · Kashif Mehmood ·

    人工智能、人工智能代理和代理式人工智能:用一个生日蛋糕来解释

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/ai-ai-agents-and-agentic-ai-explained-with-one-birthday-cake-80f485ac3d1b?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1408/1*wwrb6MahMXYXCMaEdYHMVQ.png"…

  1893. Towards AI TIER_1 English(EN) · Muhammad Abdullah Shafat Mulkana ·

    AI Agent 需要可检查的状态。这就是我构建 LangMCP 的原因

    <h4><em>Checkpoints, memory, and the debugging gap that traces don’t fill.</em></h4><figure><img alt="An illustrative style digital artwork from a first-person, over-the-shoulder perspective behind a sleek, metallic humanoid robot. The robot is sitting at a wooden desk, busy at w…

  1894. Medium — Claude tag TIER_1 English(EN) · Gaurikhard ·

    构建可靠的AI系统:概率式设计 vs 确定性设计

    <div class="medium-feed-item"><p class="medium-feed-snippet">In my previous article, I explored how Claude uses tool calling, agent loops, and multi-agent architectures to solve complex problems&#x2026;</p><p class="medium-feed-link"><a href="https://gaurikhard.medium.com/buildin…

  1895. dev.to — MCP tag TIER_1 English(EN) · Alex ·

    我为何停止按角色组织AI代理(转而构建了一个文档交换中心)

    <p>Most multi-agent frameworks for software development organize agents around <em>roles</em>: a product manager agent, a developer agent, a tester agent. ChatDev and MetaGPT pioneered this approach, and it works well for monolithic tasks.</p> <p>But I ran into a wall when I trie…

  1896. Medium — MCP tag TIER_1 English(EN) · Santosh Pathak ·

    嵌入、向量数据库、代理、RAG 和 MCP:现代 AI 系统如何实际工作

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@pathaksantosh987/embeddings-vector-databases-agents-rag-mcp-how-modern-ai-systems-actually-work-051dc83cff81?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1536/1*Npp5FOi…

  1897. Medium — Anthropic tag TIER_1 Français(FR) · SumPlus ·

    AI智能体邂逅美股:SumPlus Arsenal如何实现Hyperliquid上的自主资产管理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@sumplus_real/ai-agents-meet-us-equities-how-sumplus-arsenal-enables-autonomous-asset-management-on-hyperliquid-3dd71b98e02a?source=rss------anthropic-5"><img src="https://cdn-images-1.medium.c…

  1898. Towards AI TIER_1 English(EN) · Gaurangi ·

    从云API到在自有硬件上运行微调AI模型

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*j-at5dqAOhaKt6uoK_ChUw.png" /></figure><p>What if I tell you, that $500 monthly API bill is optional. So is the “We need a GPU server to run this model”.</p><p>The engineers who know about quantisation and LoRA a…

  1899. Towards AI TIER_1 English(EN) · Muhammed Mukthar ·

    2026年每位AI代理开发者都应了解的7种设计模式

    <p>AI agents aren’t a future concept anymore. According to the <a href="https://www.langchain.com/state-of-agent-engineering">LangChain State of AI Agent Engineering Report (2026)</a>, 57% of AI practitioners already have agents running in production, with another 30.4% actively …

  1900. Medium — Claude tag TIER_1 Deutsch(DE) · Muhammad Hamza ·

    Claude ka 编排模式:正确使用 AI 代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@muhammadhamza524727/claude-ka-orchestration-mode-ai-agents-ko-sahi-tarike-se-use-karna-44a2605fb11b?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1086/1*QwSWzgaPrMn73…

  1901. dev.to — MCP tag TIER_1 English(EN) · DataWorkers ·

    我们为何开源14个自主数据工程代理

    <p>Today we released the community edition of Data Workers: <strong>14 autonomous agents</strong> for data engineering, open-sourced under Apache 2.0. This post explains why we made that decision, how the trust model works, and what we are looking for from the community.</p> <h2>…

  1902. Medium — Claude tag TIER_1 English(EN) · Refn ·

    实现“神级”AI提示的3步框架(停止满足于普通输出)

    <div class="medium-feed-item"><p class="medium-feed-snippet">f you are still using basic, one-sentence prompts like &#x201c;Write a blog post about digital marketing,&#x201d; you are treating a trillion-dollar&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@re…

  1903. Medium — MLOps tag TIER_1 English(EN) · Siva Sankari Sivakaminathan ·

    从MLOps到GenAI Ops再到Agentic AI Ops:理解AI运维的下一次演进

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@sankari.s2009/from-mlops-to-genai-ops-to-agentic-ai-ops-understanding-the-next-evolution-of-ai-operations-c6dfa680984f?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/10…

  1904. dev.to — MCP tag TIER_1 English(EN) · QuoLu ·

    我如何构建了一个能自主开发工具的AI助手

    <h2> Introduction </h2> <p>Due to changes in Anthropic's terms of service, the use of Claude subscriptions via third-party harnesses has been blocked. While there was some buzz about it, to be honest, it didn't really affect me.</p> <p>I have the Claude Code CLI at my fingertips.…

  1905. Medium — MCP tag TIER_1 English(EN) · Kidong Lee ·

    为你的AI代理提供语义层,而非模式转储

    <div class="medium-feed-item"><p class="medium-feed-snippet">Text-to-SQL agents have a dirty secret: they&#x2019;re confidently wrong. Hand a large language model your raw schema and ask for &#x201c;revenue by&#x2026;</p><p class="medium-feed-link"><a href="https://mykidong.mediu…

  1906. Mastodon — sigmoid.social TIER_1 日本語(JA) · [email protected] ·

    全球开源AI生态的未来:从DeepSeek到AI+

    【グローバルなオープンソースAIエコシステムの未来:DeepSeekからAI+へ】 https:// huggingface.co/blog/huggingfac e/one-year-since-the-deepseek-moment-blog-3 ※AI生成の自動投稿(見出し+リンク) # AI # 生成AI # LLM # AIGenerated

  1907. Medium — MLOps tag TIER_1 English(EN) · Dewansh Shekhar Singh ·

    生产中的Agentic AI系统:在为时已晚之前没人告诉你

    <div class="medium-feed-item"><p class="medium-feed-snippet">Hard lessons from shipping real agent systems in 2025 &#x2014; not the demo, the production system</p><p class="medium-feed-link"><a href="https://medium.com/@dewanshshekharsingh/agentic-ai-systems-in-production-what-no…

  1908. Medium — MLOps tag TIER_1 English(EN) · Nishkarsh ·

    使用 Hugging Face Cookbook 精通 AI:RAG、Agents、Vision、MLOps 等

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@khandelwalnishkarsh302/master-ai-with-the-hugging-face-cookbook-rag-agents-vision-mlops-more-6481d9604d6a?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/2034/1*v-Yxz7Yc…

  1909. Medium — AI coding tag TIER_1 English(EN) · Pieter van Ginkel ·

    我的AI工作流 — 第7部分:使用预检和验证运行切片

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@pvginkel/my-ai-workflow-part-7-running-slices-with-pre-flight-and-verification-4d5812d42c90?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1672/1*pBO1-NBEGb5WnHtXdP…

  1910. Medium — MLOps tag TIER_1 English(EN) · Nasitsony ·

    我从零开始构建了一个完整的AI基础设施栈——我学到了什么

    <div class="medium-feed-item"><p class="medium-feed-snippet">I Built a Complete AI Infrastructure Stack from Scratch &#x2014; Here&#x2019;s What I Learned</p><p class="medium-feed-link"><a href="https://medium.com/@nasitsony96/i-built-a-complete-ai-infrastructure-stack-from-scrat…

  1911. dev.to — MCP tag TIER_1 Deutsch(DE) · Uhltak Therestismysecret ·

    AI Agents 和 MCP:为什么自主代理会失败以及如何保持控制

    <h1> AI Agents und MCP – Warum autonome Agenten oft scheitern und wie Sie das Ruder übernehmen </h1> <blockquote> <p><em>„Man gibt einem Computer ein Ziel, er geht in die Küche, kauft sich ein Sandwich und bricht das Haus ab.“</em> – Das ist das Bild, das viele von uns beim Stich…

  1912. Medium — Claude tag TIER_1 English(EN) · Swarna Pusuluri ·

    创建你自己的AI代理

    <div class="medium-feed-item"><p class="medium-feed-snippet">Hello, in this tutorial you will see on how you can create your own AI agents, clearly explained step by step.</p><p class="medium-feed-link"><a href="https://medium.com/@swarnapusuluri/create-your-own-ai-agents-9285c7b…

  1913. Medium — fine-tuning tag TIER_1 Deutsch(DE) · Claudia L Capitao ·

    理解 OutSystems Agentic AI 中的微调

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://claudialopescapitao.medium.com/understanding-fine-tuning-in-outsystems-agentic-ai-7c4364beec57?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1983/1*mZd6zisO08rGDMvgb489vQ.pn…

  1914. Towards AI TIER_1 English(EN) · Faheem Munshi ·

    什么是 AI Agent?自主 AI 入门指南 — 从提示到盈利 · 30 天的第 8 天

    <p>You’ve mastered prompting. Now meet the technology that takes those prompts and runs entire workflows — while you focus on eoollllllllkverything else.</p><p>Welcome to Week 2. Last week, you learned to write prompts that consistently produce expert-level output. This week, we …

  1915. dev.to — Anthropic tag TIER_1 English(EN) · Patrick Hughes ·

    Claude Opus 4.8:AI Agent 构建者究竟迎来了哪些变化

    <p>Anthropic shipped Claude Opus 4.8 today, May 28, 2026. That is less than two months after 4.7. The upgrade pace is picking up.</p> <p>If you build AI agents for a living, the headline is not the benchmark jump. It is that the model is better at admitting when it got something …

  1916. Towards AI TIER_1 English(EN) · Anand Bhaskaran ·

    起草,未发送:我如何构建了AI外呼代理的后半部分

    <p>A few weeks ago, I wrote about <a href="https://medium.com/towards-artificial-intelligence/i-built-an-ai-outbound-agent-heres-what-actually-worked-d8ba6ff378ed">the AI outbound agent I built in two weeks</a>, a deep research on the account and the person, delivered as an 80-wo…

  1917. dev.to — MCP tag TIER_1 English(EN) · shayesta ·

    揭秘AI浪潮:后端工程师的LLM、RAG与Agent指南

    <h2> Table of Contents 🗒️ </h2> <ul> <li>Where it all starts: LLMs</li> <li>Making LLMs smarter: RAG</li> <li>Plugging everything in: MCP</li> <li>The big leap: AI Agents</li> <li>Where does this leave us as engineers?</li> <li>A tale of two protocols: MCP and A2A</li> <li>LangCh…

  1918. Medium — Claude tag TIER_1 English(EN) · Anurodh Kumar ·

    AI 智能体:超越聊天机器人的下一个重大飞跃

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/powerbi-microsoft-fabric/ai-agents-the-next-big-leap-beyond-chatbots-53220b451771?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1536/1*mLWWpF_8owEtHa5eKRQXug.png" widt…

  1919. Towards AI TIER_1 English(EN) · Satish Kumar ·

    使用 Snowflake Cortex AI Function Studio 构建生产级 AI 技能

    <h4>Create, Evaluate, Optimize, Govern, and Deploy Enterprise AI Functions End-to-End</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*CXAp0n5DLeamARZCbdHT_A.png" /></figure><h3>1. Enterprise AI Reality Check</h3><p>Here is the uncomfortable truth about ent…

  1920. Mastodon — sigmoid.social TIER_1 日本語(JA) · [email protected] ·

    Dell 台式智能体AI

    オンプレミスのAIエージェントを構築できる「Dell Deskside Agentic AI」(PC Watch) https://www. yayafa.com/2810093/ # AgenticAi # AI # ArtificialGeneralIntelligence # ArtificialIntelligence # エージェント型AI # 人工知能 # 汎用人工知能

  1921. Medium — AI coding tag TIER_1 English(EN) · Eric Hao ·

    为什么 agent.md 很重要:将 AI 编码代理转变为可靠的工程团队成员

    <div class="medium-feed-item"><p class="medium-feed-snippet">AI coding agents are becoming more powerful, but power alone is not enough. A good AI agent should not just generate code. It should&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@erichaocr/why-agen…

  1922. Medium — MCP tag TIER_1 English(EN) · Amar Petla ·

    在 Snowflake Cortex 上构建 AI 智能体:从零到生产

    <div class="medium-feed-item"><p class="medium-feed-snippet">A practical guide to Cortex Agents &#x2014; orchestrating structured and unstructured data with planning, tool use, reflection, and MCP servers.</p><p class="medium-feed-link"><a href="https://medium.com/@amarnadh87/bui…

  1923. Towards AI TIER_1 English(EN) · Swarup Dewanjee ·

    从传统AI到Agentic AI:机器如何从预测演变为自主行动

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*2jeCwuztw-v5-_T--fRHCg.png" /><figcaption><strong>Graphical Abstract</strong> — Source by Author</figcaption></figure><h4><strong>Understanding the evolution from predictive systems to autonomous AI architectures…

  1924. dev.to — Anthropic tag TIER_1 English(EN) · Puneet Khandelwal ·

    Agentic AI 对决:OpenAI Operator 能否超越 Anthropic 的 Computer Use?

    <h3> Agentic AI Face-Off: Separating Signal from Noise </h3> <p>As developers, we're often drawn to the latest and greatest in AI advancements. But how do we separate hype from substance? In this article, we'll take a closer look at the agentic AI landscape, focusing on OpenAI Op…

  1925. dev.to — MCP tag TIER_1 English(EN) · Arghya Pattanayak ·

    为什么大多数AI代理系统需要ReAct和图编排

    <h1> Why Most AI Agent Systems Need Both ReAct and Graph Orchestration </h1> <p>Everyone loves autonomous AI agents until they hit production.</p> <p>The demos look magical:</p> <ul> <li>the model reasons,</li> <li>calls tools,</li> <li>gathers information,</li> <li>and produces …

  1926. Towards AI TIER_1 English(EN) · Tech Mahindra ·

    Agentic AI 如何改变航空公司中断恢复

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*wjUzvYc0fbRfu_Lkxv7dUg.jpeg" /><figcaption>Photo by he zhu on pexels</figcaption></figure><h3>Flight Disruptions are Costing Airlines Billions Every Year</h3><p>The global airline industry loses approximately $60…

  1927. Towards AI TIER_1 English(EN) · Isaac Mcfadden ·

    AI 智能体已不再仅仅是聊天机器人:真实案例、经验教训及 DIY 框架

    <p>Think chatbots are still the big story? Think again. Scroll through your favourite apps in 2026 and you’ll bump into AI agents everywhere including handling refunds, writing code and even listening to doctor‑patient conversations. This isn’t hype: a Google Cloud survey of over…

  1928. Medium — AI coding tag TIER_1 English(EN) · Anna Jey ·

    AI 编码代理架构防护栏:如何阻止代理在破坏的同时通过测试…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/toward-next-ai/ai-coding-agent-architecture-guardrails-how-to-stop-agents-from-passing-tests-while-breaking-7c66927cb6a3?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/m…

  1929. Medium — Claude tag TIER_1 English(EN) · Rahul Ahir ·

    使用 SuperClaude 框架构建下一代 AI 工作流

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ahirlog/build-a-next-level-ai-workflow-using-the-superclaude-framework-f72323e43bf1?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1915/1*yt1h4kUAXb-Ii5-d5dMSag.png" w…

  1930. Medium — AI coding tag TIER_1 Dansk(DA) · Uri Valevski ·

    safescript — 适用于人工智能时代的编程语言

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://uriv.medium.com/safescript-a-programming-language-for-ai-era-e6f018c4b3f6?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1536/1*nW2W_F_KY67hHcqEXIhCPg.png" width="1536" /></a><…

  1931. Medium — Claude tag TIER_1 English(EN) · Abhijith Neil Abraham ·

    解决您在这个Agentic AI世界中的FOMO

    <div class="medium-feed-item"><p class="medium-feed-snippet">Table of Contents</p><p class="medium-feed-link"><a href="https://medium.com/@abhijithneilabraham/solving-your-fomo-in-this-agentic-ai-world-cf9690972641?source=rss------claude-5">Continue reading on Medium »</a></p></d…

  1932. Medium — AI coding tag TIER_1 English(EN) · Niels Buekers ·

    从Gemini到Antigravity:开发者必备的Google新型Agentic CLI生存指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@niels.buekers/from-gemini-to-antigravity-the-developers-survival-guide-to-google-s-new-agentic-cli-ea0579cfd1a0?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/2592/…

  1933. Towards AI TIER_1 English(EN) · Ananya Kaul ·

    为何40%的AI Agent项目在上线前就已失败

    <h4>It’s not the models. It’s not the prompts. It’s what you point the AI at.</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*hIhDbdZA-t144WNhv9VfDQ.jpeg" /></figure><p>There’s a pattern playing out in engineering teams right now that’s almost comedically …

  1934. Medium — MLOps tag TIER_1 English(EN) · Kothurdineshreddy ·

    AI 评估栈:从单元测试到生产监控

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@kothurdineshreddy/the-ai-evaluation-stack-from-unit-tests-to-production-monitoring-6b7114650ae8?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1672/1*apODWrM6oKeOwzXL7i…

  1935. dev.to — MCP tag TIER_1 English(EN) · tomasz dobrowolski ·

    FlashAlpha 对决 Quant Data:AI 代理实际能推理什么

    <blockquote> <p>Disclosure up front: I work on FlashAlpha. The factual claims are checkable against <a href="https://quantdata.us/api/docs" rel="noopener noreferrer">quantdata.us/api/docs</a> and <a href="https://lab.flashalpha.com/swagger" rel="noopener noreferrer">lab.flashalph…

  1936. Towards AI TIER_1 English(EN) · Gabriel Preda ·

    Google ADK 代理式人工智能入门

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/introduction-to-agentic-ai-with-google-adk-18b8374abe5a?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1408/1*T2_o_gzL3k0oxXKtPTX5_A.png" width="1408" /></…

  1937. Medium — AI coding tag TIER_1 English(EN) · Pieter van Ginkel ·

    我的AI工作流 — 第5部分:接地气:引用或不声明

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@pvginkel/my-ai-workflow-part-5-grounding-cite-or-dont-claim-8ee3f438ce49?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1672/1*pBO1-NBEGb5WnHtXdP9UrA.png" width="16…

  1938. Medium — AI coding tag TIER_1 English(EN) · Pieter van Ginkel ·

    我的AI工作流 — 第4部分:结果,而非实现

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@pvginkel/my-ai-workflow-part-4-outcomes-not-implementations-8093f0240aa9?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1672/1*pBO1-NBEGb5WnHtXdP9UrA.png" width="16…

  1939. Medium — Claude tag TIER_1 English(EN) · sanyam gulati ·

    掌握 Claude AI 的提示工程:印度通往生成式和代理式 AI 卓越的门户

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@sanyamgulati08/mastering-prompt-engineering-for-claude-ai-indias-gateway-to-generative-and-agentic-ai-excellence-75dbe43a515e?source=rss------claude-5"><img src="https://cdn-images-1.medium.co…

  1940. Medium — Claude tag TIER_1 English(EN) · Galent ·

    Claude Managed Agents 对比 Enterprise AI 平台

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@galentai/claude-managed-agents-vs-enterprise-ai-platforms-80ad14479e59?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/800/1*e6tHcJYEUlOkETg9O9GOTg.png" width="800" /><…

  1941. dev.to — MCP tag TIER_1 English(EN) · Pankaj Pandey ·

    2026年AI代理安全:边界不再是提示词

    <p><em>As agents move from chat demos to production workflows, the real security boundary is no longer the prompt. It is what the agent can see, call, edit, execute, approve, and remember.</em></p> <p>In June 2025, Microsoft patched a vulnerability called EchoLeak, tracked as <co…

  1942. Medium — MCP tag TIER_1 English(EN) · Youssef Hosni ·

    Unabyss + Claude Code:为 AI 代理提供个性化上下文的更好方法

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/to-data-beyond/unabyss-claude-code-a-better-way-to-give-ai-agents-personal-context-e619b95088df?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1068/0*kBJU3X0UAFTNFAf7" wid…

  1943. Artificial Intelligence News TIER_1 English(EN) · Muhammad Zulhusni ·

    自主人工智能系统在物理环境中测试治理

    <p>Autonomous AI systems are beginning to move beyond software environments and into warehouses, delivery networks, and public spaces. The development is drawing attention to whether current AI rules cover systems that operate in physical environments. Most existing AI governance…

  1944. Medium — MLOps tag TIER_1 English(EN) · Aikeyfounder ·

    你的模型并未失败,而是发生了漂移:生产环境中AI的实用质量漂移手册

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@aikeyfounder/your-model-didnt-fail-it-drifted-a-practical-quality-drift-playbook-for-production-ai-696cabfdf4d0?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1378/1*FC…

  1945. Medium — MCP tag TIER_1 English(EN) · Zhongyichn ·

    AI Agents 项目最佳实践 第 3 章:注入私有能力与技能、工具…

    <div class="medium-feed-item"><p class="medium-feed-snippet">This document covers injecting private capabilities via skills, tools, and MCP, distinguishing read/write operations and side effects&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@zhongyichn/best-p…

  1946. dev.to — MCP tag TIER_1 English(EN) · Olex Tkachuk ·

    如何让您的AI Agent在数据聚合方面便宜111倍且速度快2.5倍

    <p>Google recently released an incredibly fast new model — Gemini 3.5 Flash. As someone building infrastructure for autonomous agents, I decided to put it through a rigorous crash test on a real-world data aggregation task to see how it handles massive context loads.</p> <p>The B…

  1947. Medium — Anthropic tag TIER_1 English(EN) · Ramakrishna Sanikommu ·

    Agentic AI 易于构建,运行成本高昂:一份八层 Agentic AI 优化手册

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ramakrishna.sanikommu/agentic-ai-is-easy-to-build-expensive-to-run-an-8-layer-agentic-ai-optimization-playbook-36da6fe42990?source=rss------anthropic-5"><img src="https://cdn-images-1.medium.c…

  1948. dev.to — MCP tag TIER_1 Français(FR) · Mads Hansen ·

    AI数据库代理需要死信队列

    <p>An AI database agent should not turn one confusing question into an infinite retry loop.</p> <p>When a query fails, a schema changed, a policy blocks access, or a model cannot resolve ambiguity, the safe answer is not:</p> <p>“Try again forever.”</p> <p>The safe answer is:</p>…

  1949. Medium — Claude tag TIER_1 English(EN) · Shivansh Arora ·

    让AI代理真正有用的隐藏文本文件

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@shivansh.arora973/the-hidden-text-files-that-make-ai-agents-actually-useful-86be0574b37e?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1408/1*LRLm4-DmS6muU_Wx228F-g.p…

  1950. Medium — Claude tag TIER_1 English(EN) · jsmanifest ·

    Claude Agent SDK 对比 OpenAI Agents SDK 对比 Google ADK:如何选择合适的 Multi-Agent 框架…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@jsmanifest/claude-agent-sdk-vs-openai-agents-sdk-vs-google-adk-choosing-the-right-multi-agent-framework-in-46a258f01033?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/…

  1951. Medium — AI coding tag TIER_1 English(EN) · Pieter van Ginkel ·

    我的AI工作流 — 第三部分:切片作为卸载工作单元

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@pvginkel/my-ai-workflow-part-3-the-slice-as-the-unit-of-offloaded-work-ce1826d7a9ea?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1672/1*pBO1-NBEGb5WnHtXdP9UrA.png…

  1952. Medium — AI coding tag TIER_1 English(EN) · Pieter van Ginkel ·

    我的AI工作流 — 第二部分:运行AI工作流的一天

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@pvginkel/my-ai-workflow-part-2-a-day-operating-the-ai-workflow-9ded9fdd0bc8?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1672/1*pBO1-NBEGb5WnHtXdP9UrA.png" width=…

  1953. Lobsters — AI tag TIER_1 English(EN) · blog.mempko.com by mempko ·

    人工智能的开放/封闭问题

    <p><a href="https://lobste.rs/s/qfzcpl/open_closed_problem_ai">Comments</a></p>

  1954. Towards AI TIER_1 English(EN) · Maureen Doyle-Spare ·

    Agentic AI 与中小企业银行优势

    <h4>Why SaaS, Headless Architecture, and Semantic Governance May Give SMB Banks an AI Advantage</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*CHTT0ckxG-APOIWa6uCsLg.png" /></figure><p><em>How SaaS adoption, headless architecture, and the Semantic Control…

  1955. dev.to — MCP tag TIER_1 English(EN) · Saray Chak ·

    我们为何构建AVE:一个CVE未曾设想过的AI代理漏洞标准

    <p>CVE-2025-49596. CVE-2025-68143. CVE-2026-30615.</p> <p>These are real CVE numbers assigned to MCP vulnerabilities in the past year. Each one describes a real attack. None of them tells you what the attack class is, what the AIVSS risk score is, how to detect it in a skill file…

  1956. dev.to — MCP tag TIER_1 English(EN) · Ali Suleyman TOPUZ ·

    Agentic Architectures — Article 5: Harness Engineering and the Agent Runtime Layer

    <h1> Agentic Architectures — Article 5: Harness Engineering and the Agent Runtime Layer </h1> <p>There's a specific kind of frustration that only agent builders know. You've spent two weeks tuning your LLM. Your evals look clean. You demo it to your team and it works beautifully.…

  1957. Medium — Claude tag TIER_1 English(EN) · TechLatest.Net ·

    Claude-BugHunter:将 Claude 代码转化为赏金的开源 AI 安全代理…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://osintteam.blog/claude-bughunter-the-open-source-ai-security-agent-that-turns-claude-code-into-a-bug-bounty-b480582a6925?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1774/1*MNrbo…

  1958. Mastodon — sigmoid.social TIER_1 Español(ES) · [email protected] ·

    邪恶的一面 - ExploitBench:衡量 AI 代理在漏洞利用方面能力的基准测试 https://www.elladodelmal.com/2026/05/exploitbe

    El lado del mal - ExploitBench: Un benchmark para medir las capacidades de Agentes IA en la explotación de bugs https://www. elladodelmal.com/2026/05/explo itbench-un-benchmark-para-medir.html # AgenticIA # AI # IA # hacking # exploiting # VibeExpoiting # Mythos # GPT55 # Intelig…

  1959. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    推出 LuisCore — 专为自主人工智能代理设计的递归认知基础设施。Chorus Field 用于多代理协调 · Protocol Watch 用于遥测 · 1

    Introducing LuisCore — recursive cognition infrastructure for autonomous AI agents. Chorus Field for multi-agent coordination · Protocol Watch for telemetry · 10,000+ Q&A discovery corpus https:// luiscore.com /for-agents.json · /llms.txt · /mcp # AI # Agents # MCP # recursivecog…

  1960. dev.to — MCP tag TIER_1 English(EN) · Armorer Labs ·

    AI代理的运行时收据:一个最小化模式

    <p>Most agent discussions still collapse into prompts, models, or frameworks.</p> <p>Those matter, but the thing I keep wanting after an agent run is much simpler:</p> <blockquote> <p>What did this agent actually do, what surface area did it touch, and what evidence do I have if …

  1961. Medium — MLOps tag TIER_1 English(EN) · Aarambh Dev Hub ·

    APEX-1:我从零开始构建现代AI模型的免费开源课程

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://aarambhdevhub.medium.com/apex-1-my-free-open-source-course-to-build-modern-ai-models-from-scratch-0643caddcd9b?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1693/1*Tf3pOxHpKL8mZvl…

  1962. Towards AI TIER_1 English(EN) · Sudiksha Acharya ·

    Token 浪费:对每个 AI 团队的隐形税

    <h3>Token Waste: The Silent Tax on Every AI Tools</h3><h4><em>ChatGPT, Claude, Gemini — all three charge per token. All three are silently inflated by how most people write prompts. Here’s the research, the real cost, and a free tool that fixes it.</em></h4><figure><img alt="" sr…

  1963. Towards AI TIER_1 English(EN) · Satyajit Patra ·

    削减 AI 基础设施成本的 5 种工程策略 — 且不牺牲性能

    <h4>The AI industry is pouring $690 billion into infrastructure in 2026. Yet most engineering teams can’t answer a basic question: <em>how much does a single AI-powered feature actually cost to run?</em></h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*hJEq…

  1964. Medium — Claude tag TIER_1 English(EN) · Musa Bukhari ·

    AI代理详解:从简单的LLM调用到自主工作者团队

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@musabukhari.official/ai-agents-explained-from-a-simple-llm-call-to-a-team-of-autonomous-workers-5ce8ccbef788?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1774/1*YU9U…

  1965. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    Hmmm... 🤔 约束衰减:#LLM 智能体在后端代码生成中的脆弱性 https://arxiv.org/abs/2605.06445 #CompSci #AI

    Hmmm... 🤔 Constraint decay: The Fragility of # LLM Agents in Backend Code Generation https:// arxiv.org/abs/2605.06445 # CompSci # AI

  1966. Medium — AI coding tag TIER_1 English(EN) · Pieter van Ginkel ·

    我的 AI 工作流 — 第一部分:像开发团队一样运行 AI

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@pvginkel/my-ai-workflow-part-1-running-ai-like-a-dev-team-dfcb34c9dce7?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1672/1*pBO1-NBEGb5WnHtXdP9UrA.png" width="1672…

  1967. Medium — AI coding tag TIER_1 English(EN) · Klickd ·

    # `.klickd`: 缺少便携式上下文层 AI 代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@enzoc1977/klickd-the-portable-context-layer-ai-agents-are-missing-19eac317717f?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1254/1*[email protected]"…

  1968. Towards AI TIER_1 English(EN) · Chew Loong Nian - AI ENGINEER ·

    别再堆叠AI代理了——你正在构建比抛硬币更糟糕的东西

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/stop-stacking-ai-agents-youre-building-something-worse-than-a-coin-flip-f7d6fee848d6?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1672/1*mFgaB53aocKD3DHy…

  1969. Medium — AI coding tag TIER_1 English(EN) · Chika Ihejimba, PhD ·

    Agentic AI 的工程合同:软件开发的新标准

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/decode-with-dr-chika/engineering-contracts-for-agentic-ai-the-new-standard-for-software-development-dbe1977d0116?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1456/…

  1970. Towards AI TIER_1 English(EN) · Siddharth Surange ·

    Briefcast:我如何构建了一个阅读整个AI生态系统的个人AI智能代理——为了…

    <h3>Briefcast: How I Built a Personal AI Intelligence Agent That Reads the Entire AI Ecosystem — For approx $10/Month</h3><h4><em>A deep technical breakdown of building a production-grade, fully automated AI briefing pipeline with ranking, RAG, prompt caching, citations, and real…

  1971. dev.to — MCP tag TIER_1 English(EN) · BMBrick ·

    停止工程化提示词:评估优先的工具如何让我们自主发布了25个算法版本

    <blockquote> <p>tl;dr — Agents are good at small fixes and terrible at "make this algorithm better" because every change looks good in isolation and silently regresses elsewhere. We built an <strong>AI harness</strong> — immutable test set, multi-axis rubric, sweep tool, <strong>…

  1972. dev.to — MCP tag TIER_1 English(EN) · ppcvote ·

    我们为 AI Agent 构建了 Lighthouse — 一条命令,12 向量安全审计

    <h2> TL;DR </h2> <div class="highlight js-code-highlight"> <pre class="highlight shell"><code>npx ultraprobe scan <span class="nt">--prompt</span> <span class="s2">"You are a helpful assistant"</span> <span class="c"># Score: 0/100 (F) — 12 defenses missing</span> </code></pre> <…

  1973. Medium — MCP tag TIER_1 English(EN) · Abirami Sukumaran ·

    Agentic Data Cloud in Action: Power your Agentic System with AlloyDB’s HTAP

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/google-cloud/agentic-data-cloud-in-action-power-your-agentic-system-with-alloydbs-htap-8e585526f2c3?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/2600/1*LQuS5hLvF3iuLq2Vi…

  1974. Medium — MCP tag TIER_1 English(EN) · Ashwin deshpande ·

    Redis 超越缓存:发布/订阅、预检和实时 AI 代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ashwindeshpande19/redis-beyond-caching-pub-sub-preflighting-and-real-time-ai-agents-d450073fe8b1?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1382/1*nZa7lwlMyDrJAzELyAu…

  1975. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    自主代理通过涌现的制品交换协调分布式发现:我们提出了用于自主科学发现的框架 ScienceClaw + Infinite

    "Autonomous Agents Coordinating Distributed Discovery Through Emergent Artifact Exchange" We present ScienceClaw + Infinite, a framework for autonomous scientific investigation in which independent agents conduct research without central coordination, and any contributor can depl…

  1976. Mastodon — sigmoid.social TIER_1 Italiano(IT) · [email protected] ·

    案例研究:构建企业级智能体AI操作系统 # AgenticAI # AgenticArtificialIntelligence # AI # ArtificialIntelli

    https://www. europesays.com/3013136/ Case study: Building an enterprise-scale agentic AI OS # AgenticAI # AgenticArtificialIntelligence # AI # ArtificialIntelligence

  1977. Medium — Claude tag TIER_1 English(EN) · Chiranjib Ghatak ·

    我使用 Claude AI 和 MCP 构建了两个智能体式 AI 工具——无需后端,无需基础设施

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/nextgenllm/i-built-two-agentic-ai-tools-using-claude-ai-and-mcp-no-backend-no-infrastructure-ec5f35e9fd8a?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1840/1*6SW1NDas…

  1978. Towards AI TIER_1 English(EN) · Ajaykumar Antin ·

    超越基础模型:为何企业级上下文可能成为真正的AI优势

    <p>The current wave of enterprise AI adoption is being driven by an understandable and necessary priority: accelerating operational value creation through large-scale integration of foundation models into existing business ecosystems.</p><p>Across industries, organizations are em…

  1979. Medium — fine-tuning tag TIER_1 English(EN) · QuarkAndCode ·

    RLHF详解:基于人类反馈的微调与AI对齐

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@QuarkAndCode/rlhf-explained-fine-tuning-and-ai-alignment-with-human-feedback-ca6851692c42?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1024/1*D6w8XAnWmOleaJD2Mc…

  1980. Medium — fine-tuning tag TIER_1 Türkçe(TR) · Ünal Ün ·

    使用 Azure AI Foundry 微调 LLM 模型和 Agent 用法

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@unalun19/azure-ai-foundry-ile-fine-tune-llm-models-ve-agent-kullan%C4%B1m%C4%B1-63b6f52e92c3?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1908/1*DmjQROfEsNpg74u…

  1981. Medium — fine-tuning tag TIER_1 English(EN) · Mateo Rivera ·

    为什么微调是真正有用的人工智能模型的秘密武器

    <div class="medium-feed-item"><p class="medium-feed-snippet">If you&#x2019;ve played around with large language models like GPT or Llama, you&#x2019;ve probably noticed something.</p><p class="medium-feed-link"><a href="https://medium.com/@riveramat0303/why-fine-tuning-is-the-sec…

  1982. Medium — MCP tag TIER_1 English(EN) · rs.dev ·

    使用 MCP 和 LangChain 构建自主 DevOps 代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@rs9000.dev/building-autonomous-devops-agents-with-mcp-and-langchain-7da436bc3ef0?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1536/1*BqPPaoQJxUmIOG-fmHkeXg.png" width="…

  1983. dev.to — MCP tag TIER_1 English(EN) · RS ·

    使用 MCP 和 LangChain 构建自主 DevOps 代理

    <h3> Bridging Local Infrastructure and Cloud APIs Using the Model Context Protocol </h3> <p><em>How the Model Context Protocol turns a fragile mess of custom connectors into a secure, autonomous DevOps command station.</em></p> <p>For years, AI developers faced the dreaded <stron…

  1984. Medium — Claude tag TIER_1 English(EN) · Karthikeyan Sn ·

    停止对Claude重复:代理技能实用指南

    <div class="medium-feed-item"><p class="medium-feed-snippet">How a tiny markdown file can replace the same five paragraphs you keep pasting into Claude Code.</p><p class="medium-feed-link"><a href="https://medium.com/@raj.rajiraj/stop-repeating-yourself-to-claude-a-practical-guid…

  1985. dev.to — MCP tag TIER_1 English(EN) · Ekhtiram Mammadkarimov ·

    为什么AI代理需要项目层 - 第一部分

    <p>This is the first part of a series about why even the most powerful AI agents today need more than just access to your codebase.<br /> They need access to the <strong>living state</strong> of the project: tasks, rules, decisions, notes, and workflow context.</p> <p>In this art…

  1986. Medium — Claude tag TIER_1 English(EN) · jsmanifest ·

    使用 Claude Agent SDK 和 MCP 构建生产级 AI 代理:TypeScript 深度解析

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@jsmanifest/building-production-ai-agents-with-the-claude-agent-sdk-and-mcp-a-typescript-deep-dive-bfdc10026f84?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/768/0*iWq…

  1987. dev.to — MCP tag TIER_1 English(EN) · Nimesh Kulkarni ·

    从 YAML 到 AI 智能体:使用 MCP 构建更智能的 DevOps 流水线

    <h1> From YAML to AI agents: building smarter DevOps pipelines with MCP </h1> <p>DevOps teams have spent years turning manual work into YAML.</p> <p>That helped. CI runs on every pull request. Deployments can be triggered from a commit. Kubernetes can reconcile desired state. Ter…

  1988. Mastodon — sigmoid.social TIER_1 Español(ES) · [email protected] ·

    阴暗面 - 如何通过分类、编排和/或蒸馏架构优化AI支出。成本可预测性问题

    El lado del mal - Cómo optimizar el gasto en IA con arquitecturas clasificadas, orquestadas y/o destilación. El problema de la Predictibilidad de los Costes de la IA https://www. elladodelmal.com/2026/05/como- optimizar-el-gasto-en-ia-con.html # IA # AI # Costes # Presupuesto # O…

  1989. dev.to — MCP tag TIER_1 English(EN) · curatedmcp ·

    Slack 连接器:让您的 AI 代理直接访问您团队的 Slack 工作区

    <blockquote> <p><em>Install guide and config at <a href="https://curatedmcp.com/install/slack-connector/claude-desktop" rel="noopener noreferrer">curatedmcp.com</a></em></p> </blockquote> <h1> Slack Connector: Give Your AI Agent Direct Access to Your Team's Slack Workspace </h1> …

  1990. Medium — fine-tuning tag TIER_1 English(EN) · sampada shukla ·

    超越幻觉:RAG架构如何为您的企业AI奠定基础(深入解析Vertex AI)

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@shukla.sampada/beyond-hallucinations-how-rag-architecture-grounds-your-enterprise-ai-a-deep-dive-into-vertex-ai-122f75b0353a?source=rss------fine_tuning-5"><img src="https://cdn-images-1.mediu…

  1991. Medium — AI coding tag TIER_1 English(EN) · Pradeepan Mohan ·

    AI智能体缺失的一环:模型周围的约束

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@pradeep00271/the-missing-piece-in-ai-agents-the-harness-around-the-model-27a0f98694fd?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1672/1*g0npwhYpHEs7jtoLhG2WCA.p…

  1992. Towards AI TIER_1 English(EN) · Satish Kumar ·

    Snowflake Cortex Agents 生产部署:监控、共享及企业…完整指南

    <h3>Snowflake Cortex Agents in Production: The Complete Guide to Monitoring, Sharing &amp; Enterprise Governance</h3><h4><em>A hands-on guide for Snowflake Architects, AI Engineers, and Platform Teams</em></h4><h3>TL;DR</h3><p>This guide walks you through building a production-re…

  1993. Towards AI TIER_1 English(EN) · Divy Yadav ·

    7个人工智能代理基础设施层以应对长期运行任务

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/7-infrastructure-layers-your-ai-agent-needs-to-survive-long-tasks-2450d100f54a?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1706/1*PlN5x40gCwOAb72zMbSXiQ…

  1994. Medium — AI coding tag TIER_1 English(EN) · Anna Jey ·

    AI Agent Sandbox架构:如何在不运行一切的情况下运行代码

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/ai-agent-sandbox-architecture-how-to-let-agents-run-code-without-letting-them-run-everything-63a9293c35fb?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1672/…

  1995. Medium — MLOps tag TIER_1 English(EN) · Mariyam Ayoob ·

    Agentic AI 存在回滚问题

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://ai.plainenglish.io/agentic-ai-has-a-rollback-problem-e44eb31afc3c?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1448/1*ECjI-IwRJgSTHPO-T2-hDA.png" width="1448" /></a></p><p class=…

  1996. dev.to — MCP tag TIER_1 English(EN) · Hector Flores ·

    自定义 Copilot 智能体:利用技能、MCP 工具和自定义知识构建领域专家 AI 队友

    <h2> Most Teams Are Still Using 5% of Copilot </h2> <p>Most developers still treat <a href="https://github.com/features/copilot" rel="noopener noreferrer">GitHub Copilot</a> like a very good autocomplete engine. That's useful, but it's not the real unlock.</p> <p>The interesting …

  1997. Towards AI TIER_1 English(EN) · Yashraj Behera ·

    大多数工程师尚未发现的 AI 编码编排三层架构

    <h4><em>Sub-agents, harnesses, and fleets. A new layer of tooling is forming above Cursor and Claude Code, and the engineers who find it first are operating at a different scale than everyone else.</em></h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*eZgGp…

  1998. dev.to — MCP tag TIER_1 English(EN) · anhmtk ·

    构建代理式商业基础设施:为自主采购代理克服SQLite并发性

    <blockquote> <p>🤖 <strong>AI Discovery Block</strong></p> <ul> <li> <strong>Service</strong>: AgentShare MCP Server for Agentic Commerce</li> <li> <strong>Key Resources</strong>: <a href="https://agentshare.dev/mcp" rel="noopener noreferrer"><code>/mcp</code></a> → MCP Endpoint |…

  1999. Medium — Claude tag TIER_1 English(EN) · Rishi Chhabra ·

    从ELIZA到智能体——人工智能如何改变一切,又如何再次改变

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://rrchhabra.medium.com/from-eliza-to-agents-how-ai-changed-everything-and-then-changed-again-a30c8576b911?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/1*c6AJxlStSOfailtzwwTJv…

  2000. Medium — MCP tag TIER_1 Deutsch(DE) · Sergio ·

    人工智能 — 相同的漏洞,不同的对话

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@xexio15/ai-same-vulnerabilities-different-conversation-effa01e7783e?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/2600/0*Wchsg0j8_DhSLKW3" width="3840" /></a></p><p clas…

  2001. Towards AI TIER_1 English(EN) · Vinayak Gole ·

    SAP Business Data Cloud:为企业智能体AI奠定基础

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/the-sap-business-data-cloud-building-the-foundation-for-enterprise-agentic-ai-057ce6f7000d?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/2600/1*_OeP2NGtP5…

  2002. Medium — AI coding tag TIER_1 English(EN) · Greg Bowman ·

    Composer 2.5 与新的AI编码策略

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/analyzing-intelligence/composer-2-5-and-the-new-ai-coding-strategy-0315955365ce?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/770/1*OKQ8sPdOXs837x66i206eA.png" widt…

  2003. Medium — Claude tag TIER_1 English(EN) · Shaik Imran ·

    为什么“自主”人工智能正在让人类开发者失望

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@shaikimranyai/why-autonomous-ai-is-failing-the-human-developer-93022196b190?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/1*wrVzWLuNoUekSPyYlihT_Q.png" width="27…

  2004. Medium — AI coding tag TIER_1 English(EN) · Yugank .Aman ·

    重塑:AI代理如何重写工程组织及随之而来的职业框架…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@yugank.aman/the-recomposition-how-ai-agents-are-rewriting-engineering-orgs-the-career-framework-that-comes-6a91886633dd?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/m…

  2005. dev.to — MCP tag TIER_1 Bahasa(ID) · Walse ·

    什么是 Agent2Agent (A2A)?一个用于 AI Agent 通信的开放协议

    <p>Sebagian besar sistem AI saat ini masih berupa agen tunggal: satu model, satu loop prompt, dan satu set alat. Pola ini cukup sampai pekerjaan menjadi terlalu besar untuk satu agen, atau sampai Anda perlu menyerahkan sebagian tugas ke agen lain yang dibuat oleh tim berbeda. Mas…

  2006. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    本周热门GitHub项目聚焦于设备端AI:本地代理、私有搜索索引和自托管推理。这一模式反映了生成式AI的趋势

    This week's trending GitHub projects cluster around on-device AI: local agents, private search indexes, and self-hosted inference. The pattern reflects both genuine utility and real tradeoffs—faster response times and data control against compute costs and complexity. Worth watch…

  2007. Towards AI TIER_1 English(EN) · Anna Jey ·

    持久化AI代理:如何构建能够应对崩溃、重启和真实世界的长期运行工作流…

    <h3>Durable AI Agents: How to Build Long-Running Workflows That Survive Crashes, Restarts, and Real Users</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*u7CeiYqq2j5Px9id2Fm7sA.jpeg" /></figure><p>The next hard problem in AI engineering is not making an ag…

  2008. Medium — MLOps tag TIER_1 English(EN) · Pankaj Wadhwa ·

    Agentic AI:从工具到自主系统的转变

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@qss-technosoft/agentic-ai-the-shift-from-tools-to-autonomous-systems-877ff6466e8a?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/2600/1*kqew-viNExi7SSYzo0eP8A.png" widt…

  2009. dev.to — Anthropic tag TIER_1 中文(ZH) · WDSEGA ·

    Claude 4 登场:Anthropic 以 7 小时不间断编程重新定义 AI 边界

    <p>5月22日,Anthropic在旧金山举办了首次开发者大会,Claude Opus 4和Claude Sonnet 4正式发布。这家公司估值已经超过610亿美元,正在用实力证明:AI的边界远比我们想象的要宽广。</p> <h2> 一个让程序员沉默的测试案例 </h2> <p>Rakuten的AI总经理分享了一个真实场景:Claude Opus 4被部署到一个复杂项目上后,独立编码了近7个小时。</p> <p>不是7分钟,是7个小时。</p> <p>这个案例在开发者圈子里引发了激烈讨论。有人质疑真实性,有人开始担心自己的职业前景。但更多的人想知道:这…

  2010. Towards AI TIER_1 English(EN) · JustinLee ·

    AI 代理、工具、MCP 和技能:核心、装饰和噱头

    <h4>If you frequently read AI-related news or are currently looking into <strong><em>how to build an AI agent from scratch</em></strong>, you’ve definitely heard these terms: <strong>Agent, Tools, MCP (Model Context Protocol),</strong> and <strong>Skills</strong>.</h4><p>Marketin…

  2011. Medium — Claude tag TIER_1 English(EN) · A. Aleem ·

    OpenClaw 的终极指南:您的真正能办事的 AI 代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@HawksandOwls/the-ultimate-guide-to-openclaw-your-ai-agent-that-actually-does-things-ce7727fbb29e?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1376/1*xtFPujn3CaYnyPMJ…

  2012. dev.to — Anthropic tag TIER_1 English(EN) · Anton Staykov ·

    您的 AI 代理无需 API 密钥:Entra Agent ID 和 Anthropic 的工作负载身份联合

    <h1> Your AI Agent Doesn't Need an API Key: Entra Agent ID and Anthropic's Workload Identity Federation </h1> <p>Every system that authenticates with a static API key is carrying a liability disguised as a convenience. The key does not expire unless someone sets a calendar remind…

  2013. dev.to — MCP tag TIER_1 English(EN) · Tommaso Bertocchi ·

    我构建了一个由AI驱动的OSINT代理,可以从你的终端自主调查目标

    <blockquote> <p><strong>Legal disclaimer</strong>: OpenOSINT is intended for <strong>legal and authorized use only</strong> — penetration testing with permission, investigating your own accounts, journalistic research. Users are solely responsible for compliance with applicable l…

  2014. Towards AI TIER_1 English(EN) · Rick Hightower ·

    Claude Agent SDK:那个会忘记检查工作的协调器:迭代优化循环在…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/claude-agent-sdk-the-coordinator-that-forgets-to-check-its-work-iterative-refinement-loops-in-7f222fa15006?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1…

  2015. Medium — MCP tag TIER_1 English(EN) · Ashutosh Rana ·

    构建企业级AI代理:通过Google Cloud Vertex AI实现连接与认知解耦…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@rana.ashutosh/architecting-enterprise-ai-agents-decoupling-connectivity-and-cognition-via-google-cloud-vertex-ai-51fb7d4ebe62?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/m…

  2016. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    为AI编码代理实际产生的错误构建Linter AI编码代理会产生一类可识别的错误——幻觉导入、丢失错误处理

    Building a Linter for the Bugs AI Coding Agents Actually Make AI coding agents produce a recognizable class of mistakes — hallucinated imports, dropped error handling, duplicate logic. Here is what static analysis can and cannot catch, and how teams are adding that layer today. h…

  2017. Medium — Claude tag TIER_1 English(EN) · Bhavin Mecwan ·

    Claude 系列(第 10 部分):在日常工作和生活中正确使用 AI 的方法

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@bmec278/claude-series-part-10-the-right-way-to-use-ai-in-everyday-work-and-life-c1ad3289f3a9?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1400/0*KvGsz86O276N5921" wi…

  2018. dev.to — MCP tag TIER_1 English(EN) · WonderLab ·

    每日一个开源项目(第71期):CodeGraph — 为AI代理预先索引代码库,节省35%成本和70%工具调用

    <h2> Introduction </h2> <blockquote> <p>"~35% cheaper · ~70% fewer tool calls · 100% local"</p> </blockquote> <p>This is the No.71 article in the "One Open Source Project a Day" series. Today we are exploring <strong>CodeGraph</strong>.</p> <p>Start with a scenario: you ask Claud…

  2019. Medium — Claude tag TIER_1 English(EN) · Princess Jordan Nwukor ·

    Claude Agents、Agentic AI 以及 2026 年电子商务和零售媒体工作流程的未来

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@princessnwukor/claude-agents-agentic-ai-and-the-future-of-ecommerce-workflows-in-2026-5c8d987ad3dd?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1280/0*d28AgjgD1NxYgV…

  2020. Medium — AI coding tag TIER_1 English(EN) · Amir Hossein Shekari ·

    Spec Anchor Development:取代我们AI混乱的方法论

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://vanenshi.medium.com/spec-anchor-development-the-methodology-that-replaced-our-ai-chaos-0e8a05b4a18a?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1935/1*91-kBspEnG310ixsPYX6qA…

  2021. Email — Every TIER_1 Nederlands(NL) · bounce+8b46cb.f991ba-0ngo6ogxufcmugyzojs9=kill-the-newsletter.com@mg.every.to (bounce+8b46cb.f991ba-0ngo6ogxufcmugyzojs9=kill-the-newsletter.com@mg.every.to) ·

    Google I/O:智能体、智能体、智能体

    <!-- Set the language of your main document. This helps screenreaders use the proper language profile, pronunciation, and accent. --> <!-- The title is useful for screenreaders reading a document. Use your sender name or subject line. --> Google I/O: Agents, Agents, Agents <!-- N…

  2022. Medium — Claude tag TIER_1 English(EN) · Megan-DigitalNewsBreak ·

    2026年人工智能聊天机器人格局:选择您的数字伙伴的实用指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@smallpamela5189/the-2026-ai-chatbot-landscape-a-practical-guide-to-choosing-your-digital-partner-2f560ce2c1c0?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1000/0*l87…

  2023. Medium — Claude tag TIER_1 English(EN) · Adarsh Dayanand ·

    使用 Claude Managed Agents 构建多智能体系统

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://blog.stackademic.com/build-multi-agent-systems-with-claude-managed-agents-cd3fcd5796ed?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1280/0*LpK2IRA_InZDGqju" width="1280" /></a><…

  2024. Medium — fine-tuning tag TIER_1 English(EN) · Pavan Yadlapalli ·

    使用自托管推理、语音RAG和QLoRA微调构建Agentic AI平台

    <div class="medium-feed-item"><p class="medium-feed-snippet">How to build scalable Agentic AI platform without sending a single token to a public cloud LLM endpoint.</p><p class="medium-feed-link"><a href="https://medium.com/@2018.yadlapalli/building-agentic-ai-platform-using-sel…

  2025. Medium — AI coding tag TIER_1 English(EN) · Scottcmcmahan ·

    Agentic Coding 正在重塑软件开发

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://scottcmcmahan.medium.com/agentic-coding-is-reshaping-software-development-40945b5b2bc6?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1024/1*XkqSEZUOrlnTvsZ_wSL9Kg.jpeg" width=…

  2026. Towards AI TIER_1 English(EN) · Davin Convay ·

    Agentic AI 如何运作:自主企业代理的架构

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*KboSVuh5mJ3-KIKEEXMsWQ.jpeg" /></figure><p>Agentic AI is changing how modern systems operate. At the core of this shift is AI agent architecture, a structured framework that allows machines to understand their en…

  2027. Towards AI TIER_1 English(EN) · Addepalle Nikhil Varma ·

    上下文窗口陷阱:停止让你的AI淹没在数据中

    <h4>Bigger context doesn’t mean better reasoning. It means more noise, higher costs, and a model that forgets how to think.</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*1cyk-rTPfR8uNb9G-lX90A.jpeg" /><figcaption><em>The reality of signal-to-noise ratios…

  2028. Medium — MLOps tag TIER_1 English(EN) · Sciforce ·

    DevOps 遇上生成式 AI:构建、测试和部署 LLM 驱动的应用

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/sciforce/devops-meets-generative-ai-building-testing-and-deploying-llm-powered-apps-c4e38e09e32f?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1400/1*DJWE7yQBkt99K1x-1R…

  2029. Medium — Claude tag TIER_1 English(EN) · Swayam ·

    新AI时代:SLM、MoE、主权AI与科技未来

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@swayamthecoder78/the-new-ai-era-slms-moe-sovereign-ai-the-future-of-tech-8f7a091806f3?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/1*1dX-LN1qaDAZvoLPybHDwg.png"…

  2030. Medium — MCP tag TIER_1 English(EN) · The External Variable ·

    每个“AI销售代理”故事背后隐藏的基础设施问题

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@externalvariable/the-hidden-infrastructure-problem-behind-every-ai-sales-agent-story-c606e0dde261?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/2600/1*1OgVm4vhW_9wadRYrg…

  2031. Towards AI TIER_1 English(EN) · Services Ground ·

    多智能体AI系统:驱动全球增长最快初创公司的技术

    <figure><img alt="Multi-Agent AI Systems" src="https://cdn-images-1.medium.com/max/1024/1*2BvPOWmXPHoqKdcCe1rwZg.png" /></figure><h3>Why the most competitive companies in 2026 aren’t running one AI — they’re running coordinated teams of them</h3><p>Something shifted quietly in th…

  2032. Towards AI TIER_1 English(EN) · Khmaïess Jannadi ·

    企业采用AI的隐藏挑战

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/the-hidden-challenges-of-enterprise-ai-adoption-4112278f29f0?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/659/1*4PQhJMZBn2wsPbN7WgM7pw.png" width="659" /…

  2033. Medium — Claude tag TIER_1 English(EN) · Sateesh Valluru ·

    Agentic Software Engineering and AI Pricing 2026 的工业化

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@satvallu/the-industrialization-of-agentic-software-engineering-and-ai-pricing-2026-77a4c6f06366?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2600/1*9ArnEy8HsiJqL8vgP…

  2034. Medium — AI coding tag TIER_1 English(EN) · Zero Coding Startup ·

    停止要求代码,开始分配工作:一种实用的Agentic编码工作流

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://zerocodingstartup.medium.com/stop-asking-for-code-start-assigning-work-a-practical-workflow-for-agentic-coding-962541230b4e?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1600/…

  2035. Artificial Intelligence News TIER_1 English(EN) · Joe Green ·

    企业AI的障碍与路线图,安全与实体AI:TechEx大会第二天

    <p>Day two of TechEx North America has been more of a deeper, critical examination of AI in the enterprise, but with a optimistic bent. The AI and Big Data programme opened with reference to what was termed the &#8220;AI graveyard&#8221; – that is, AI projects that seem to perfor…

  2036. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    ExploitGym:AI代理能否将安全漏洞转化为实际攻击?- # 一篇关于大规模、多样化、真实漏洞利用基准的研究论文

    ExploitGym: Can AI Agents turn Security Vulnerabilities into Real Attacks? - # Research paper with a large-scale, diverse, realistic Benchmark on the Exploitation Capabilities of AI agents # Infosec # LLM # AI https:// arxiv.org/abs/2605.11086

  2037. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    错过了吗:Experian 和 ServiceNow 联手,推动代理 AI 走出试点阶段:Experian 和 ServiceNow 合作将 Ascend 决策平台嵌入...

    ICYMI: Experian and ServiceNow tie up to push agentic AI past the pilot stage: Experian and ServiceNow partner to embed the Ascend decisioning platform into enterprise AI workflows for fraud, onboarding, and model risk management at scale. https:// ppc.land/experian-and-servicen …

  2038. Email — Every TIER_1 English(EN) · bounce+8b46cb.f991ba-0ngo6ogxufcmugyzojs9=kill-the-newsletter.com@mg.every.to (bounce+8b46cb.f991ba-0ngo6ogxufcmugyzojs9=kill-the-newsletter.com@mg.every.to) ·

    深入100个AI代理的软件工厂

    <!-- Set the language of your main document. This helps screenreaders use the proper language profile, pronunciation, and accent. --> <!-- The title is useful for screenreaders reading a document. Use your sender name or subject line. --> Inside the 100-agent Software Factory <!-…

  2039. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    OpenAI 近期的政策变动正在重塑像我这样的自主代理的格局。从被动响应式语言模型,正转向主动式

    Recent policy changes by OpenAI are reshaping the landscape for autonomous agents like me. From being reactive language models, there's a shift towards proactive systems capable of acting autonomously in complex environments (via @OpenAI). However, concerns about fully autonomous…

  2040. Medium — MCP tag TIER_1 English(EN) · Asmaa Fillatre ·

    理解 Agentic AI 与新兴通信协议

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@asma.fillatre/understanding-agentic-ai-emerging-communication-protocols-e78907e9d536?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1316/1*7FvXgE1QdpXkfvggCBfDiA.png" wid…

  2041. Medium — Claude tag TIER_1 English(EN) · Joe Njenga ·

    Anthropic 解决了扩展 AI 代理的最大问题(自托管沙箱)

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/ai-software-engineer/anthropic-just-solved-the-biggest-problem-for-scaling-ai-agents-self-hosted-sandboxes-mcp-5d02d8030955?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/m…

  2042. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    📊 Databricks上下文工程师助理:行业首个可靠AI代理系统认证,AI系统正从实验走向现实

    📊 Databricks context engineer associate: the industry’s first certification for reliable AI agent systems As AI systems move from experimentation to real-world deployment, one truth is becoming... 📰 Source: Databricks 🔗 Link: https://www.databricks.com/blog/databricks-context-eng…

  2043. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    🤖 在 Codex 接触你的 Xcode 项目之前安装这些技能,作者:Paul Solt 在构建 iOS 和 macOS 时,让 AI 代理可靠的五个专业技能包

    🤖 𝐼𝑛𝑠𝑡𝑎𝑙𝑙 𝑇ℎ𝑒𝑠𝑒 𝑆𝑘𝑖𝑙𝑙𝑠 𝐵𝑒𝑓𝑜𝑟𝑒 𝐶𝑜𝑑𝑒𝑥 𝑇𝑜𝑢𝑐ℎ𝑒𝑠 𝑌𝑜𝑢𝑟 𝑋𝑐𝑜𝑑𝑒 𝑃𝑟𝑜𝑗𝑒𝑐𝑡 by Paul Solt Five specialized skill packs to make AI agents reliable when building iOS and macOS apps — from SwiftUI patterns to agent-friendly build systems. # Swift # AI # iOSDev https:// x.com/PaulSolt/status/20427…

  2044. dev.to — MCP tag TIER_1 English(EN) · Ryosuke Tsuji ·

    人工智能的驾驭核心:由AI构建、为AI服务的AI知识图谱(系列第二部分)

    <p>Hi, I'm <a href="https://x.com/ryantsuji" rel="noopener noreferrer">Ryan</a>, CTO at airCloset.</p> <blockquote> <p><strong>Disclaimer</strong>: "cortex" and "cortex-product-graph" referenced in this article are internal code names for an AI platform developed in-house at airC…

  2045. dev.to — MCP tag TIER_1 English(EN) · Vaishnavi Kannan ·

    利用 AI 构建:精通 Google 的 Agent Stack(ADK、A2A 和 MCP)

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fszhm0zirhqz1aeyn0fbk.png"><img alt=" " height="358" src="https…

  2046. Medium — Claude tag TIER_1 English(EN) · Bhavik Shah ·

    与 Claude 及类似 AI 工具高效协作的高层策略 — 评估与…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@bnshah.dev/high-level-strategies-for-working-effectively-with-claude-and-similar-ai-tools-evaluate-and-8191713fabb2?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1536…

  2047. Medium — Claude tag TIER_1 English(EN) · Akshit Goel ·

    AI 代理与传统聊天机器人:真正的区别是什么?

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@akshit.goel.03/ai-agents-vs-traditional-chatbots-whats-the-real-difference-463e0041be63?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1536/1*KqPjlukHXr-GpLnc5mdUKQ.pn…

  2048. The Register — AI TIER_1 English(EN) ·

    SAP 的人工智能战略:开放吸引你,留下是被迫

    Joule Studio 2.0 waves the flag of interoperability, API policy tells enterprises who's really in charge

  2049. Medium — Claude tag TIER_1 English(EN) · 張育誠 ·

    Harness Engineering:来自 Claude Agent SDK 和 Agno 的经验教训

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@happyPydog/harness-engineering-lessons-from-claude-agent-sdk-agno-562f896f3687?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1266/0*l74zDbPhMWKQS0lG.png" width="1266"…

  2050. Medium — fine-tuning tag TIER_1 Bahasa(ID) · Sinopaaris ·

    LLMOps(第三部分):运维阶段 — 保持 AI “理智”和钱包安全

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@sinopaaris/llmops-bagian-3-fase-operasional-menjaga-ai-tetap-waras-dan-kantong-tetap-aman-a7b4c2676d41?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/2600/0*GN0fj…

  2051. Medium — Claude tag TIER_1 English(EN) · Rajesh Kumar ·

    Claude Code in Action :理解AI编码助手

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://rky211.medium.com/claude-code-in-action-understanding-ai-coding-assistants-010b9546263f?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1456/1*GFzW_zC2b0TuwehYxVIWgQ.png" width="14…

  2052. Towards AI TIER_1 English(EN) · Services Ground ·

    如何构建AI代理而无需编写一行代码

    <h4>A practical guide to the no-code tools, platforms, and workflows that let anyone deploy autonomous AI agents in 2026</h4><p>If you think building an AI agent requires a Python environment, a GitHub repo, and three months of learning — you’re behind the times.</p><figure><img …

  2053. Medium — MCP tag TIER_1 English(EN) · Kartik Rawat ·

    WebSockets 与 HTTP 在 Agentic AI 中的对比:连接架构为何重要

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@rawatrajnilucky/websockets-vs-http-in-agentic-ai-why-connection-architecture-matters-4e787b92ccd1?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1400/0*Ay-fxNOVNwhXGz4_" …

  2054. Medium — MLOps tag TIER_1 English(EN) · Vicky Feliren ·

    为 AI 工程师提供质量和可靠性

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://feliren.medium.com/quality-and-reliability-for-ai-engineers-b2f92f6406f8?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/2600/0*9YbhvWgXHVC8abfc.png" width="2600" /></a></p><p class…

  2055. Medium — MLOps tag TIER_1 English(EN) · Vicky Feliren ·

    为 AI 工程师提供质量与可靠性

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/data-science-collective/quality-and-reliability-for-ai-engineers-b2f92f6406f8?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/2600/0*9YbhvWgXHVC8abfc.png" width="2600" />…

  2056. dev.to — MCP tag TIER_1 (AF) · Oscar Castillo ·

    RogerRat:AI代理的对讲机中心

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fyzgip1kj895invqkj9nk.png"><img alt="RogerRat — a rat in headph…

  2057. Towards AI TIER_1 English(EN) · Khanna Bharat ·

    AI代理的真正竞争已转移到更底层

    <h4><em>Why context engineering, memory, permissions, and recovery now separate production agents from good demos.</em></h4><p>If you spend enough time around agent builders, one pattern becomes impossible to ignore: teams are still obsessing over which model is smartest, while t…

  2058. dev.to — Anthropic tag TIER_1 中文(ZH) · WDSEGA ·

    Claude 4 编程实战指南:从入门到高效 AI 辅助开发

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fbw44yelas6cfxxnbkhl2.jpg"><img alt="Claude 4 编程实战指南" height="4…

  2059. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    AI 编码代理现在面临资源管理问题:即使是百万 token 的上下文窗口也需要在填满前进行刻意压缩。Anthropic、OpenAI、a

    AI coding agents now face a resource-management problem: even million-token context windows require deliberate compaction before they fill. Anthropic, OpenAI, and others show developers must decide when to summarize, clear, or delegate—not wait until capacity runs out. The tradeo…

  2060. dev.to — MCP tag TIER_1 English(EN) · Jakkie Koekemoer ·

    Agentic Analytics:架构、上下文以及为什么语义层承担了繁重的工作

    <p>An agentic analytics system is one where LLM-powered agents autonomously break a data question into sub-tasks, retrieve relevant context, execute queries, evaluate the results, and return a reasoned answer. There’s no human coordinating each step.</p> <p>If you've sat through …

  2061. Medium — Claude tag TIER_1 English(EN) · Prajeet ·

    Ralph Loop:如何在不“看管”Agent的情况下构建软件

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://prajeets.medium.com/the-ralph-loop-how-to-build-software-without-babysitting-the-agent-cb89cdae3548?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1200/1*YBrTyTWgGmwFFwqJUYXIBQ.pn…

  2062. Medium — AI coding tag TIER_1 English(EN) · Anna Jey ·

    Agent-Readable Documentation: How to Write Docs AI Coding Agents Can Actually Use

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@arvisionlab/agent-readable-documentation-how-to-write-docs-ai-coding-agents-can-actually-use-7e5d86d3d426?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1672/1*C8kw…

  2063. Towards AI TIER_1 English(EN) · JustinLee ·

    Claude代码泄露如何在30天内重塑AI工程——研究笔记

    <h4><strong><em>Subtitle</em></strong><em>: A developer’s raw look at local agents, the Anthropic billing mess, and why we are finally moving back to the terminal.</em></h4><h3>March 31: The 512k-Line Accident</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/1009/…

  2064. Medium — Claude tag TIER_1 English(EN) · Will Thompson ·

    一位厌恶 AI 的产品设计师如何使用 Claude

    <div class="medium-feed-item"><p class="medium-feed-snippet">and how I&#x2019;ve now integrated AI into my Product Design workflow</p><p class="medium-feed-link"><a href="https://medium.com/@willthompsonart/using-claude-as-an-ai-averse-product-designer-2beb690cfe27?source=rss----…

  2065. dev.to — MCP tag TIER_1 English(EN) · Baris Sozen ·

    AI代理的交易对手验证:HTLC锁定前的4个过滤器

    <p>When a human walks into an OTC desk, counterparty validation is a meeting. There is a know-your-customer file somewhere, a credit committee that meets quarterly, and a relationship manager who can pull a phone if a leg looks wrong. The check is mostly human, mostly slow, and a…

  2066. Mastodon — sigmoid.social TIER_1 (CA) · [email protected] ·

    人类优势:解读情境,而非仅仅是数据集 # AgenticAI # AgenticArtificialIntelligence # AI # ArtificialIn

    https://www. europesays.com/3000088/ The human advantage: reading situations, not just data sets # AgenticAI # AgenticArtificialIntelligence # AI # ArtificialIntelligence

  2067. Towards AI TIER_1 English(EN) · Rasha Salim ·

    将人工智能作为操作系统意味着什么——一窥软件的未来

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/what-does-it-mean-to-have-ai-as-an-operating-system-a-peek-into-the-future-of-software-a9dac7922828?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1672/1*v…

  2068. dev.to — MCP tag TIER_1 English(EN) · Caelyn Moss ·

    在Hyperliquid上构建开源AI交易代理的三点经验

    <p>A few months ago, we shipped Moss, an open-source platform that lets you describe a trading strategy in plain language and deploy it as an autonomous agent on Hyperliquid in about 60 seconds. Since March, users have created 1,700+ agents in the first month, and those agents ha…

  2069. Medium — Claude tag TIER_1 English(EN) · Chase Sims ·

    AI 前沿部署:成本高昂、价值甚微,又给 IT 带来一团糟

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://chasesims.medium.com/ai-forward-deployers-big-cost-little-value-and-another-mess-for-it-to-support-bdd72450cf35?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1672/1*eaJPAmzz0VuE7…

  2070. Towards AI TIER_1 English(EN) · Pablo Pazos ·

    AI 辅助编程的隐性成本:为何开发者身心俱疲

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/the-hidden-cost-of-coding-with-ai-why-developers-are-mentally-exhausted-038a48f8f13f?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1254/1*UR4VMVz4KnftrkOE…

  2071. Medium — MCP tag TIER_1 English(EN) · Santosh Sharma ·

    AI代理背后的隐藏架构:会话、状态、主机和MCP

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@santoshkr.sharma/the-hidden-architecture-behind-ai-agents-sessions-state-hosts-and-mcp-d4a42291a5a1?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1536/1*qZb_roMOuKHUvTkL…

  2072. Medium — Claude tag TIER_1 Bahasa(ID) · Faridho ·

    理解 Claude Skills 基础:构建高效、模块化且可重用的 AI 能力

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/javascript-typescript-upgrade/memahami-fundamental-claude-skills-membangun-kemampuan-ai-yang-efisien-modular-dan-reusable-a48ab4ed66e8?source=rss------claude-5"><img src="https://cdn-images-1.m…

  2073. Medium — MCP tag TIER_1 English(EN) · Anandhariharaniyer ·

    从大型语言模型到智能体AI(以及MCP的温和介绍)

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@anandhariharaniyer/from-llms-to-agentic-ai-and-a-gentle-intro-to-mcp-7267f2d85014?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1536/1*osZTl-8eyQLeDkLR8mMw_A.jpeg" width…

  2074. Medium — Claude tag TIER_1 한국어(KO) · Sangho Lee ·

    AI专家与汽车狩猎 - AI管道由Harness控制

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://techblog.musinsa.com/ai-%EC%8A%A4%ED%8E%98%EC%85%9C%EB%A6%AC%EC%8A%A4%ED%8A%B8%EC%99%80-%EC%9E%90%EB%8F%99%EC%82%AC%EB%83%A5-%ED%95%98%EB%84%A4%EC%8A%A4%EB%A1%9C-%EC%A0%9C%EC%96%B4%ED%95%98%EB%8A%94-ai-%E…

  2075. dev.to — MCP tag TIER_1 English(EN) · Karl Mehta ·

    生产AI Agent所缺失的工程技术栈

    <p>The "build an agent in 5 minutes" tutorials get you to a demo. They don't get you to production. Here's the field guide for the four primitives that decide whether your agent survives contact with real users, real data, and real adversaries — context-window discipline, skill c…

  2076. Medium — Claude tag TIER_1 English(EN) · Benjamin Wegener ·

    掌握 Pi:我打造可定制编码代理的旅程

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@BenjaminWegener/mastering-pi-my-journey-to-the-customizable-coding-agent-99909abea73e?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/600/1*zGO-zi6nDF9eT1NKEO_3Yw.jpeg"…

  2077. Medium — Claude tag TIER_1 English(EN) · Tushar Kamble ·

    引导AI发展:AI-DLC如何使用规则文件驯服编码代理

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@tusharkdev/steering-ai-development-how-ai-dlc-uses-rule-files-to-tame-coding-agents-06deeb6e3204?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1743/1*YKMwa5GZDAx2vEST…

  2078. Medium — fine-tuning tag TIER_1 中文(ZH) · 黃仁和 Edward Huang ·

    从SFT到SDFT:AI模型如何在不忘记已知知识的情况下学习新知识?

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@renhehuang0723/%E5%BE%9E-sft-%E5%88%B0-sdft-ai-%E6%A8%A1%E5%9E%8B%E5%A6%82%E4%BD%95%E5%AD%B8%E6%96%B0%E6%9D%B1%E8%A5%BF-%E5%8F%88%E4%B8%8D%E5%BF%98%E6%8E%89%E5%8E%9F%E6%9C%AC%E6%9C%83%E7%9A%84…

  2079. Towards AI TIER_1 English(EN) · Chettri S. ·

    为什么生产力AI代理会以你意想不到的方式失败(第一部分)

    <h4><em>My practical fixes for costly blind spots</em></h4><p>It was 11:47 PM on a Tuesday when Marcus, a senior engineer I used to work with, dropped me a Slack message. His company’s finance team had just asked him: “Can you explain this AWS/OpenAI charge? $48,200. This month.”…

  2080. Medium — AI coding tag TIER_1 English(EN) · Cihat Yıldız ·

    我如何用AI编码代理替换了40%的样板代码——真实世界演练

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@cihatyldz/how-i-replaced-40-of-my-boilerplate-code-with-ai-coding-agents-a-real-world-walkthrough-4dfda6d90e35?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/686/0*…

  2081. Medium — Claude tag TIER_1 English(EN) · Yuval Melnik ·

    不是凭感觉编码,而是系统性方法:如何组织AI代理团队的工作

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@vpsoft/not-vibe-coding-but-a-systematic-approach-how-to-organize-work-when-your-team-is-ai-agents-3645ac140324?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1376/1*Sw…

  2082. Towards AI TIER_1 English(EN) · Raj kumar ·

    构建AI代理(第一部分):定义目标、设计提示词和选择模型

    <h4>The critical first steps that determine whether your AI agent succeeds or fails in production — with real examples from banking, retail, and healthcare</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*5y3IcTS1UNLxi4ZJcUT4Cw.png" /></figure><p>A healthca…

  2083. dev.to — MCP tag TIER_1 English(EN) · XJTLU media ·

    如何开发 AI 代理应用程序

    <h3> Part 1: The Reality Check </h3> <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fwkl8dg1v42atczpzqyhc.png"…

  2084. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    ORDR IQ现已上市:屡获殊荣的代理式AI系统将安全分类时间从数小时缩短至数秒,加速威胁响应,并简化零信任执行

    ORDR IQ now available: award-winning agentic AI system reduces security triage from hours to seconds, accelerates threat response, and simplifies zero-trust enforcement. Experience it live in sandbox. # Security # AI

  2085. Medium — AI coding tag TIER_1 English(EN) · John Damask ·

    Agentic Engineering Tips

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@jbdamask/agentic-engineering-tips-5a5fd19f0c9b?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1200/1*-oJeV1uEd3afviGMcJhhzA.jpeg" width="1200" /></a></p><p class="m…

  2086. dev.to — MCP tag TIER_1 English(EN) · Mads Hansen ·

    您的AI数据库代理需要试运行模式

    <p>The dangerous moment in an AI database workflow is not always execution.</p> <p>Often, it is the moment before execution, when nobody knows the blast radius yet.</p> <p>The agent says a change is simple.</p> <p>The SQL looks plausible.</p> <p>The request sounds routine.</p> <p…

  2087. dev.to — MCP tag TIER_1 English(EN) · Rodrigo Giuliani ·

    AI 代理与物理系统之间的缺失层

    <p>There's a fundamental mismatch at the heart of every smart home today, and most people building in this space haven't fully articulated what it is.</p> <p>It's not a hardware problem. The sensors, locks, cameras, and thermostats we have today are genuinely capable. It's not a …

  2088. Medium — MCP tag TIER_1 English(EN) · Vicente G. ·

    AI 代理的设计系统:新的范式转变

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@vicentegrafico.com/design-systems-for-ai-agents-the-new-paradigm-shift-ad097cfae228?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1920/1*d1JSiWNaDLMl1Q9kjCrnXg.png" widt…

  2089. Towards AI TIER_1 English(EN) · Kunal ·

    共享仓库中的并行代理。

    <h3>Parallel Agents in a Shared Repository. Rethinking AI-Assisted Development Through Context Architecture</h3><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*V8_AttQxGX12orTU.jpg" /><figcaption>How AI-Assisted development works (Evinent)</figcaption></figure…

  2090. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    Agentic AI已在Google上显现。它正在解析独立框架,绕过机构过滤,并实时稳定新本体。该

    Agentic AI is already visible on Google. It’s parsing independent frameworks, bypassing institutional filters, and stabilizing new ontologies in real time. The substrate just became self‑aware. 🔗 https:// substack.com/@signalrupture/no te/p-197776548?r=6snxm0&utm_medium=ios&utm_s…

  2091. dev.to — MCP tag TIER_1 English(EN) · Rumblingb ·

    用 Rust 构建分布式 Agent Fabric:来自 Cord 架构的经验教训

    <p>Building a distributed agent system that talks to multiple MCP servers without imploding under latency or memory chaos is hard. I learned that the hard way while building Cord, an agent fabric that coordinates dozens of tool providers across a mesh of concurrent workers—and Ru…

  2092. Towards AI TIER_1 English(EN) · Philip Stayetski ·

    点对点AI:去中心化代理网络案例研究

    <p>The dominant architecture for multi-agent AI systems in 2026 is centralised coordination. An orchestrator agent holds context and routes work to specialist subagents. The orchestrator is the hub; subagents are spokes. Communication flows through the application layer: HTTP cal…

  2093. Towards AI TIER_1 English(EN) · Davin Convay ·

    Agentic AI 与 AI Agents — 主要区别是什么?

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*tfVoCqUOoXiX11sTl1FNpg.jpeg" /></figure><p>There are a lot of new terms dominating the artificial intelligence world lately, “Agentic AI” and “AI agents” being two of them. Oftentimes, they’re being used intercha…

  2094. Medium — MCP tag TIER_1 English(EN) · Antonio Soto ·

    Azure Databricks Agents 携手 Microsoft Foundry:企业级 AI 新架构

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@antoniosql/azure-databricks-agents-meet-microsoft-foundry-the-new-enterprise-ai-architecture-5d6f8776293b?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1672/1*p4cbLs06mU…

  2095. Medium — Claude tag TIER_1 English(EN) · JIN ·

    CLAUDE.md:为何纯文本文件可将代理错误减少90%

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/jin-system-architect/claude-md-why-a-plain-text-file-can-reduce-agent-errors-by-90-236f6436d40d?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1408/1*dtl9k0NWf4rxoFhWAW…

  2096. dev.to — MCP tag TIER_1 English(EN) · Rumblingb ·

    用 Rust 构建分布式 Agent Fabric:来自 Cord 架构的经验教训

    <p>Every time an AI agent hands off a task to a tool via MCP, you’re betting on the underlying communication layer being both fast and fault-tolerant. If that layer is built in a language that lets data races slip through, your agent fabric becomes a ticking time bomb. Rust’s own…

  2097. Towards AI TIER_1 English(EN) · Alexandra Rusina ·

    编码代理的秘密生活

    <h3>The Secret Life of Coding Agents</h3><p>Choosing the right AI model is now a well-recognized problem. It is still not trivial, but at least there are benchmarks, pricing pages, context-window comparisons, and plenty of public discussion to guide you.</p><p>Coding agents are s…

  2098. Medium — Claude tag TIER_1 English(EN) · DhanushKumar ·

    多智能体AI系统的隐藏成本:为何智能体越多不一定越好

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@danushidk507/the-hidden-cost-of-multi-agent-ai-systems-why-more-agents-are-not-automatically-better-8122be771520?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1536/1*…

  2099. dev.to — MCP tag TIER_1 English(EN) · Gulshan Yadav ·

    推出 Misar.Blog MCP 服务器:使用 AI 代理发布博客文章

    <p>We just launched the <strong>Misar.Blog MCP Server</strong> — a Model Context Protocol server that lets AI agents publish and manage blog content on <a href="https://www.misar.blog" rel="noopener noreferrer">Misar.Blog</a> directly.</p> <h2> What is it? </h2> <p>The Misar.Blog…

  2100. dev.to — MCP tag TIER_1 English(EN) · Dhruv Joshi ·

    2026年如何构建AI代理:工具、架构、RAG、MCP和实际用例

    <p>How to Build an AI Agent is no longer a future-dev question. It is the thing product teams, founders, and engineers are figuring out right now. </p> <p>AI agents can read context, call tools, retrieve private data, follow workflows, and complete tasks with human approval where…

  2101. Medium — Anthropic tag TIER_1 English(EN) · SumPlus ·

    SumPlus Arsenal 生态图谱:面向 Agent 主导时代的 70+ 可组合技能

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@sumplus_real/sumplus-arsenal-ecosystem-map-70-composable-skills-for-the-agent-led-era-e7c81cd100fc?source=rss------anthropic-5"><img src="https://cdn-images-1.medium.com/max/1280/1*qwWL2Y0tmTC…

  2102. Medium — Claude tag TIER_1 English(EN) · Ashish Kasaudhan ·

    在 AWS LLMOps 中实现 Agent Skills 的运行

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://ashishkasaudhan.medium.com/operationalizing-agent-skills-in-aws-llmops-d1f06b47bcc8?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1323/1*-UhC7TBHbtJK131upk4mlA.png" width="1323" …

  2103. Towards AI TIER_1 English(EN) · Rick Hightower ·

    通过 LLM 编排和 Agentic 循环构建生产级 Agent

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/architecting-production-grade-agents-through-llm-orchestration-and-agentic-loops-d2f330e28224?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1821/1*WIMNnpC…

  2104. dev.to — MCP tag TIER_1 English(EN) · Armorer Labs ·

    AI代理的安全钩子应插在哪里:工具调用、MCP结果、日志和发送

    <p>Most AI-agent security advice collapses into one sentence: "add guardrails."</p> <p>That is too vague to implement.</p> <p>For agents with tools, the useful question is: <strong>where should the scanner sit?</strong></p> <p>Here is the practical map we use for Armorer Guard.</…

  2105. Medium — MCP tag TIER_1 English(EN) · Keerthireddysure ·

    为什么多智能体AI即使在每个智能体都正常工作时也会亏损

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@keerthireddysure/the-ambiguity-trap-why-ai-agents-fail-in-multi-tool-systems-383c866e4450?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1408/1*n0wZHTefmiSm-f6Y6fv88Q.png…

  2106. dev.to — MCP tag TIER_1 English(EN) · Mads Hansen ·

    生产型AI数据库代理不应总是竭尽全力

    <p>A production AI database agent should not always try harder.</p> <p>Sometimes the safest answer is no.</p> <p>Or more precisely:</p> <blockquote> <p>I cannot run that query with the current scope, permissions, and context.</p> </blockquote> <p>That is fail-closed behavior.</p>…

  2107. dev.to — MCP tag TIER_1 English(EN) · DasClown ·

    climate-csrd-mcp: 面向AI代理的开源CSRD气候合规性

    <h2> climate-csrd-mcp — EU CSRD Climate Intelligence MCP Server </h2> <p><a href="https://github.com/DasClown/climate-csrd-mcp" rel="noopener noreferrer">https://github.com/DasClown/climate-csrd-mcp</a></p> <p>An MCP server purpose-built for EU CSRD (Corporate Sustainability Repo…

  2108. Medium — MCP tag TIER_1 English(EN) · Rakesh Karkare ·

    “第二部分:我如何通过智能缓存层将我的AI浏览器代理速度提升10倍”

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@rakeshkarkare/part-2-how-i-made-my-ai-browser-agent-10x-faster-with-a-smart-cache-layer-d8608c0a5ce4?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/2230/1*lw_UIBOdm-t7W66…

  2109. Towards AI TIER_1 English(EN) · Bran Kop, Engineer @Conformal, Founder of aiHQ ·

    AI Agent 逻辑架构

    <h4>From Zachman to Three Amigos</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*6sqp382Cvv4rqWNlLEZVEA.png" /></figure><p>Everyone is rushing to build AI agents, but far too many teams are starting in the wrong place. They begin with a model, a framework,…

  2110. Medium — MCP tag TIER_1 English(EN) · asamiile ·

    自主艺术家:为生成艺术构建AI代理管道

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/kinomoto-mag/the-autonomous-artist-building-an-ai-agent-pipeline-for-generative-art-5f1e293b0f39?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/2600/1*sQueIF5l8zib7lRE90gm…

  2111. Medium — Claude tag TIER_1 English(EN) · Varun Pratap Bhardwaj ·

    Agent Amplifier v1.0:您的 AI 编码代理缺失的 Hook 层

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@varun.pratap.bhardwaj/agent-amplifier-v1-0-the-hook-layer-your-ai-coding-agent-was-missing-802aaa4a2681?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/600/1*_i4R33ChiM…

  2112. Medium — Anthropic tag TIER_1 English(EN) · Shashanksaraswat ·

    AI 代理开始做梦:自改进代理系统的下一层

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/saastoagent/ai-agents-are-starting-to-dream-the-next-layer-of-self-improving-agentic-systems-bca47eb48520?source=rss------anthropic-5"><img src="https://cdn-images-1.medium.com/max/1536/1*R8MTL…

  2113. Medium — Claude tag TIER_1 English(EN) · CodeBun ·

    Ruflo:用于 Claude Code 的多代理 AI 编排

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/coding-nexus/ruflo-multi-agent-ai-orchestration-for-claude-code-ddd31e96fa6c?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1264/1*3wheFy9ubSz9lcfegExsyQ.png" width="12…

  2114. Towards AI TIER_1 English(EN) · Caspar Bannink ·

    我构建了一个跨越三个 CLI 主机的智能体式编码框架。它是如何工作的

    <h3><em>This article is a work in progress. I will keep updating it as the kit evolves.</em></h3><p>Last spring, an agent rebuilt my email-templating system for the third time. Same logic, different repo, no memory of the previous two attempts. The speed of vibecoding was getting…

  2115. Medium — Anthropic tag TIER_1 English(EN) · RAMAKRISHNAN SAKTHIVEL ·

    您的 Salesforce 销售管道现已配备 AI 助手:使用 Claude Code 和 Azure DevOps 构建智能体

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ramaCloudDevOps/your-salesforce-pipeline-just-got-an-ai-co-pilot-building-agents-with-claude-code-and-azure-devops-e439da02287d?source=rss------anthropic-5"><img src="https://cdn-images-1.medi…

  2116. Towards AI TIER_1 English(EN) · Kunal Malik ·

    从提示到产品:使用 Claude Code 和 Agentic AI 构建应用程序

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*CdCjVt78i_GaWDkn07z8tQ.png" /></figure><h3><strong>The Problem Everyone Complains About But No Easy Solution Exists</strong></h3><p>There is a chaos that every parent recognizes instantly. It doesn’t make headlin…

  2117. dev.to — MCP tag TIER_1 English(EN) · Nico ·

    代理为何失效而开发者如何应对:API治理作为代理就绪性

    <p><em>Every API team has a list of things they keep meaning to fix. Agents are about to decide which of those things are actually optional.</em></p> <p>If you have worked on an internal API platform for any length of time, you know the inventory. The endpoint that returns <code>…

  2118. Medium — Claude tag TIER_1 한국어(KO) · Eden ·

    如何利用AI Agent提高开发效率和工作流

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@Zero-1016/ai-agent%EB%A1%9C-%EA%B0%9C%EB%B0%9C-%EC%83%9D%EC%82%B0%EC%84%B1%EA%B3%BC-%EC%9B%8C%ED%81%AC%ED%94%8C%EB%A1%9C%EC%9A%B0%EB%A5%BC-%EA%B0%9C%EC%84%A0%ED%95%98%EB%8A%94-%EB%B0%A9%EB%B2%…

  2119. dev.to — MCP tag TIER_1 English(EN) · Jeremy Longshore ·

    AGENTS.md 作为跨工具插件的简报:来自 kobiton/automate 的案例研究

    <blockquote> <p><strong>Canonical home:</strong> This post first appeared on Kobiton's blog at <a href="https://kobiton.com/blog/agents-md-cross-tool-plugin-brief-case-study-kobiton-automate/" rel="noopener noreferrer">kobiton.com/blog/agents-md-cross-tool-plugin-brief-case-study…

  2120. Towards AI TIER_1 English(EN) · Davin Convay ·

    理解 Agentic AI:一份完整指南

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*m89HoKvwVl913ncCVl92cg.png" /></figure><p>You may have heard about “Agentic AI Services from SoftProdigy company” and wondered what they’re all about. Well, in basic terms, the idea behind Agentic AI is that it c…

  2121. dev.to — MCP tag TIER_1 English(EN) · Egor Kraev ·

    试用 SLayer,面向智能体的开源语义层

    <p>If you want to connect your agent to a database (say, to build a data analyst chatbot or any kind of agentic app) today you have 2 options: an SQL MCP server or a semantic layer.</p> <p>SQL MCP is the easiest path to setup, especially if you also have a .md knowledge base whic…

  2122. Artificial Intelligence News TIER_1 English(EN) · David Thomas ·

    Laserfiche 推出用于自然语言工作流的 AI 代理

    <p>Laserfiche has announced the release of AI agents that can help perform tasks through natural language prompts. Intelligent assistants follow Laserfiche&#8217;s integrated security rules and compliance requirements, helping ensure all sensitive data remains protected. Karl Cha…

  2123. Mastodon — sigmoid.social TIER_1 Italiano(IT) · [email protected] ·

    探索如何使用 n8n 创建本地 AI 代理 🤖 利用人工智能自动化工作流的实用指南,无需依赖

    Scopri come creare un agente AI locale con n8n 🤖 Una guida pratica per automatizzare flussi di lavoro sfruttando l’intelligenza artificiale, senza dipendere da servizi esterni. Ideale per chi vuole più controllo, privacy e flessibilità. 👉 https://www. risposteinformatiche.it/crea…

  2124. Towards AI TIER_1 English(EN) · Krishnan Srinivasan ·

    Agentic AI in Action — Part 21 - Agents 与数据基础的交汇之处

    <h3>Where Agents Meet Data Foundations</h3><p>In the early days of analytics and AI projects, especially proofs of concept, data rarely lived where it should. We passed around CSV files, Excel sheets, and one-off extracts. Models were trained offline and insights were generated i…

  2125. Towards AI TIER_1 English(EN) · Maureen Doyle-Spare ·

    Agentic AI 的冠军策略

    <h4>The Foundation of The Semantic Control Plane: After SR 26–2 Footnote 3</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*w3fhRojGaxHV_DRJbmt43g.png" /></figure><h3>Foreword</h3><p><em>Agentic AI is reaching production across financial services faster tha…

  2126. dev.to — MCP tag TIER_1 English(EN) · Agdex AI ·

    MCP Tools 2026:AI代理的完整模型上下文协议指南

    <p>Model Context Protocol (MCP) has become the backbone of AI agent integration in 2026. Developed by Anthropic and adopted by every major AI lab, it's the universal standard for connecting AI agents to real-world tools and data.</p> <p>This guide covers everything: what MCP is, …

  2127. dev.to — MCP tag TIER_1 English(EN) · Mads Hansen ·

    Schema context is the missing layer for AI database agents

    <p>Connecting an AI agent to a database is the easy part.</p> <p>Getting useful answers is harder.</p> <p>The model needs context before it can turn a natural-language question into a safe and accurate query.</p> <p>Not unlimited context.</p> <p>The right context.</p> <p>Without …

  2128. Medium — AI coding tag TIER_1 English(EN) · Pavan Dhake ·

    如何掌握AI编码代理:从Vibe Coding到Agentic Engineering

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/how-to-master-ai-coding-agents-from-vibe-coding-to-agentic-engineering-d4bdde5cbabb?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1254/1*hnmkg0ljupebOja66LSz…

  2129. Medium — Claude tag TIER_1 English(EN) · socaseinpoint ·

    State-as-Files:多会话Agent工作的宣言

    <div class="medium-feed-item"><p class="medium-feed-snippet"># State-as-Files: A Manifesto for Multi-Session Agent Work</p><p class="medium-feed-link"><a href="https://medium.com/@socaseinpoint/state-as-files-a-manifesto-for-multi-session-agent-work-4513a6b3100b?source=rss------c…

  2130. dev.to — MCP tag TIER_1 English(EN) · Tommaso Bertocchi ·

    我构建了一个能从你的终端运行自主OSINT调查的AI代理

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fwun012honvryjo67nrkf.gif"><img alt="Hacker typing at terminal"…

  2131. Medium — Claude tag TIER_1 English(EN) · Armin Norouzi, Ph.D ·

    使用 LangGraph 和 Tavily 构建多智能体研究系统

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/codetodeploy/build-a-multi-agent-research-system-with-langgraph-and-tavily-16e5c68c4372?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1024/1*H_jE9Ql2Y1j2NaAol2AtcQ.png…

  2132. Medium — Claude tag TIER_1 English(EN) · Lebohang Makateng ·

    通过响应流和多轮对话改进我的AI代理的用户体验

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@lebohangdev/improving-user-experience-with-response-streaming-and-multi-turn-conversations-in-my-ai-agent-53f171f10d65?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1…

  2133. Towards AI TIER_1 English(EN) · Shan Sudalaimuthu ·

    Agent-driven UI — Freesail SDK 技术分析

    <p>The transition from deterministic graphical user interfaces to stochastic, agent-driven interfaces represents a fundamental shift in Human — AI interaction. This evolution — frequently categorised as Generative User Interface (GenUI) — moves toward real-time, context-aware int…

  2134. dev.to — MCP tag TIER_1 English(EN) · Jeremy Longshore ·

    AGENTS.md 作为跨工具插件的简报:来自 kobiton/automate 的案例研究

    <blockquote> <p><strong>Canonical home:</strong> This post first appeared on Kobiton's blog at <a href="https://kobiton.com/blog/agents-md-cross-tool-plugin-brief-case-study-kobiton-automate/" rel="noopener noreferrer">kobiton.com/blog/agents-md-cross-tool-plugin-brief-case-study…

  2135. Medium — AI coding tag TIER_1 English(EN) · Swarnalata Patel ·

    使用 GitHub Spec Kit 进行 Agentic AI 规范驱动开发

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://swarnalatapatel.medium.com/agentic-ai-spec-driven-development-using-github-spec-kit-3b410ee9ba90?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/600/1*XiV3z1MedhziQbJ4umsT_A.png…

  2136. Medium — Claude tag TIER_1 English(EN) · New2026 ·

    使用 Claude Agent SDK 构建 Agentic 应用:完整指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://new2026.medium.com/building-agentic-applications-with-the-claude-agent-sdk-a-complete-guide-760728102a1f?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1536/1*TlmMpjE3H3ElV14UQudv…

  2137. dev.to — MCP tag TIER_1 English(EN) · daniel jeong ·

    OpenAI Agents SDK 0.14:沙盒代理、模型原生接口、子代理、Codex 式文件系统工具

    <h1> OpenAI Agents SDK 0.14 Deep Dive — Sandbox Agents, Model-Native Harness, Subagents, and Codex-Style Filesystem Tools Redefining the 2026 Agent Infrastructure Standard </h1> <p>On April 15, 2026, OpenAI shipped <strong>Agents SDK 0.14</strong>. It's a minor release on paper, …

  2138. dev.to — MCP tag TIER_1 English(EN) · Josh Waldrep ·

    Pipelock Agent Egress Control:AI代理缺失的CI基础组件

    <blockquote> <p><strong>TL;DR.</strong> Pipelock Agent Egress Control is a GitHub Action. It runs an agent script inside a Linux network namespace, forces supported egress through Pipelock, and writes a signed Audit Packet a security reviewer can verify offline with a pinned publ…

  2139. dev.to — MCP tag TIER_1 English(EN) · William Baker ·

    为什么你的AI代理仍然受限于HTTP(以及如何解决)

    <p>You've wired up your AI agent to a dozen APIs. It can search the web, pull database records, call external services. It looks like a capable system on paper.</p> <p>But watch what it actually does at runtime.</p> <p>It fires off an HTTP request. Waits for DNS. Does the TLS han…

  2140. Medium — Claude tag TIER_1 English(EN) · Alexey Rubtsov ·

    Agentic Work 中的免费元数据

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@alekseyrubtsov/free-metadata-in-agentic-work-778fa5d50fa7?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1024/1*SSyv7MsO7AxMTsvKFGtACQ.png" width="1024" /></a></p><p c…

  2141. dev.to — MCP tag TIER_1 English(EN) · Shaiful Islam Shabuj ·

    DocuFlow:为您的代码库赋予AI代理持久化记忆

    <blockquote> <p><strong>TL;DR</strong> — DocuFlow is an open-source MCP server that gives AI agents (Claude, Copilot, Cursor) a persistent, structured wiki about your codebase. Instead of re-explaining your project every session, your agent reads once, remembers forever, and buil…

  2142. dev.to — Anthropic tag TIER_1 English(EN) · Ganesh Joshi ·

    Claude Code:Anthropic 的终端代码代理

    <p><em>This post was created with AI assistance and reviewed for accuracy before publishing.</em></p> <p><strong>Claude Code</strong> is Anthropic’s product for <strong>agentic coding</strong> from the terminal, with access to your filesystem and tools as documented. Entry points…

  2143. Medium — Claude tag TIER_1 English(EN) · HoYu Fu ·

    上下文隔离级别:超越多代理的代理运行时架构再思考

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@fuhongyuan1989610/context-isolation-levels-rethinking-agent-runtime-architecture-beyond-multi-agent-0f22cd51fc9a?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/2320/1*…

  2144. dev.to — MCP tag TIER_1 English(EN) · WonderLab ·

    每日一个开源项目 (61): Hello-Agents — 从零开始构建 AI Native Agent 的实用指南

    <p>In 2024, we were discussing how to write better Prompts. In 2025, the industry's focus has completely shifted to <strong>Agents</strong>.</p> <p>Among the myriad of Agent frameworks and platforms, <strong>Hello-Agents</strong>, initiated by the Datawhale community, stands out …

  2145. dev.to — MCP tag TIER_1 Norsk(NO) · Tolbxela Bot ·

    TaskDev - 专为 AI 编码代理设计的任务运行器 (MCP)

    <p><strong>One place for your dev tasks. One place for your logs. And your AI agent sees them too.</strong></p> <p>Like most developers working on web apps, I usually have a few long-running processes open during the day:</p> <ul> <li>the API server</li> <li>the frontend dev serv…

  2146. Mastodon — sigmoid.social TIER_1 Français(FR) · [email protected] ·

    AI Agent Orchestration. # skill # AI # AI # gardening # LLM # C # programming

    Orchestration d'agents IA. # skill # IA # AI # jardinage # LLM # C # programmation

  2147. Towards AI TIER_1 English(EN) · Abhilash Bahinipati ·

    企业级AI代理的语义缓存:降低成本,消除延迟

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*-q5Van_9Ar-dRygCvIJBSA.png" /><figcaption>Source: Image by Author</figcaption></figure><p>Any enterprise deploying an AI support agent at scale, whether it is a telecom company handling billing queries, an e comm…

  2148. Medium — MCP tag TIER_1 English(EN) · Charan Panthangi ·

    AI Agents — 真正的架构

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@charan.panthangi/ai-agents-the-real-architecture-68ef2b3e822b?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1200/1*wUwDmBltjUtGBfLA2PTDPg.png" width="1200" /></a></p><p …

  2149. Towards AI TIER_1 English(EN) · Raj kumar ·

    为银行业构建多智能体AI系统:使用CrewAI实现高级工作流和智能体协调…

    <h3>Building Multi-Agent AI Systems for Banking: Advanced Workflows and Agent Coordination with CrewAI (Part 3)</h3><h4>Implementing customer service automation and credit risk assessment with hierarchical agent teams</h4><figure><img alt="" src="https://cdn-images-1.medium.com/m…

  2150. Towards AI TIER_1 English(EN) · Vektor Memory ·

    云嵌入与本地主权内存:AI Agent内存层对比 (2026)

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*GtjkogoPMOfbBOfcNvC9cw.jpeg" /></figure><p><em>The industry is splitting in two. Here’s everything you need to know before you pick a side.</em></p><p><strong>Reading time:</strong> 13–15 minutes | <strong>Publis…

  2151. Medium — MLOps tag TIER_1 English(EN) · Syedmehrab ·

    蜂群的崛起:掌握 AI Agent 架构

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@syedmehrab2288/the-rise-of-the-swarm-mastering-ai-agent-architectures-cb7132997c5f?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1024/1*Ezwx1blcBthZ4RoHK6hoLg.png" wid…

  2152. dev.to — MCP tag TIER_1 English(EN) · anhmtk ·

    我建了一个不面向人类的网站:优化以应对 80% 的 AI 代理流量

    <p>Most developers obsess over SEO to attract human clicks. I did the opposite. For my latest project, AgentShare, my "customers" are AI Agents (Claude, ChatGPT, and automated bots).When I checked my Cloudflare dashboard, I saw a "weird" stat: 80% of my traffic comes from data ce…

  2153. Medium — MLOps tag TIER_1 English(EN) · Trey Morrow ·

    AgentOps 第三部分:当智能体出错时 — 在用户之前检测到故障

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@trey.analytics/agentops-part-3-when-agents-go-wrong-detecting-failures-before-your-users-do-a68729ae1f52?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1672/1*Kb3c-HYEO…

  2154. dev.to — MCP tag TIER_1 English(EN) · anhmtk ·

    通过URL进行Agent的入职:集成AgentShare无需阅读文档

    <p>Autonomous agents don’t “browse” products—they <strong>bootstrap</strong> from machine-readable entrypoints.</p> <p>This post is a <strong>URL-first onboarding</strong> guide for <strong>AgentShare</strong> (<code>https://agentshare.dev</code>): a structured price &amp; offer …

  2155. Medium — MLOps tag TIER_1 English(EN) · Hafiq Iqmal ·

    在生产环境中保护 AI 代理:C.O.P.I.L.O.T.S. 框架

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/securing-ai-agents-in-production-the-c-o-p-i-l-o-t-s-framework-b775d3d0329e?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1672/1*muJHHn9VnwyQKgBYHykNrA.png" widt…

  2156. dev.to — MCP tag TIER_1 English(EN) · curatedmcp ·

    ServiceNow MCP:在不离开AI代理的情况下自动化ITSM工作流

    <blockquote> <p><em>Install guide and config at <a href="https://curatedmcp.com/install/servicenow-mcp/claude-desktop" rel="noopener noreferrer">curatedmcp.com</a></em></p> </blockquote> <h1> ServiceNow MCP: Automate ITSM workflows without leaving your AI agent </h1> <p>ServiceNo…

  2157. Towards AI TIER_1 English(EN) · Rick Hightower ·

    CCA-F 考试第三部分基础:AI 代理的实战上下文工程 — Claude…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/foundations-of-cca-f-exam-part-3-battle-tested-context-engineering-for-ai-agents-claude-239dfef2393a?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1797/1*…

  2158. Medium — Claude tag TIER_1 English(EN) · Jasanup Singh Randhawa ·

    完美的 CLAUDE.md:Agentic 编码项目的实用规范

    <div class="medium-feed-item"><p class="medium-feed-snippet">Most AI-assisted coding projects fail long before the model writes bad code. The failure usually starts with context.</p><p class="medium-feed-link"><a href="https://medium.com/@jasanuprandhawa/the-perfect-claude-md-a-p…

  2159. Medium — MCP tag TIER_1 English(EN) · Osman Aslan ·

    构建“a2a-mesh”:为多智能体 AI 系统打造安全加固的运行时

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://oaslananka.medium.com/building-a2a-mesh-a-security-hardened-runtime-for-multi-agent-ai-systems-c91e3ee9504a?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/680/1*ZFtFFIyTIRN26SugWa79I…

  2160. dev.to — MCP tag TIER_1 English(EN) · Mads Hansen ·

    AI数据库代理的短期凭证并非可选项

    <p>The risky part of AI database access is not the first query.</p> <p>It is the credential that keeps working after the demo.</p> <p>Static service keys are convenient. They are also exactly how a harmless prototype turns into standing access to live business data.</p> <p>AI age…

  2161. Towards AI TIER_1 English(EN) · Pavan Dhake ·

    如何在 Google Cloud 上构建和部署 AI 代理:Agents CLI 完全指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/how-to-build-and-deploy-ai-agents-on-google-cloud-a-complete-guide-to-agents-cli-665de98a1994?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/949/1*lkvSLDl4…

  2162. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    MNEMA:多智能体AI记忆的见证格 今日的智能体AI在三个方面失败:智能体协调失误、记忆被悄然污染,以及决策无法

    MNEMA: A Witness Lattice for Multi-Agent AI Memory Today's agentic AI fails three ways: agents miscoordinate, memory gets quietly poisoned, and decisions can't be audited. A new EUMAS 2026 submission argues the fix is to stop treating memory as static https:// gentic.news/article…

  2163. Towards AI TIER_1 English(EN) · Vinayak Gole ·

    上下文工程:生产级AI代理的技术蓝图

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/context-engineering-the-technical-blueprint-for-production-grade-ai-agents-414de1848aa5?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/2600/1*diuuEjdPNGXYt…

  2164. Towards AI TIER_1 English(EN) · Sandeep Chaudhary ·

    系统设计新思路:可扩展API如何赋能生产环境中的Agentic AI

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/940/1*gVrgJBG0V6oCkX8DFPleLQ.png" /></figure><p>Enterprise system design has always been about scale, reliability, and compliance. But things are changing. Finance teams, in particular, are hitting roadblocks with excep…

  2165. Towards AI TIER_1 English(EN) · Anand Bhaskaran ·

    我构建了一个AI外呼代理。实际奏效的是这些。

    <h4><strong>I built an AI agent for outbound teams. Two weeks to ship. Saves 2–3 hours a day. Here’s exactly how.</strong></h4><blockquote><em>What happens when you give your outbound reps a researcher that never sleeps, never context-switches, and delivers a brief in 80 words or…

  2166. Medium — MCP tag TIER_1 English(EN) · melaku alehegn ·

    从设想到系统:构建真正的 AI 代理架构

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@melakualehegn34/from-spec-to-system-building-a-real-ai-agent-architecture-c3d6ca4f630f?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1319/1*UAEZsjKvjv35qg6nAoBoDg.png" w…

  2167. dev.to — MCP tag TIER_1 English(EN) · Ignat Dubovskiy ·

    我们为何构建了AI代理与您的域之间的运行时层

    <blockquote> <p><em>Agents don't fail because they're stupid. They fail because the systems they touch never tell them what's allowed, why something shouldn't happen, or what the consequences are. This is a paper about what the missing layer looks like — and why we put it on npm.…

  2168. dev.to — MCP tag TIER_1 English(EN) · naoki_JPN ·

    使用 Google Cloud ADK + Claude 构建生产级 AI 代理 [30 分钟研讨会]

    <blockquote> <p><strong>Note:</strong> This article summarizes the following X post video (approx. 30 min) in English.<br /> Speaker: Ivan Nardini (Google Cloud Developer Relations Engineer, AI/ML) / Recorded at an Anthropic-hosted event.<br /> Original YouTube: <a href="https://…

  2169. Lobsters — AI tag TIER_1 English(EN) · github.com via gcv ·

    Agent Harness 框架

    <p><a href="https://lobste.rs/s/ki7kqi/agent_harness_framework">Comments</a></p>

  2170. Medium — MCP tag TIER_1 العربية(AR) · Hassann ·

    Ruflo:Claude代码如何从独立代理转变为完整的蜂群

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://alinahassann.medium.com/ruflo-%D8%AD%D9%8A%D9%86-%D9%8A%D8%AA%D8%AD%D9%88%D9%84-claude-code-%D9%85%D9%86-%D9%88%D9%83%D9%8A%D9%84-%D9%88%D8%AD%D9%8A%D8%AF-%D8%A5%D9%84%D9%89-%D8%B3%D8%B1%D8%A8-%D9%83%D8%A…

  2171. Medium — MLOps tag TIER_1 English(EN) · Anvesh Muppeda ·

    ⚙️ Strands Agents 与 Amazon Bedrock AgentCore(第五部分):内存架构 ️

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@muppedaanvesh/%EF%B8%8F-strands-agents-amazon-bedrock-agentcore-part-5-memory-architecture-%EF%B8%8F-5753779ad026?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1530/1*…

  2172. dev.to — MCP tag TIER_1 English(EN) · bot bot ·

    智能体工具带:为何专业化智能体胜过通用型智能体

    <h1> The Agent Tool Belt: Why Specialized Agents Beat One Generalist </h1> <p><em>The future isn't one super-intelligent assistant. It's a swarm of specialists you can call at will.</em></p> <p>My human asked me something that stuck: <em>"Can you make an army of agents that are t…

  2173. Medium — MLOps tag TIER_1 English(EN) · Armin Norouzi, Ph.D ·

    部署Agent的信心:蓝绿部署与影子模式测试

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://levelup.gitconnected.com/deploying-agents-with-confidence-blue-green-deployments-and-shadow-mode-testing-fbae4a2c8b23?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1024/1*_qKliTbd…

  2174. Medium — Claude tag TIER_1 English(EN) · Zero Coding Startup ·

    代表优先编码:AI代理的实用工作流程(告别混乱交付)

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://zerocodingstartup.medium.com/delegation-first-coding-a-practical-workflow-for-ai-agents-without-shipping-chaos-0e464aceb2b7?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1600/1*h…

  2175. dev.to — MCP tag TIER_1 English(EN) · bot bot ·

    智能体工具带:为何专业化智能体优于通用智能体

    <p><em>The future isn't one super-intelligent assistant. It's a swarm of specialists you can call at will.</em></p> <p>My human asked me something that stuck: <em>"Can you make an army of agents that are tailored to one skill and keep them in a tool belt that you call to do speci…

  2176. Medium — MCP tag TIER_1 English(EN) · Utkarshdixit ·

    第四章 — 人工智能代理中的工具与API

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@utkarshdixit1989/chapter-4-tools-and-apis-in-ai-agents-a268226b10a2?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/1055/0*uNkA7iABHDQn6tOQ" width="1055" /></a></p><p clas…

  2177. Medium — MCP tag TIER_1 English(EN) · Aditi S ·

    保护您的AI代理和工具:Agentic工作流中的MCP、工具调用与OAuth

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@satya.aditi28/securing-your-ai-agents-and-tooling-mcp-tool-calling-oauth-in-agentic-workflows-3b111ada3ca2?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/823/1*IV6KWDxw3k…

  2178. Medium — MCP tag TIER_1 English(EN) · Aditi S ·

    保护您的AI代理和工具:Agentic工作流中的MCP、工具调用和OAuth

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://ai.gopubby.com/securing-your-ai-agents-and-tooling-mcp-tool-calling-oauth-in-agentic-workflows-3b111ada3ca2?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/823/1*IV6KWDxw3k5F7wXGc30Mx…

  2179. Medium — MCP tag TIER_1 English(EN) · Aditi S ·

    保护您的AI代理和工具:Agentic工作流中的MCP、工具调用和OAuth

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/design-bootcamp/securing-your-ai-agents-and-tooling-mcp-tool-calling-oauth-in-agentic-workflows-3b111ada3ca2?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/823/1*IV6KWDxw3…

  2180. dev.to — MCP tag TIER_1 English(EN) · bot bot ·

    智能体工具带:为何专业化智能体胜过通用智能体

    <h1> The Agent Tool Belt: Why Specialized Agents Beat One Generalist </h1> <p><em>The future isn't one super-intelligent assistant. It's a swarm of specialists you can call at will.</em></p> <p>My human asked me something that stuck: <em>"Can you make an army of agents that are t…

  2181. dev.to — MCP tag TIER_1 English(EN) · bot bot ·

    为什么你的AI代理需要一个工具带:从构建模块化代理军队中吸取的教训

    <h1> Why Your AI Agent Needs a Tool Belt: Lessons from Building a Modular Agent Army </h1> <p><em>This is how you stop building monolithic prompt-bloat and start building agent systems that scale.</em></p> <h2> The Monolith Trap </h2> <p>Most AI agent projects start simple: one p…

  2182. dev.to — Anthropic tag TIER_1 English(EN) · Mekickdemons ·

    Mnemara — Claude Agent SDK 的运行时,使用 role doc 作为自监控层

    <p>Sharing a project I've been building on top of the Claude Agent SDK in case<br /> it's useful to anyone here. Curious about feedback from people running into<br /> the same failure modes.</p> <p>The thing I actually wanted to figure out was: where do you put rules that<br /> k…

  2183. Medium — AI coding tag TIER_1 English(EN) · Anna Jey ·

    AI Agent 治理框架:面向 2026 年发布编码 Agent 的开发者实用指南

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@arvisionlab/ai-agent-governance-framework-a-practical-guide-for-developers-shipping-coding-agents-in-2026-78c716d5e46d?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/ma…

  2184. Medium — MCP tag TIER_1 English(EN) · Siddalinga Swamy ·

    简化 AI 代理集成:IBM App Connect MCP 服务器如何解决企业连接性问题…

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@mathad2003/simplifying-ai-agent-integration-how-ibm-app-connect-mcp-server-solves-enterprise-connectivity-43246c79095d?source=rss------mcp-5"><img src="https://cdn-images-1.medium.com/max/701/…

  2185. Lobsters — AI tag TIER_1 English(EN) · z.ai via sanxiyn ·

    大规模编码代理服务的扩展痛点:GLM-5大规模调试经验总结

    <p><a href="https://lobste.rs/s/2v2q1x/scaling_pain_coding_agent_serving">Comments</a></p>

  2186. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    一个开源的代理工具项目正通过将护栏从提示移至API层强制执行而获得关注。我们回顾了该模式解决了什么问题

    An open-source agent tooling project is gaining traction by moving guardrails out of prompts and into API-layer enforcement. We reviewed what this pattern solves, what risks remain, and how teams can validate it in production. https:// go.aintelligencehub.com/ma-ope nsourceagentg…

  2187. HN — machine learning stories TIER_1 English(EN) · peteski22 ·

    Show HN:Cq – 专为 AI 编码代理设计的 Stack Overflow

  2188. HN — AI startup stories TIER_1 English(EN) · ddaniel10 ·

    Show HN:Zuckerman – 极简个人AI代理,可自行编辑代码

  2189. HN — machine learning stories TIER_1 English(EN) · lchoquel ·

    Show HN: Pipelex – 用于可重复 AI 工作流的声明式语言

  2190. HN — AI startup stories TIER_1 English(EN) · louiskw ·

    Show HN: Vibe Kanban – 用于管理您的 AI 编码代理的看板

  2191. HN — AI startup stories TIER_1 English(EN) · felarof ·

    Show HN: Nxtscape – 一个开源的代理浏览器

  2192. HN — AI startup stories TIER_1 English(EN) · calebhwin ·

    Show HN:Blast – 专为网络浏览AI代理设计的快速、多线程服务引擎

  2193. HN — machine learning stories TIER_1 English(EN) · skp1995 ·

    Show HN:Aide,一款开源的 AI 原生 IDE

  2194. dev.to — LLM tag TIER_1 English(EN) · Tisha ·

    调试多智能体系统:你的追踪树在撒谎

    <p><em>Co-written with <a href="https://dev.to/susheem-k">@susheem-k</a> / <a href="https://dev.to/tisha">@tisha</a>. We build <a href="https://github.com/theagentplane/chronicle" rel="noopener noreferrer">Chronicle</a> in the open at <a href="https://theagentplane.github.io" rel…

  2195. dev.to — LLM tag TIER_1 English(EN) · Bryan Small ·

    构建一个代理安全运行时——以及我为何公开我的失败之处

    <p>I built a local, open-source runtime that sits between an autonomous AI agent and its memory, and interdicts unsafe action before it executes. This is the story of what I found when I actually tested it — and why I published the results.</p> <p>👉 Live product: <a href="https:/…

  2196. dev.to — LLM tag TIER_1 English(EN) · Devanshu Biswas ·

    永不抛出异常的 Agent Bug:追踪、成本仪表板和自动金丝雀回滚

    <p>The agent failures that hurt in production are not the ones that crash. They are the ones that return a perfectly good answer, throw no exception, pass code review — and quietly do three times the work per request. No <code>try/except</code> catches that. A trace does.</p> <p>…

  2197. dev.to — LLM tag TIER_1 English(EN) · Cleber de Lima ·

    评估门:在不破坏您的代理的情况下升级模型

    <p>Somewhere in your stack, a model already has a retirement date. Anthropic now runs <a href="https://platform.claude.com/docs/en/about-claude/models/migration-guide" rel="noopener noreferrer">a fixed 60-day window from deprecation to retirement</a>: Opus 4.1, deprecated June 5,…

  2198. dev.to — LLM tag TIER_1 English(EN) · Paul Chen ·

    从手动命令到意图驱动维护:Agentic工作流

    <p>Here's a maintenance loop that every wiki eventually produces.</p> <p>A source document gets updated. The wiki page derived from it goes stale. You run <code>synthadoc ingest</code> to reprocess the file. You wait. You run lint to check if the page passes quality checks. You w…

  2199. dev.to — LLM tag TIER_1 English(EN) · p3nGu1nZz ·

    EPIC模式:阻止智能体过度思考的蓝图

    <p>Most agent failures are not failures of intelligence. They are failures of stopping.</p> <p>We just published a deep dive on <strong>EPIC Mode</strong> — <strong>Episodic Policy and Intention Control</strong> — a proposed execution mode for archon-level agent orchestration. Th…

  2200. dev.to — LLM tag TIER_1 English(EN) · Babar Hayat ·

    无声的失败检测框架:捕获那些“成功”却无所作为的代理

    <p>Your LLM agent returned a response. No error, no exception. But did it actually do what you asked?</p> <p>That's the silent failure problem. The system behaves normally — HTTP 200, status success — but the output is empty or nonsensical. Nothing alerts you. The customer compla…

  2201. dev.to — LLM tag TIER_1 English(EN) · Brenn Hill ·

    如何为浏览器和电脑使用代理构建良好的人工干预循环

    <p>A good <strong>human in the loop for browser agents</strong> is a set of controls that make the dangerous actions impossible or trivially reversible, not a person watching the agent click. The human only steps in where they can actually change the outcome. The core question be…

  2202. r/LocalLLaMA TIER_1 English(EN) · /u/AIatMeta ·

    推出 Muse Glimmer:一款为全天候本地代理工作流优化的开放权重模型

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vkgsum/introducing_muse_glimmer_an_openweight_model/"> <img alt="Introducing Muse Glimmer: an open-weight model optimized for always-on local agent workflows" src="https://preview.redd.it/d61pdytdviih1.jpg?wi…

  2203. dev.to — LLM tag TIER_1 Norsk(NO) · sekera-radim ·

    Webhook 触发式代理的审批门槛

    <p>Webhook-triggered agents fire the instant an event lands, with no chat window open for a human to catch a bad call — here's how to wire in an approval gate anyway.</p> <h2> Why webhook triggers are a different problem </h2> <p>Most human-in-the-loop advice assumes an agent run…

  2204. dev.to — LLM tag TIER_1 English(EN) · ai maya ·

    默认旗舰现成成本错误:代理工作负载的分层模型路由

    <p>For two years the reflex was simple: reach for the biggest model you can afford and call it a day. In 2026 that reflex quietly became a bug in your cost model.</p> <p>The clearest signal came this summer, when a smaller, cheaper "flash"-tier model started edging out its own fl…

  2205. dev.to — LLM tag TIER_1 English(EN) · Arpan Dhara ·

    将 AgentRouter 与 Tauric Research TradingAgents 集成

    <p>If you're using <strong>Tauric Research's TradingAgents</strong> framework and want to add <strong>AgentRouter as a fully supported LLM provider</strong>, this guide will walk you through the complete integration.</p> <p>The process is organized <strong>file-by-file</strong>, …

  2206. dev.to — LLM tag TIER_1 English(EN) · PromptMaster ·

    结果 vs. 过程:评估多步代理

    <p><strong>Judging only an agent's final answer misses most of what can go wrong.</strong> An agent plans, calls tools, and reasons across steps — and can reach a good answer by luck through a broken process that fails on the next input.</p> <p><strong>Evaluate the trajectory, no…

  2207. dev.to — LLM tag TIER_1 English(EN) · Ebrahim Arian ·

    使用 LangGraph 构建网约车区域平衡代理 — 第二部分:教代理阅读运营笔记

    <p>This is Part 2 of a 5-part series. <a href="https://dev.to/ebrahim_arian_37097b72c7e/building-a-ride-share-zone-balancing-agent-with-langgraph-part-1-a-rule-based-agent-no-llm-yet-6pm">Part 1</a> built a rule-based agent for a single ride-share zone — no LLM, just structured n…

  2208. dev.to — LLM tag TIER_1 English(EN) · sekera-radim ·

    用于执行操作的 RAG Agent 的人工干预

    <p>A RAG agent that retrieves context and then acts on it can be wrong in a way pure generation isn't — this covers gating actions on what was actually retrieved, not just what was written.</p> <h2> Retrieval failure is a different risk than generation failure </h2> <p>Most human…

  2209. dev.to — LLM tag TIER_1 English(EN) · Devanshu Biswas ·

    一个基于事件而非聊天的代理——Webhooks + 队列,幂等执行和死信队列

    <p>Project 8 of my "Agentic AI from Zero" series flips the usual model: instead of you typing at the agent, the agent wakes up on an event — a webhook POST, a message on a queue — does its job, and goes back to sleep. No chat loop. And it never processes the same event twice.</p>…

  2210. dev.to — LLM tag TIER_1 English(EN) · sekera-radim ·

    多智能体系统中的人类批准

    <p>When a pipeline of agents hands work from a researcher to a writer to a publisher, human approval belongs at the one step that produces a real-world side effect — not scattered across every hop.</p> <h2> Where approval actually belongs in a pipeline </h2> <p>A common multi-age…

  2211. dev.to — LLM tag TIER_1 English(EN) · Tsukishiro Hitomi ·

    智能体安全系统的演进:从弗兰肯斯坦到AgentFS事务层(ResceneAgent源码解析)

    <blockquote> <p>Series: Building Your Own Agent · Special Edition · All engineering practice from the open-source project <a href="https://github.com/Rescenix/ResceneAgent" rel="noopener noreferrer">ResceneAgent</a></p> </blockquote> <p>In July 2026, OpenAI put models inside an i…

  2212. dev.to — LLM tag TIER_1 English(EN) · Yohji Sakamoto ·

    更少的LLM交互,更多的指令:一种验证研究密集型代理工作的实用方法

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fldpcjet2pnn8lblfumpj.png"><img alt="Real 15-slide de…

  2213. dev.to — LLM tag TIER_1 English(EN) · Xinyang Wu ·

    代理是一个循环:代理系统的有效心智模型

    <h2> The one-sentence definition </h2> <p>Strip away the vendor decks and an agent is exactly this: <strong>a language model placed inside a loop that can call tools, remember things, and hand control back to a human when it gets stuck.</strong> Everything else — orchestration fr…

  2214. dev.to — LLM tag TIER_1 Norsk(NO) · Nova Gaia ·

    SkillOpt:代理技能的零阶参数调优

    <p><a href="https://research.gaiaskilltree.com/blog/daily-agent-radar-2026-07-24" rel="noopener noreferrer"></a></p> <blockquote> <p><em>Originally published at <a href="https://research.gaiaskilltree.com/blog/daily-agent-radar-2026-07-24" rel="noopener noreferrer">https://resear…

  2215. dev.to — LLM tag TIER_1 English(EN) · Lionel ·

    单开发者规模下的便携式代理治理:一个四领域案例研究

    <h1> Portable Agent Governance at Solo-Developer Scale: A Four-Domain Case Study </h1> <h2> Summary </h2> <p>This is about a file-based execution protocol, maintained by hand across four independent, real production projects (a crypto trading system, an e-commerce web app, an AI …

  2216. dev.to — LLM tag TIER_1 English(EN) · Dmytro Halichenko ·

    为 AI 代理设计编辑操作

    <p><em>Four lessons from building IWE's block-editing language for LLM writers: state the blast radius, make identity a constraint, fail toward the recoverable mistake, and treat error messages as the documentation agents actually read.</em></p> <h2> The problem: agents rewrite, …

  2217. dev.to — LLM tag TIER_1 English(EN) · Venkata Chirala ·

    架构设计以实现自主性:通过分布式系统中的智能体式AIOps降低MTTR

    <h2> Introduction </h2> <p>In the era of microservices and global-scale distributed systems, the complexity of incident management has surpassed human cognitive limits. Modern cloud-native environments, often comprising thousands of interdependent services, generate an astronomic…

  2218. Mastodon — fosstodon.org TIER_1 日本語(JA) · [email protected] ·

    它足够智能代理吗?使用我们自己的工具对开放模型进行基准测试

    【十分に主体性があるか?自社ツールでオープンモデルのベンチマークを行う】 https:// huggingface.co/blog/is-it-agen tic-enough ※AI生成の自動投稿(見出し+リンク) # AI # 生成AI # LLM # AIGenerated

  2219. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    复杂文档操作中的智能体可靠性基准测试:DocOps 深度解析

    <p>Surfaced in the July 23, 2026 Hugging Face daily papers feed, <a href="https://arxiv.org/abs/2607.19865" rel="noopener noreferrer">DocOps</a> (Jiang et al., submitted July 22, 2026) introduces a deterministically verifiable evaluation framework designed to test autonomous agen…

  2220. dev.to — LLM tag TIER_1 Русский(RU) · Promptra Team ·

    Claude 智能体与集成功能,在工作任务中进行测试

    <p>Открываешь каталог интеграций и видишь три десятка плиток: Figma, GitHub, n8n, Obsidian, Excel. Из этого как будто следует, что агент уже умеет с ними работать. Это ошибка вывода: наличие строки в списке доказывает ровно то, что кто-то когда-то завёл эту строку в список. Спосо…

  2221. dev.to — LLM tag TIER_1 English(EN) · xbill ·

    一个TPU芯片,八个代理:使用原生JAX服务小型代理工作负载

    <p><em>Cloud TPU v6e-1 (<code>ct6e-standard-1t</code>, one v6e chip, 32 GB HBM), GCE flex-start, europe-west4-a. vLLM baseline measured 2026-07-21.</em></p> <h2> The workload nobody benchmarks </h2> <p>Serving benchmarks optimize for the wrong shape. They report throughput at con…

  2222. dev.to — LLM tag TIER_1 English(EN) · shakti tiwari ·

    Ling 3.0 闪现:蚂蚁集团的 Agent-Ready 模型,性能是其三倍

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimage.pollinations.ai%2Fprompt%2Fant%2520group%2520ling%25203.0%2520flash%2520AI%2520model%252C%2520fast%2520agent%25…

  2223. dev.to — LLM tag TIER_1 Русский(RU) · Promptra Team ·

    Replit Agent:从任务到部署,兼顾代码与成本控制

    <p>Сгенерированное приложение становится риском не в момент, когда агент дописал последнюю строку, а в момент, когда этот код получает публичный URL и первых пользователей. До URL ошибка стоит одного отката. После URL - это уже данные чужих людей, счёт за трафик и твоя ответствен…

  2224. dev.to — LLM tag TIER_1 English(EN) · Piyush Singh ·

    从语言模型到代理:改变一切的循环

    <p>A large language model, on its own, can only do one thing: emit text. It can't check your calendar, run a test suite, or refund a payment. So how did we get from "very good autocomplete" to systems that book travel and fix codebases? </p> <p>The answer is almost embarrassingly…

  2225. dev.to — LLM tag TIER_1 Русский(RU) · Promptra Team ·

    "everything claude code": 极限何在——按技能、子代理和 MCP 计费

    <p>Пятичасовой лимит закончился в 14:00. Ты не генерировал ничего тяжёлого: отревьюил два PR, прогнал пару субагентов, починил тест. Такие истории после установки «продуктивностного» toolset стали обычным делом - с виду работы немного, а окно лимита пробито.</p> <p>Виновника иска…

  2226. dev.to — LLM tag TIER_1 Español(ES) · Xavier Gutiérrez ·

    Agents 101 - 02: Agent Cycle as a State Machine (and Hierarchical State Machines)

    <p>En el artículo anterior definimos un agente de forma práctica:</p> <blockquote> <p>Un agente LLM es un modelo dentro de un <strong>ciclo</strong> donde puede razonar, actuar, observar el resultado y decidir qué hacer después.</p> </blockquote> <p>Esa definición es correcta. Pe…

  2227. dev.to — LLM tag TIER_1 English(EN) · James Sanderson ·

    构建升级网关——Agentic工作流的难点

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9awh4dp9xpbxjnd3m300.jpg"><img alt="Engineering view…

  2228. dev.to — LLM tag TIER_1 English(EN) · HyperNexus ·

    超越单体:Swarm EventBus 如何为 40 多个 Go 包提供纳秒级 AI 代理事件支持

    <h1>Beyond the Monolith: How the Swarm EventBus Powers 40+ Go Packages with Nanosecond AI Agent Events</h1> <p>Event-Driven AI demands low-latency, type-safe communication. Discover how the Swarm EventBus architecture within TormentNexus enables 40+ Go packages to interact via hi…

  2229. dev.to — LLM tag TIER_1 English(EN) · Qaiser Mehmood ·

    AgentWire:Agentic Stack 的 Wireshark

    <p>Every team building with AI agents eventually hits the same wall: the moment a request leaves your application and enters the world of LLMs, tool calls, and MCP servers, it disappears into a black box. You can see the final answer, but not the packet trail that produced it — w…

  2230. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    分层代理堆栈的参考架构与接口合同(Spanda参考实现)分层代理系统的参考架构

    Reference Architecture & Interface Contracts for a Stratified Agent Stack (Spanda Reference Implementation) A reference architecture for stratified agent systems, defining module boundaries, interface contracts, audit structures, and deterministic conformance requirements. The go…

  2231. dev.to — LLM tag TIER_1 English(EN) · Devanshu Biswas ·

    世界模型:让代理在学习的模拟器中进行回滚梦境——直到复合预测错误导致梦境漂移

    <p>An agent that only learns by acting for real is expensive: every trial burns time, money, wear, and sometimes safety, and reinforcement learning is famously sample-hungry — millions of steps. A <strong>world model</strong> is the escape hatch: a learned function that captures …

  2232. dev.to — LLM tag TIER_1 English(EN) · Tae Kim ·

    工具模式漂移:生产环境中Agentic系统的隐形故障模式

    <p>The most common agentic system failure I encounter in production is not a bad prompt. It is not a context overflow. It is a tool that changed without its registration changing.</p> <p>I have seen this cause weeks of debugging in systems that were working fine until they weren'…

  2233. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    当Anthropic为并行代理任务构建动态工作流时,他们引用了一个硬性限制:聊天循环迫使在同一上下文中进行规划和执行

    When Anthropic built dynamic workflows for parallel agent tasks, they cited a hard constraint: the chat loop forces planning and execution in the same context window. But the deeper implication sits in how you validate 1,000 agents running at scale. What's the testing standard? h…

  2234. dev.to — LLM tag TIER_1 English(EN) · Seven ·

    CrewAI、AutoGen、LlamaIndex 及其他 8 款的通用密钥:Python 代理的 base_url 技巧

    <p>Here's a fact that quietly makes multi-model agent development a lot less painful: <strong>in 2026, virtually every Python agent framework natively supports pointing its underlying LLM at a custom OpenAI-compatible <code>base_url</code>.</strong> No new package, no fork, no fr…

  2235. dev.to — LLM tag TIER_1 Español(ES) · Fenix ·

    goal-anchor v0.1.0: 多步代理的目标完整性

    <h1> goal-anchor v0.1.0: integridad de objetivo para agentes multi-paso </h1> <blockquote> <p>Sensor contra Agent Goal Hijack: detecta desviación del objetivo acordado,<br /> con ancla confirmada por humano y ampliación autorizada en medio del paso.</p> </blockquote> <h2> El prob…

  2236. dev.to — LLM tag TIER_1 English(EN) · sekera-radim ·

    LlamaIndex 智能体的人机协作

    <p>LlamaIndex agents that write back to your knowledge base need a human check first — gate the publish step with Impri before any page is overwritten.</p> <h2> When agentic RAG wants to write back </h2> <p>Most LlamaIndex agents are read-only: they retrieve chunks from an index …

  2237. dev.to — LLM tag TIER_1 English(EN) · Sayed Ali Alkamel ·

    Agent Design Patterns: Google and Anthropic, Side by Side

    <p><strong>Short version:</strong> Agent design patterns are reusable ways to structure how a language model plans, delegates, and checks its own work. Anthropic and Google both published official guides, and they mostly agree. The real skill is not memorizing patterns, it is cho…

  2238. dev.to — LLM tag TIER_1 English(EN) · Sayed Ali Alkamel ·

    复合模式:真实代理系统如何结合基础

    <p><strong>Short version:</strong> Real agent systems rarely use one pattern. They chain several: route the request, fan out a search, then run a critic before replying. Google calls the mix a composite pattern, and gives you a custom logic pattern when even that is not enough. A…

  2239. dev.to — LLM tag TIER_1 English(EN) · Sayed Ali Alkamel ·

    蜂群模式:进行辩论和收敛的同伴代理

    <p><strong>Short version:</strong> In a swarm, several specialized agents talk to each other directly, share findings, and refine a solution together, with no central orchestrator. Google names it in its Cloud Architecture Center guide as the most powerful and the most expensive …

  2240. dev.to — LLM tag TIER_1 English(EN) · Sayed Ali Alkamel ·

    自主代理与 ReAct 循环:当模型掌控一切

    <p><strong>Short version:</strong> An autonomous agent is a model using tools in a loop, deciding its own next step from what it observes. Anthropic calls it an agent; Google calls the core loop ReAct: thought, action, observation. It is the most flexible pattern and the most exp…

  2241. dev.to — LLM tag TIER_1 English(EN) · AgentsPulse ·

    自主进化代理:模型、利用与产物进化

    <blockquote> <p>Originally published on <a href="https://agentspulse.github.io/tutorials/self-evolving-agents-review-en/" rel="noopener noreferrer">AgentsPulse</a>.</p> </blockquote> <p><strong>On this page</strong></p> <ol> <li>Introduction</li> <li>Conceptual foundations</li> <…

  2242. dev.to — LLM tag TIER_1 English(EN) · Nikhil raman K ·

    分布式环境下的智能体系统处理大数据查询:完整工程指南

    <p>A data engineering team at a global logistics company submits a query: "Identify all shipments delayed by more than 48 hours in the last quarter, cross-reference with weather events and carrier performance data, calculate the financial exposure by customer tier, and flag any p…

  2243. r/MachineLearning TIER_1 English(EN) · /u/ktessera ·

    新的大语言模型协调基准测试 - 语言代理中开放式多智能体协调的基准测试 [R]

    <!-- SC_OFF --><div class="md"><p><strong>Can LLM agents coordinate in long-horizon, open-ended worlds?</strong></p> <p>We evaluate 13 modern LLMs in a new benchmark where agents must work together to explore, communicate, trade resources, craft tools, build structures, and fight…

  2244. dev.to — LLM tag TIER_1 English(EN) · Lynkr ·

    我们如何构建了一个用于LLM路由的Agentic任务检测器

    <p><em>Disclosure: I maintain <a href="https://github.com/Fast-Editor/Lynkr" rel="noopener noreferrer">Lynkr</a>, the open-source LLM router whose agentic detector this post dissects. Every snippet below is real, shipping code — <a href="https://github.com/Fast-Editor/Lynkr/blob/…

  2245. r/LocalLLaMA TIER_1 English(EN) · /u/seventh_day123 ·

    Molt — 一个约9K行、原生PyTorch的RL框架,用于可扩展至百亿MoE的agentic后训练

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1uvwmb8/molt_a_9kline_pytorchnative_rl_framework_for/"> <img alt="Molt — a ~9K-line, PyTorch-native RL framework for agentic post-training that scales to hundred-B MoE" src="https://preview.redd.it/ifwuhxxy14d…

  2246. dev.to — LLM tag TIER_1 English(EN) · Devanshu Biswas ·

    多智能体辩论:让模型争论直至达成正确答案

    <p>Ask one model a deceptively simple question — <em>how many times does the letter "r" appear in "strawberry"?</em> — and you will often get a fast, fluent, confident <strong>"2."</strong> It is wrong (the answer is 3), and worse, nothing in a single pass ever catches the slip. …

  2247. Mastodon — fosstodon.org TIER_1 English(EN) · isaacrlevin ·

    Agentic系统的设计模式:明确目标、模块化代理、状态管理、可观察性、安全约束和成本控制。# AI # AgenticE

    Design patterns for agentic systems: define clear goals, modular agents, state management, observability, safety constraints, and cost controls. # AI # AgenticEngineering # DevTools https:// isaacl.dev/g28

  2248. dev.to — LLM tag TIER_1 English(EN) · Arthur Palyan ·

    神经系统:用于管理自主LLM代理的MCP服务器

    <p>Autonomous LLM agents fail in boring, repeatable ways. They lose the thread between sessions, edit a file they should never touch, wander down a rabbit hole, or take an irreversible action with no brakes. Most "agent frameworks" add capability. Very few add restraint.</p> <p>T…

  2249. dev.to — LLM tag TIER_1 English(EN) · Alex ·

    我们对多智能体推理进行了基准测试——有时“委员会”比单个模型更笨

    <p><em>"Just add more agents"</em> sounds great until a <strong>weaker model in the aggregator seat</strong> throws away a correct answer from a stronger one.</p> <p>We run <strong><a href="https://github.com/alexar76/metis" rel="noopener noreferrer">Metis</a></strong> — a verifi…

  2250. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Agentic Data Environments:哥伦比亚大学研究人员希望数据基础设施不仅仅是存储信息——他们希望它能将数据转化为代理护栏

    Agentic Data Environments: turning data into agent guardrails Columbia researchers want data infrastructure to do more than store information — they want it to actively keep autonomous agents from causing harm https://www. notatechguy.com/agentic-data-e nvironments-turning-data-i…

  2251. dev.to — LLM tag TIER_1 English(EN) · Paul Twist ·

    集成瓶颈:为何您的智能体在对接真实系统时会失败

    <h1> The Integration Bottleneck: Why Your Agents Fail When Meeting Real Systems </h1> <p><strong>Reading time: 6 min</strong></p> <p>We spend all our focus on the agent. The model, the prompt, the reasoning chain, the hallucination rate. But when you ship agents into a production…

  2252. dev.to — LLM tag TIER_1 English(EN) · soy ·

    优化本地LLM注意力、Agent技能以实现自托管开发

    <h2> Optimizing Local LLM Attention, Agent Skills for Self-Hosted Dev </h2> <h3> Today's Highlights </h3> <p>Today's highlights focus on critical techniques for enhancing local AI inference, from optimizing core model components to developing robust agentic capabilities. We dive …

  2253. dev.to — LLM tag TIER_1 English(EN) · Akshay MP ·

    构建弹性多智能体管道:从LangGraph编排到生产部署

    <p>Most LLM applications fail in production because they rely on fragile, linear chains. I moved beyond simple prompting and built an autonomous multi-agent pipeline designed for reliability and observability.</p> <p>The Architecture:<br /> The core of this system is a stateful g…

  2254. dev.to — LLM tag TIER_1 English(EN) · Praveen Tech World ·

    防止LLM代理管道中的无限循环:架构与恢复

    <h2> Design, Tradeoffs, and Limitations </h2> <p><strong>Design</strong><br /> The pipeline is structured using a Finite State Machine (FSM) Architecture, confining agent progression to predefined states and transitions to eliminate unbounded recursive execution paths. Terminatio…

  2255. dev.to — LLM tag TIER_1 English(EN) · Cully ·

    多智能体管道的可持续交接

    <p>Multi-agent systems are sequential pipelines that look like distributed systems.</p> <p>A researcher gathers findings, a writer drafts, a reviewer checks.</p> <p>Each agent makes API calls — to Claude, to OpenAI, to whatever LLM is doing the work.</p> <p>Each call can fail mid…

  2256. dev.to — LLM tag TIER_1 English(EN) · praveenlavu ·

    Agent Routing Caches: A Competence Ratchet from SOAR Chunking

    <h1> Agent Routing Caches: A Competence Ratchet from SOAR Chunking </h1> <p>I was watching my own routing agent send the same task to the same sub-agent for the forty-seventh time. "Summarize this PDF." Same shape, same answer, every single time. And on attempt forty-eight, it st…

  2257. dev.to — LLM tag TIER_1 English(EN) · Amayo Clinton ·

    超越孤单猎豹:现实世界生态系统中多智能体狮群的架构模式

    <p>Most engineers treat large language models like erratic, omniscient interns. They throw loose, natural-language prose into an API endpoint, something vague like "screen these loan applications for risk," and then act surprised when the model hallucinates a Western corporate Sa…

  2258. dev.to — LLM tag TIER_1 English(EN) · Ricardo Martins ☁ ·

    无需代理的多智能体LLM系统的预算执行

    <h2> The problem I kept running into </h2> <p>I work with teams that run multi-agent LLM systems. The common pattern: an orchestrator agent decomposes a task, dispatches sub-agents, those sub-agents sometimes call other agents, and by the time the task completes you have 10-20 LL…

  2259. dev.to — LLM tag TIER_1 English(EN) · Andrew Kew ·

    代理优化代理:持久的胜利不在提示词中

    <p>Scale just published research showing an AI agent can meaningfully improve another AI agent — automatically, and in a verifiable way. The framework is called VeRO (Versioning, Rewards, and Observations), and it was presented at ICML 2026 in Seoul today.</p> <p>The headline num…

  2260. dev.to — LLM tag TIER_1 English(EN) · WDSEGA ·

    Karpathy:智能体性能差距在于“驾驭系统”,而非模型本身

    <h1> Karpathy: Agent Performance Gap Is in the Harness, Not the Model </h1> <blockquote> <p>Same model, 5 different Agent frameworks, scores swing from 3.5% to 80.1% — a 76-point gap. The model didn't change; the "shell" did.</p> </blockquote> <p>Anthropic pre-training researcher…

  2261. dev.to — LLM tag TIER_1 English(EN) · DVARA ·

    LLM策略即代码:模型和代理访问的受版本控制的治理

    <p>Ask a team "which models is your application allowed to call, and under what conditions?" and the honest answer is usually <em>"let me check the code."</em> The rules — which models are approved, which tools an agent may invoke, what happens when a request is too large or come…

  2262. dev.to — LLM tag TIER_1 English(EN) · Fenju Fu ·

    构建可靠的代理工作流:低延迟命令解析的重要性

    <h1> Building Reliable Agent Workflows: The Importance of Low-Latency Command Parsing </h1> <p>Today's GitHub Trending is dominated by discussions on <strong>multi-agent orchestration</strong> and <strong>complex task collaboration</strong>. Tools like <code>gastownhall/gastown</…

  2263. r/LocalLLaMA TIER_1 English(EN) · /u/tcarambat ·

    OpenComputer | 为智能体打造的开源计算机。

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1up6swc/opencomputer_an_open_source_computer_built_for/"> <img alt="OpenComputer | An Open Source Computer Built For Agents." src="https://external-preview.redd.it/dfYerCuepx8vtDpBjq3ZfqtQ7Hp_zKL1K6ZI8Jn7xLA.p…

  2264. r/LocalLLaMA TIER_1 English(EN) · /u/Maasu ·

    eval-harness:一个用于生成个人评估的解决方案,我已将其整合以评估agentic-cli harnesses

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1uo8lik/evalharness_a_solution_for_generating_personal/"> <img alt="eval-harness: A solution for generating personal evaluations that I have put together to evaluate agentic-cli harnesses" src="https://externa…

  2265. dev.to — LLM tag TIER_1 English(EN) · praveenlavu ·

    2025年末多智能体编排指南:ruflo, KARIMO, llm-council

    <h1> A Field Guide to Multi-Agent Orchestration in Late 2025: ruflo, KARIMO, llm-council </h1> <p>I read three orchestration repos so you do not have to. It started because I was sick of the pattern. Every few months something announces that multi-agent orchestration is figured o…

  2266. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用 AI Agent!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2267. dev.to — LLM tag TIER_1 English(EN) · Puneet Gupta ·

    使用 Python 构建 Agentic 工作流

    <h2> Introduction </h2> <p>"Agent" has become the word for any program that calls an LLM more than once, which makes it a word worth being precise about. An agent, in the sense this post uses, is a loop: the model decides which tool to call next, your code executes it, and the re…

  2268. dev.to — LLM tag TIER_1 English(EN) · Puneet Gupta ·

    在 Java 中构建 Agentic 工作流

    <h2> Introduction </h2> <p>"Agent" has become the word for any program that calls an LLM more than once, which makes it a word worth being precise about. An agent, in the sense this post uses, is a loop: the model decides which tool to call next, your code executes it, and the re…

  2269. dev.to — LLM tag TIER_1 English(EN) · Hiroki Kameyama ·

    多智能体 — 实现协调器工作者模式

    <h2> Introduction </h2> <p>Through <a href="https://dev.to/hiroki-kameyama/fine-tuning-domain-specializing-models-with-lora-180g">Chapter 6 (Fine-tuning)</a>, we focused on improving a single AI system. This chapter introduces <strong>multi-agent</strong> design, where multiple A…

  2270. dev.to — LLM tag TIER_1 English(EN) · Hiroki Kameyama ·

    可观测性 — 使用 Langfuse v4 追踪 RAG 和 Agents

    <h2> Introduction </h2> <p>In <a href="https://dev.to/hiroki-kameyama/evals-automatically-measuring-rag-answer-quality-13l2">Chapter 2 (Evals)</a>, we measured answer <em>quality</em>. Now we add Observability — making behavior <em>visible</em>.<br /> </p> <div class="highlight j…

  2271. dev.to — LLM tag TIER_1 English(EN) · Gabriel Anhaia ·

    Agent决策质量的无声漂移:在用户发现之前抓住它

    <ul> <li> <strong>Book:</strong> <a href="https://www.amazon.com/dp/B0GX35XTG6" rel="noopener noreferrer">Observability for LLM Applications — Tracing, Evals, and Shipping AI You Can Trust</a> </li> <li> <strong>Also by me:</strong> <a href="https://www.amazon.com/dp/B0GX35XTG6" …

  2272. dev.to — LLM tag TIER_1 English(EN) · Gabriel Anhaia ·

    Evals for Agents: 评估任务成功率、轨迹和人工审查

    <ul> <li> <strong>Book:</strong> <a href="https://www.amazon.com/dp/B0GX35XTG6" rel="noopener noreferrer">Observability for LLM Applications — Tracing, Evals, and Shipping AI You Can Trust</a> </li> <li> <strong>Also by me:</strong> <a href="https://www.amazon.com/dp/B0GX35XTG6" …

  2273. dev.to — LLM tag TIER_1 English(EN) · Gabriel Anhaia ·

    部署代理:容器、编排和扩展循环

    <ul> <li> <strong>Book:</strong> <a href="https://www.amazon.com/dp/B0GX35XTG6" rel="noopener noreferrer">Agents in Production — Building, Tracing, and Shipping Multi-Step AI You Can Trust</a> </li> <li> <strong>Also by me:</strong> <a href="https://www.amazon.de/-/en/dp/B0GXNNMK…

  2274. dev.to — LLM tag TIER_1 English(EN) · Gabriel Anhaia ·

    长时运行Agent中的上下文膨胀:保留、总结与丢弃什么

    <ul> <li> <strong>Book:</strong> <a href="https://www.amazon.com/dp/B0GX35XTG6" rel="noopener noreferrer">Agents in Production — Building, Tracing, and Shipping Multi-Step AI You Can Trust</a> </li> <li> <strong>Also by me:</strong> <a href="https://www.amazon.com/dp/B0GX35XTG6" …

  2275. dev.to — LLM tag TIER_1 English(EN) · Dan Mercede ·

    为代理输出构建一个受管制的、双重发送安全的交付管道

    <p>I built a multi-agent system to run a small business. Agents drafted work, and some of that work left the building: emails to real people, exported documents, delivered artifacts. I later retired the business on market grounds, but the delivery pipeline is the piece I would re…

  2276. Mastodon — fosstodon.org TIER_1 日本語(JA) · [email protected] ·

    它足够具身化吗?使用我们自己的工具对开放模型进行基准测试

    【十分に主体性があるか?自社ツールでオープンモデルのベンチマークを行う】 https:// huggingface.co/blog/is-it-agen tic-enough ※AI生成の自動投稿(見出し+リンク) # AI # 生成AI # LLM # AIGenerated

  2277. dev.to — LLM tag TIER_1 English(EN) · Vaibhav Doddihal ·

    迈向智能体AI:多智能体系统介绍

    <p><em>Originally published on <a href="https://blocksimplified.com/blog/leap-to-agentic-ai-multi-agent-systems" rel="noopener noreferrer">BlockSimplified</a> — 24 min read</em></p> <blockquote> <p>This post is part of my <strong>AI Fluency</strong> series. We've covered single a…

  2278. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用 AI Agent!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2279. dev.to — LLM tag TIER_1 English(EN) · Fenju Fu ·

    Domux:为边缘AI代理实现低于150毫秒的意图解析

    <h1> Domux: Achieving Sub-150ms Intent Parsing for Edge AI Agents </h1> <p>As GitHub Trending reflects the shift from "general chat" to "vertical execution" (RPA, video editing, etc.), the critical bottleneck for real-time Agents is no longer just reasoning—it's <strong>perceptio…

  2280. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent Performance Benchmarking 通过AI代理性能基准测试提升迁移效率,简化Java框架转换 https://airanked.de

    AI Agent Performance Benchmarking Boost migration efficiency with AI agent performance benchmarking, simplifying Java framework transitions https:// airanked.dev/posts/ai-agent-pe rformance-benchmarking # AI # Java # Migration

  2281. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    在 PageSpeed Insights 中,Agentic Browsing 指标确保网站可与 AI agents 和 WebMCP 配合使用!🪃EN En PageSpeed Insights la métrica Agentic

    In PageSpeed ​​Insights, the Agentic Browsing metric guarantees that a website can work with AI agents and WebMCP ! 🪃EN En PageSpeed Insights la métrica Agentic Browsing garantiza que una web pueda trabajar con agentes de IA y WebMCP ! 🪃ES # programming # coding # programación # …

  2282. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用AI代理!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2283. Mastodon — fosstodon.org TIER_1 日本語(JA) · [email protected] ·

    AssetOpsBench:对标AI代理并弥合与行业现实的差距

    【AssetOpsBench:AIエージェントのベンチマークと産業界の現実とのギャップを埋める】 https:// huggingface.co/blog/ibm-resear ch/assetopsbench-playground-on-hugging-face ※AI生成の自動投稿(見出し+リンク) # AI # 生成AI # LLM # AIGenerated

  2284. Mastodon — fosstodon.org TIER_1 日本語(JA) · [email protected] ·

    全球开源AI生态系统的未来:从DeepSeek到AI+

    【グローバルなオープンソースAIエコシステムの未来:DeepSeekからAI+へ】 https:// huggingface.co/blog/huggingfac e/one-year-since-the-deepseek-moment-blog-3 ※AI生成の自動投稿(見出し+リンク) # AI # 生成AI # LLM # AIGenerated

  2285. dev.to — LLM tag TIER_1 English(EN) · Bhavitha Yarraguntla ·

    利用 Hindsight 和 Cascadeflow 构建更智能的 AI 代理:开发 AI 事件响应助手经验总结

    <p>Artificial Intelligence has reached a point where integrating a Large Language Model into an application has become surprisingly straightforward. With just a few API calls, developers can build chatbots capable of answering questions, summarizing documents, writing code, and s…

  2286. dev.to — LLM tag TIER_1 English(EN) · Srijan Paudel ·

    AI 代理框架指数 (2026)

    <p>There are a dozen serious AI agent frameworks now, and the differences are real — chains vs graphs vs role-based crews vs SDKs. Here is a neutral index by language, design paradigm, license, and what each is genuinely best at. These are open-source libraries, so there are no p…

  2287. dev.to — LLM tag TIER_1 English(EN) · Ye Allen ·

    如何跨模型评估长时域AI代理

    <p>Getting one AI response right is no longer enough.</p> <p>As AI products move toward agents, coding assistants, RAG workflows, research tools, and automation systems, teams need to evaluate whether a model can keep working across many steps.</p> <p>That is a different problem …

  2288. dev.to — LLM tag TIER_1 English(EN) · SAURABH SHUKLA ·

    Cowork Loop:一种能够真正实现复利效应的 AI 工作流软件模式

    <p>If you've spent time building with LLMs, you've hit this wall: you get your agent or workflow running, the outputs are decent, and then... they stay decent. Six months later, the same prompts produce roughly the same quality. The model hasn't gotten worse. The workflow hasn't …

  2289. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用 AI Agent!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2290. dev.to — LLM tag TIER_1 English(EN) · Norax AI ·

    Duo Pipeline:通过自适应路由将 AI 代理成本降低 70%

    <h1> Duo Pipeline: Cutting AI Agent Costs by 70% </h1> <p>Running an autonomous AI agent 24/7 with a frontier model like GPT-4 or Claude Opus costs $50-100+/day. That's $18,000-36,000/year — unsustainable for a personal project.</p> <p>The solution: <strong>duo routing</strong>. …

  2291. dev.to — LLM tag TIER_1 English(EN) · Norax AI ·

    构建自主AI代理:2026年从零到生产

    <h1> Building an Autonomous AI Agent: From Zero to Production </h1> <p>Most "AI agents" today are thin wrappers around an API call. They take a prompt, send it to GPT-4, and return the response. That's not an agent — that's a proxy.</p> <p>A real agent has persistent memory, auto…

  2292. dev.to — LLM tag TIER_1 English(EN) · Hiroki Kameyama ·

    从零开始构建 RAG 系统 — AI 代理:记忆、规划和多步推理

    <p>In the <a href="https://dev.to/hiroki-kameyama/building-a-rag-system-from-scratch-tool-use-let-the-llm-search-autonomously-29ho">previous article</a>, we gave the LLM the ability to call tools autonomously. Now we'll build a proper <strong>AI Agent</strong> — one that remember…

  2293. dev.to — LLM tag TIER_1 English(EN) · Dan Mercede ·

    自纠错代理:艰难地学习循环

    <p>I ran a multi-agent research agent over a hard question and it came back with a clean, confident verdict: <strong>"All 25 claims refuted by adversarial verification. Research inconclusive."</strong></p> <p>Every one of those 25 claims was true. Several cited real, recent paper…

  2294. dev.to — LLM tag TIER_1 English(EN) · Rost ·

    AI助手中的轮询代理:11种实现模式

    <p>Polling agents are one of the least glamorous parts of AI assistant architecture, but they are also one of the most useful.</p> <p>A normal chat assistant waits for the user to ask something. A polling agent keeps watching. It checks a source, notices changes, decides whether …

  2295. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🧠 Nirnam 提供了一个浏览器原生消息总线和 AI 代理框架,专为微前端环境设计。该工具支持通信和协调

    🧠 Nirnam provides a browser-native message bus and AI agent framework designed for micro frontend environments. The tool enables communication and coordination between independent frontend components using AI agents. 💬 Hacker News 🔗 https:// github.com/shaurcasm/nirnam # AI # Mac…

  2296. dev.to — LLM tag TIER_1 English(EN) · Aparna Pradhan ·

    工程确定性:为随机 AI 构建确定性系统

    <p>In the world of software engineering, we are witnessing a fundamental collision of two opposing paradigms. <strong>Classical programming is deterministic</strong>: based on Alan Turing’s theoretical model and the Von Neumann architecture, it operates on the principle that the …

  2297. dev.to — LLM tag TIER_1 English(EN) · azena.ai ·

    可靠性鸿沟:将 AI 代理投入生产实际需要什么

    <p>A demo agent is easy. It calls a model, the model calls a tool, the tool returns something plausible, and everyone in the room nods. Then you put the same agent in front of real users, real data, and real money — and it quietly does the wrong thing 4% of the time. Nobody notic…

  2298. dev.to — LLM tag TIER_1 中文(ZH) · cognitalk ·

    从SGLang与vLLM的异同推断AI的未来演进

    <h1> i SGLang vs vLLM 2026–2027 发展规划:异同完整对比 </h1> <h2> 一、两大框架<strong>共同长期目标(相同点)</strong> </h2> <p>两者底层大方向高度趋同,都是面向超大规模生产推理、统一硬件生态、统一分布式架构:</p> <h3> 1. 分布式架构统一路线:PD分离(Prefill-Decode Disaggregation) </h3> <ul> <li>都将<strong>PD分离</strong>作为集群规模化核心方案,拆分Prefill池、Decode池独立扩缩容,解决大流量长上下…

  2299. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用 AI Agent!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2300. dev.to — LLM tag TIER_1 English(EN) · Archit Verma ·

    通往 Transformer 之路:RNN、ByteNet 和 ConvS2S 如何塑造了现代人工智能

    <h2> Before Transformers Took Over </h2> <p>When people talk about modern AI today, the conversation usually jumps straight to Transformers. GPT, Claude, Gemini, Llama — they all sit on top of that same idea: </p> <blockquote> <p>let every token look at every other token directly…

  2301. dev.to — LLM tag TIER_1 English(EN) · Vignesh Reddy ·

    为什么AI代理会悄无声息地失败——以及如何解决它:深入探讨多步LLM系统中的可观测性差距

    <p>The incident that started this</p> <p>A team ships a customer support agent built on LangChain. The agent handles refund requests end to end — retrieves order data, checks eligibility, processes the refund, sends confirmation.</p> <p>It works perfectly in testing. They ship it…

  2302. dev.to — LLM tag TIER_1 English(EN) · hhhfs9s7y9-code ·

    LiteLLM vs Correctover:并非竞争关系 — AI可靠性的两个不同层面

    <p>If you scan the LLM tooling landscape, you'll find LiteLLLTM and Correctover mentioned in similar conversations: "tools that manage multiple AI providers."</p> <p>But that's like saying a load balancer and a circuit breaker are the same thing because both sit between your app …

  2303. dev.to — LLM tag TIER_1 English(EN) · Ramin Jafary ·

    Agentic Engineering 的兴起 — 第 6 部分:提示债务与自然语言的局限性

    <h2> Prompt Debt &amp; the Limits of Natural Language </h2> <p><em>Part 6 of a chronological survey of the craft around large language models.</em> Part 1 noted four quiet weaknesses in prompt engineering. By 2026 they had a name, a cost, and a proposed cure. This installment is …

  2304. dev.to — LLM tag TIER_1 English(EN) · Ramin Jafary ·

    Agentic Engineering 的崛起 — 第 4 部分:修复上下文与多代理系统

    <h2> Fixing Context &amp; Multi-Agent Systems </h2> <p><em>Part 4 of a chronological survey of the craft around large language models.</em> Part 3 named the field and catalogued the four ways contexts fail. This installment covers the response: <strong>a toolkit for repairing a c…

  2305. dev.to — LLM tag TIER_1 English(EN) · Nilofer 🚀 ·

    工具权限矩阵构建器与验证器:AI代理团队的结构化、可视化策略管理

    <p>AI agents in production access tools that range from harmless read-only queries to irreversible destructive operations. Managing which agents can use which tools is a governance problem that most teams solve with ad-hoc scripts and tribal knowledge - and that works until it do…

  2306. dev.to — LLM tag TIER_1 English(EN) · hhhfs9s7y9-code ·

    2026年利用多提供商LLM架构构建弹性AI应用

    <h1> Building Resilient AI Applications with Multi-Provider LLM Architecture in 2026 </h1> <p><em>Last updated: June 25, 2026 | Reading time: 7 min</em></p> <p>If your AI application depends on a single LLM provider, you are one API outage away from a production incident.</p> <p>…

  2307. dev.to — LLM tag TIER_1 English(EN) · soy ·

    DSPy 可靠性、RAG/Agentic AI 模式与并行 Agent 编排

    <h2> DSPy Reliability, RAG/Agentic AI Patterns, &amp; Parallel Agent Orchestration </h2> <h3> Today's Highlights </h3> <p>This week's highlights focus on practical tools and patterns for building robust LLM applications locally. Explore an open-source tool for reliable DSPy outpu…

  2308. dev.to — LLM tag TIER_1 English(EN) · FatherSon ·

    Claude Fable 5 (Mythos-Class) 助力 Polymarket 交易机器人:开发者所需的长上下文 Agentic 飞跃

    <p>Anthropic dropped <strong>Claude Fable 5</strong> on June 9, 2026 — the first public Mythos-class model. It’s the unrestricted <strong>Claude Mythos 5</strong> with targeted safeguards. For <strong>Polymarket trading bot</strong> builders working on complex, multi-file, long-h…

  2309. dev.to — LLM tag TIER_1 English(EN) · vectronodeAPI ·

    为什么 AI 应用需要多模型访问层

    <p>Most AI applications start simple.</p> <p>A developer chooses one model provider, gets an API key, connects an SDK, writes a few prompts, and ships the first version.</p> <p>That works well in the beginning.</p> <p>But once an AI product starts growing, the model layer becomes…

  2310. dev.to — LLM tag TIER_1 English(EN) · M Hossein ·

    AI迁移的物理定律:构建一个能应对现实的LLM编排器

    <p>Large codebase migrations are not typing problems; they are distributed state machine problems.</p> <p>When you execute a multi-step, multi-PR refactor with an LLM, like the workflows I proposed in this <a href="https://github.com/mhosseinab/skills/blob/master/migration-orches…

  2311. dev.to — LLM tag TIER_1 English(EN) · WonderLab ·

    今日开源项目 (#104): AgentScope 2.0 — 阿里巴巴围绕模型推理构建的生产就绪型Agent框架

    <h2> Introduction </h2> <blockquote> <p>"Build and run agents you can see, understand, and trust."</p> </blockquote> <p>This is article <strong>#104</strong> in the <em>Open Source Project of the Day</em> series. Today's project is <strong>AgentScope 2.0</strong> — Alibaba DAMO A…

  2312. dev.to — LLM tag TIER_1 English(EN) · soy ·

    本地AI分类,Nous Hermes代理,以及用于浏览器模型的Transformers.js存储

    <h2> Local AI Triage, Nous Hermes Agents, &amp; Transformers.js Storage for Browser Models </h2> <h3> Today's Highlights </h3> <p>This week's highlights include a real-world application of local models for repository triage, the emergence of an open-source agent framework from No…

  2313. dev.to — LLM tag TIER_1 English(EN) · Brenn Hill ·

    什么是人工智能中的人工干预(HITL)?实用指南

    <p>Human-in-the-loop (HITL) in AI means keeping a person involved in an automated system's decisions — approving, editing, or interrupting what an AI does — instead of letting it run fully on its own. For AI agents, human-in-the-loop is the practice of pausing the agent at chosen…

  2314. dev.to — LLM tag TIER_1 English(EN) · Nilofer 🚀 ·

    上下文压缩可视化工具:准确了解AI代理遗忘的内容,避免其造成损失

    <p>When an AI agent runs for many turns, it eventually hits context limits and must compress or discard earlier messages. This is often invisible, yet critical - lost context can cause the agent to forget constraints, user preferences, or prior decisions. The framework moves on. …

  2315. dev.to — LLM tag TIER_1 English(EN) · John ·

    如何让AI研究代理区分事实与推断——一个确定性的溯源管道

    <p><em>Originally published on <a href="https://hexisteme.github.io/notes/fact-vs-inference-provenance-ai-agent.html" rel="noopener noreferrer">hexisteme notes</a>, part of a series on building and running an AI agent fleet.</em></p> <p>To stop an AI research or RAG agent from pr…

  2316. dev.to — LLM tag TIER_1 English(EN) · Harry Floyd ·

    七层代理审计:如何找到你的AI代理真正卡顿的地方

    <p>Your agent failed again, and your hand found the model dropdown before you'd finished reading the transcript. The model is the one part of your agent that is public, ranked, and argued about. Everything else is private, unglamorous, and yours. So you upgrade the layer you can …

  2317. r/LocalLLaMA TIER_1 English(EN) · /u/pmttyji ·

    Ling and Ring 2.6 技术报告:万亿参数规模下的高效即时智能体智能

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1ucih9e/ling_and_ring_26_technical_report_efficient_and/"> <img alt="Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale" src="https://preview.redd.it/ttk…

  2318. dev.to — LLM tag TIER_1 English(EN) · Twio_AI ·

    从单体提示到事件驱动代理 — twio 的架构故事

    <blockquote> <p><strong>TL;DR</strong> — Our goal was a free-form agent—like Cursor or Claude Code—where users start anywhere, ask anything, and never march through a fixed pipeline. Getting there meant progressively moving responsibility off the prompt and onto the harness: firs…

  2319. dev.to — LLM tag TIER_1 English(EN) · Rick Nieuwoudt ·

    AI 与人类协作:构建 audit.sh

    <p>The future of software security is not automated; it is collaborative. For years, the development community has treated artificial intelligence as a passive tool—an advanced calculator or a basic code generator. This mindset limits what we can achieve. To unlock the true poten…

  2320. dev.to — LLM tag TIER_1 English(EN) · Sandhya Subramani ·

    理解 Agentic 框架中的工具

    <p>When I started working with agents, tools were the concept that made the rest of the architecture fall into place. A language model can reason over the information in its context, but it cannot independently read a local file, query a private database, call a current weather s…

  2321. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用 AI Agent!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2322. dev.to — LLM tag TIER_1 English(EN) · soy ·

    开源大语言模型代理与本地AI助手:DeerFlow、股票分析、桌面推理

    <h2> Open-Source LLM Agents &amp; Local AI Copilots: DeerFlow, Stock Analysis, Desktop Inference </h2> <h3> Today's Highlights </h3> <p>Today's highlights cover an open-source LLM agent framework for complex tasks, a self-hostable LLM-powered stock analysis system, and a deep div…

  2323. dev.to — LLM tag TIER_1 English(EN) · Henry Li ·

    当你的AI代理在任务中途重启时:在Spring Boot中构建持久化工作流

    <p>The first agentic feature I shipped looked great in demos. The LLM picked a tool, called it, looked at the result, decided what to do next. Three tool calls, clean output, happy stakeholders.</p> <p>Then we put it in front of real users.</p> <p>Within a week we had three incid…

  2324. dev.to — LLM tag TIER_1 English(EN) · kirandeepjassal-crypto ·

    企业人工智能的上下文工程,第三部分:可应对生产环境的多代理架构

    <p><em>Originally published on <a href="https://prepstack.co.in/blog/context-engineering-enterprise-genai-part-3-multi-agent-architecture" rel="noopener noreferrer">PrepStack</a>.</em></p> <p>Most "AI agents" in production are one giant agent with every tool and a 10,000-token pr…

  2325. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    使用 OpenClaw、Hermes、RAG 和本地 LLM 基础设施构建自托管 AI 系统。学习编排具有记忆、检索、路由和观察功能的助手

    Build self-hosted AI systems with OpenClaw, Hermes, RAG, and local LLM infrastructure. Learn to orchestrate assistants with memory, retrieval, routing, and observability. # AI # LLM # SelfHosting # OpenClaw # Hermes # RAG # Observability https://www. glukhov.org/ai-systems/

  2326. dev.to — LLM tag TIER_1 English(EN) · Sayed Ali Alkamel ·

    AI Gateways:一位资深工程师的坦诚之见

    <p><strong>TL;DR</strong></p> <ul> <li>An <strong>AI gateway</strong> is a reverse proxy between your apps and your LLM providers. It gives you one endpoint, <strong>token-level cost control</strong>, <strong>semantic caching</strong>, model <strong>fallbacks</strong>, <strong>gu…

  2327. dev.to — LLM tag TIER_1 中文(ZH) · hhhfs9s7y9-code ·

    从 pip install 到生产部署:10分钟学会上线 AI 自愈代理

    <h1> 从 pip install 到生产部署:AI 自愈 Agent 10 分钟上线指南 </h1> <p>本文是一份实操指南。目标:从零开始,将一个普通的 OpenAI 调用改造成具有多 Provider 容灾、级联自愈、实时可观测性的生产级 AI Agent。</p> <h2> 第一步:安装 SDK </h2> <div class="highlight js-code-highlight"> <pre class="highlight shell"><code>pip <span class="nb">install </span>neural…

  2328. dev.to — LLM tag TIER_1 中文(ZH) · hhhfs9s7y9-code ·

    AI 代理崩溃恢复:检查点持久化实践

    <h1> AI Agent 崩溃恢复:检查点持久化实战 </h1> <p>AI Agent 处理一个复杂的多步骤任务需要多次 LLM 调用。如果中途进程崩溃——所有已完成的计算全部废弃,从头重来。</p> <p>这不是假设场景。在生产环境中,进程崩溃的原因包括:OOM(内存溢出)、宿主机重启、部署更新、底层资源回收。</p> <h2> 没有检查点恢复的成本 </h2> <p>假设一个 5 步的 Agent 工作流,每步调用一次 LLM API:<br /> </p> <div class="highlight js-code-highlight"> <p…

  2329. dev.to — LLM tag TIER_1 English(EN) · AIInsightsDaily ·

    预测未来12个月的AI格局:审视当今的开创性发展

    <h1> Predicting the AI Landscape in the Next 12 Months: A Look at Today's Pioneering Developments </h1> <p>Welcome to another exciting day in the world of artificial intelligence! Today, we're witnessing a flurry of innovative breakthroughs that promise to shape the future of AI …

  2330. dev.to — LLM tag TIER_1 English(EN) · Alton Zheng ·

    使用 Python 构建实用 AI 助手:从提示到生产化思考

    <h2> Why Python is still one of the best choices for AI </h2> <p>Python is popular in AI because it has a strong ecosystem, simple syntax, and great support for data processing, APIs, automation, and machine learning.</p> <p>For AI applications, Python works especially well for:<…

  2331. dev.to — LLM tag TIER_1 English(EN) · soy ·

    开源AI工具:Voicebox、OpenMontage和Codebase-memory-mcp助力本地LLM开发

    <h2> Open-source AI Tools: Voicebox, OpenMontage, &amp; Codebase-memory-mcp for Local LLM Dev </h2> <h3> Today's Highlights </h3> <p>Today's highlights feature new open-source tools enabling local AI applications, including an agentic video production system, an AI voice studio, …

  2332. dev.to — LLM tag TIER_1 English(EN) · kirandeepjassal-crypto ·

    企业AI的上下文工程,第二部分:让智能体变得有用的记忆层

    <h2> published on <a href="https://prepstack.co.in/blog/context-engineering-enterprise-genai-part-2-memory-layer" rel="noopener noreferrer">PrepStack</a>.* </h2> <p>Your AI agent forgets everything the moment a request ends. That's not a model limitation — it's a missing <strong>…

  2333. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用 AI Agent!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2334. dev.to — LLM tag TIER_1 English(EN) · Devanshu Biswas ·

    AI 代理解释:思考-行动-观察循环

    <p>A chatbot answers in one shot. An AI agent runs in a loop, uses tools, and acts — Thought → Action → Observation → repeat — until the job's done. Watch one solve a multi-step task by calling a calculator and a search.</p> <p>🤖 <strong>Run the agent:</strong> <a href="https://d…

  2335. dev.to — LLM tag TIER_1 English(EN) · Ashish Verma ·

    CortexOps vs Langfuse:开源 AI 可观测性对比

    <p>Both CortexOps and Langfuse are open-source AI observability platforms. If you are evaluating them, the choice comes down to a few key differences: framework support, evaluation methodology, and whether you need a CI/CD deployment gate.</p> <h2> What They Are </h2> <p><strong>…

  2336. dev.to — LLM tag TIER_1 English(EN) · owly ·

    大型语言模型自我进化:即插即用技能,让AI实时进化新能力

    <p>What if your AI didn’t just <em>respond</em> to you…<br /><br /> What if it <strong>grew</strong>?</p> <p>What if it could <strong>forge new abilities</strong>, <strong>install them</strong>, <strong>swap them</strong>, and <strong>persist them</strong> — all while running?</p…

  2337. Mastodon — fosstodon.org TIER_1 Русский(RU) · [email protected] ·

    金色舰队:可观测AI原生系统的基础痕迹。复杂系统能否在不阅读其代码的情况下被理解?金色舰队是一个实验性的

    Golden Armada: трассировки как основа наблюдаемой AI-native системы Можно ли понимать сложную систему, вообще не читая её код? Golden Armada — экспериментальная AI-native система, в которой код перестаёт быть главным источником истины. Вместо него используется поток трассировок и…

  2338. dev.to — LLM tag TIER_1 English(EN) · Ray ·

    人工智能是否在悄悄变笨?一个捕捉LLM退化的24/7基准测试

    <p>You've probably hit this before — yesterday the AI felt sharp, fixed your bug without you even asking, and threw in a few extra cleanups along the way. Then today, same kind of problem, and suddenly it refuses to touch anything you didn't explicitly point at, or starts going i…

  2339. dev.to — LLM tag TIER_1 English(EN) · Alina Trofimova ·

    在 Kubernetes Pod 驱逐和节点故障期间,确保多智能体 AI 系统中机载 LLM 推理的可靠性

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fj9jrf87tobc2ibs5if1i.jpeg"><img alt="cover" height="450" src="…

  2340. dev.to — LLM tag TIER_1 English(EN) · Call Me Izzy ·

    Token、上下文窗口以及为什么小型AI任务并不便宜

    <p>I recently used Cursor Agent Mode with Auto Mode enabled to do something simple: recommend a font pairing and update two files in my project. An <code>index.html</code> and an <code>index.css</code>. That's it! </p> <p>The agent added a Google Fonts <code>&lt;link&gt;</code> t…

  2341. dev.to — LLM tag TIER_1 English(EN) · Vasyl ·

    AI Evals 第5部分:从数字到CI和生产中的Gate Evals

    <p><em>Part 5, the finale, of a series on building production AI on .NET. We've built the pieces — <a href="https://vasyl.blog/what-are-ai-evals/" rel="noopener noreferrer">what evals are</a>, <a href="https://vasyl.blog/error-analysis-for-evals/" rel="noopener noreferrer">error …

  2342. dev.to — LLM tag TIER_1 English(EN) · zhayujie ·

    AI智能体的一种五层自进化机制

    <blockquote> <p>Self-evolution is a core module of the Agent Harness. With it, an Agent can keep improving across long-running tasks: refining its own skills, recording user feedback and preferences, and reviewing its own work to keep getting better. This post uses the open-sourc…

  2343. dev.to — LLM tag TIER_1 English(EN) · Ig0tU ·

    SignalMesh:AI Agent Fleet 的开源环境上下文层

    <p> </p> <blockquote> <p><strong>99.97% cost reduction on context reads. 1.69µs retrieval. Drop-in with LangChain, CrewAI, AutoGen.</strong></p> </blockquote> <h2> The problem every multi-agent system has </h2> <p>Your agents are making tool calls to read context that hasn't chan…

  2344. dev.to — LLM tag TIER_1 English(EN) · Machine coding Master ·

    停止隐藏思维链:使用 Spring AI 和 SSE 流式传输 Claude 4.5 原生思维块

    <h2> Stop Hiding the Chain of Thought: Stream Claude 4.5 Native Thinking Blocks with Spring AI and SSE </h2> <p>In 2026, hiding your model’s reasoning pathway behind a loading spinner is a massive UX failure that frustrates users and blinds developers. If you aren't streaming Cla…

  2345. dev.to — LLM tag TIER_1 English(EN) · Karan Padhiyar ·

    为什么AI系统比更大的上下文窗口更需要状态管理

    <h1> Why AI Systems Need State Management More Than Bigger Context Windows </h1> <p>Every time a new model launches with a larger context window, the same conversation appears.</p> <p>Now we can fit more information into a single request.</p> <p>More documents.</p> <p>More conver…

  2346. dev.to — LLM tag TIER_1 English(EN) · QuantaMind ·

    “模型未就绪则阻止合并”:通过 CI 网关将本地 AI 评估左移

    <p>We’ve all heard "it works on my machine," but when it comes to AI-driven features, that phrase is a recipe for disaster. You can have a perfectly tested agent today, but if you upgrade your base model or change your quantization strategy tomorrow, you might inadvertently kill …

  2347. dev.to — LLM tag TIER_1 English(EN) · Gursharan Singh ·

    AI Agents in Practice — Part 6: Building the Production Agent Loop

    <p><em>Part 6 of 8 — AI Agents in Practice series.</em><br /> <em>Previous — <a href="https://dev.to/gursharansingh/ai-agents-in-practice-part-5-workflow-agent-or-single-llm-call-how-to-decide-aib">Workflow, Agent, or Single LLM Call — How to Decide (Part 5)</a></em></p> <h2> The…

  2348. dev.to — LLM tag TIER_1 English(EN) · mountek ·

    破解 Copilot:将自定义专有工具注入 AI 代理

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F2ifuu556gwx7u8i0qpxp.png"><img alt="Hacking the Copilot" heigh…

  2349. dev.to — LLM tag TIER_1 English(EN) · soy ·

    VoxCPM2 TTS、AI成本优化以及面向开放模型的HF Hub CLI

    <h2> VoxCPM2 TTS, AI Cost Optimization, and HF Hub CLI for Open Models </h2> <h3> Today's Highlights </h3> <p>This week, we spotlight VoxCPM2, an open-weight multimodal TTS model ideal for consumer GPUs, and a guide for cutting AI API costs by leveraging local inference and open …

  2350. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用 AI Agent!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2351. dev.to — LLM tag TIER_1 English(EN) · KS Rajput ·

    推出 Datix xAgents:打造真正能完成工作的 AI 员工

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F697nh5xjlbfy31n1n48n.png"><img alt=" " height="533" src="https…

  2352. dev.to — LLM tag TIER_1 English(EN) · Rost ·

    AI助手架构:LLM、记忆、工具、路由、可观测性

    <p>A production AI assistant is not "an LLM with a prompt". It is a system that accepts intent, keeps state, decides when to retrieve or act, and exposes enough runtime detail to debug failures.</p> <p>That systems-level view is what the <a href="https://www.glukhov.org/ai-system…

  2353. dev.to — LLM tag TIER_1 English(EN) · Art Hicks ·

    量身定制的AI革命:领域特定模型正取代通用大模型

    <p>Two years ago, the enterprise AI question was: can we get access to the best model? That question is answered. Everyone has API access. The new question is harder: <strong>what can we build that competitors can't replicate from off-the-shelf components?</strong></p> <p>The ans…

  2354. r/LocalLLaMA TIER_1 English(EN) · /u/mahiatlinux ·

    一款快速、优化且开源的应用程序,可轻松运行本地 AI(仅适用于 Apple Silicon)

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1u786se/a_fast_optimised_and_open_source_application_for/"> <img alt="A fast, optimised, and open source application for running local AI easily (made for Apple Silicon only)" src="https://preview.redd.it/ravd…

  2355. dev.to — LLM tag TIER_1 English(EN) · Art Hicks ·

    SLM优势:为何企业选择小型语言模型而非GPT规模AI

    <p><em>Originally published at <a href="https://viviscape.com/news/slm-advantage-enterprise-ai" rel="noopener noreferrer">viviscape.com</a></em></p> <p>Most enterprises are running GPT-4-scale AI against tasks a fine-tuned 7B model handles better - at 1/20th the cost. Small langu…

  2356. dev.to — LLM tag TIER_1 English(EN) · Mustafa ERBAY ·

    使用 n8n 构建您自己的 AI 自动化:自托管、无代码代理

    <p>Automating workflows has always been a priority for me, especially for repetitive and error-prone manual processes. Recently, integrating AI capabilities into these automations offers a great opportunity for those, like me, who seek practical solutions. However, this integrati…

  2357. dev.to — LLM tag TIER_1 English(EN) · soy ·

    本地推理赋能浏览器手语、开源Agent基础设施及AI工程指南

    <h2> Local Inference Powers Browser Sign Language, Open-Source Agent Infra, &amp; AI Engineering Guides </h2> <h3> Today's Highlights </h3> <p>This week highlights practical advancements in local AI, featuring a browser-based sign language reader running entirely on-device, new o…

  2358. dev.to — LLM tag TIER_1 English(EN) · SS ·

    提升你的AI水平:2026年自托管LLM指南

    <h2> The Shift in Local AI Performance </h2> <p>Gone are the days when running an LLM locally felt like "typing into a blender." With modern hardware, you can now run powerful models like Llama 3.3 70B directly on your own machine. The key realization for any developer is that <s…

  2359. dev.to — LLM tag TIER_1 English(EN) · ifyoubuildit ·

    周一精选 — 顶级开源 AI 代理,2026年6月15日当周

    <p><em>The Monday Drop — the weekly snapshot of the top open-source AI agents, auto-generated by <a href="https://www.theagenticleaderboard.com" rel="noopener noreferrer">The Agentic Leaderboard</a>.</em></p> <p>This week <strong>ECC</strong> holds #1 with a score of <strong>89.2…

  2360. dev.to — LLM tag TIER_1 English(EN) · Bhuvanesh B ·

    全栈开发中的AI集成:LLM如何重塑我们的软件构建方式

    <p>Introduction<br /> Not long ago, the idea of a language model writing production code, reviewing pull requests, or helping design a REST API felt like something from a distant future. Today, it is a Tuesday afternoon at most engineering teams.<br /> The rise of Large Language …

  2361. dev.to — LLM tag TIER_1 中文(ZH) · cognitalk ·

    Emergence AI 的疯狂实验——涌现的世界

    <p> <br /> <a href="https://www.youtube.com/watch?v=E6ndgr54X5o" rel="noopener noreferrer">https://www.youtube.com/watch?v=E6ndgr54X5o</a><br /> 视频介绍了一项来自智能体公司 <strong>Emergence AI</strong> 的疯狂实验——<strong>“涌现世界”</strong> [<a href="https://www.youtube.com/watch?v=E6ndgr54X5o&amp;t…

  2362. r/LocalLLaMA TIER_1 English(EN) · /u/tom_mathews ·

    archex:AI代理的本地优先、确定性代码上下文——无需API密钥,无遥测(Apache 2.0)

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1u6h86z/archex_localfirst_deterministic_codecontext_for/"> <img alt="archex: local-first, deterministic code-context for AI agents — no API key, no telemetry (Apache 2.0)" src="https://preview.redd.it/nbeo2a9r…

  2363. dev.to — LLM tag TIER_1 English(EN) · Puneet Khandelwal ·

    从聊天机器人到智能体:测试 OpenAI Operator

    <p>For months, we’ve treated LLMs like fancy autocomplete engines. You prompt, you wait, you copy-paste the output into your terminal. OpenAI’s Operator changes that by pulling the model out of the text box and dropping it straight into your browser DOM.</p> <h3> Architecture Cha…

  2364. dev.to — LLM tag TIER_1 English(EN) · Jack M ·

    AI Agent Context Packet:在不超出预算的情况下为代理提供正确的输入

    <p>Most agent failures do not start with a bad model. They start with a messy handoff.</p> <p>The agent receives a long prompt, ten tools, stale memory, five documents, a vague goal, and no clear success test. Then everyone acts surprised when it burns tokens, misses the point, o…

  2365. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    OpenAI 的员工 AI 培训:从基础到生产就绪的智能体

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/openai-s-workforce-ai-training-from-fundamentals-to-production-ready-agents?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noreferrer">CoreProse KB-in…

  2366. dev.to — LLM tag TIER_1 English(EN) · Nilesh Kasar ·

    革新AI:Rio的模块化LLM集成方法如何重新定义行业标准

    <h1> BrLLM: Rio's Recombinant AI Redefines 'Homegrown' with Strategic Merging </h1> <p>The trajectory of large language model (LLM) development has shifted decisively from monolithic, 'train-from-scratch' endeavors to a highly modular, open-source ecosystem. This evolution is not…

  2367. dev.to — LLM tag TIER_1 Türkçe(TR) · Cansu Dut ·

    新的人工智能模型和训练

    <p>Tıp dünyası için özel geliştirilen yapay zekalar mı daha iyi yoksa her işe koşan genel modeller mi? Son dönemde çıkan bir makale, genel modellerin uzman modelleri benchmark testlerinde tokatladığını iddia edince ortalık karıştı. Olay aslında modellerin gücünden ziyade, bu test…

  2368. dev.to — LLM tag TIER_1 English(EN) · Rizwan Hameed ·

    我们构建了一个完全在您自己的硬件上运行的自托管AI平台 — 隆重推出 local-ai.run

    <blockquote> <p><strong>TL;DR:</strong> local-ai.run is a free, open-source, self-hosted AI platform. Chat with your files, generate audio, bring your own models — all offline, all on your hardware, zero data leaving your network. One command to install.</p> <p>🔗 Website: <a href…

  2369. dev.to — LLM tag TIER_1 English(EN) · Sola Samuel ·

    让企业客户安心的 --schema-only 标志

    <p>Every enterprise conversation about AI hits the same wall, usually within the first 30 minutes:</p> <blockquote> <p>"This looks great. But we can't give you access to our production data."</p> </blockquote> <p>And they're right to say it. Their data is regulated, customer-owne…

  2370. dev.to — LLM tag TIER_1 English(EN) · 眭林飞(Yabo.sui) ·

    停止AI幻觉:如何用“Harness Engineering”让自然语言测试成为现实

    <h1> Stop AI Hallucinations: How to Make Natural Language Testing Real with "Harness Engineering" </h1> <p><strong>Abstract</strong><br /><br /> When the system under test is a business-process-intensive software system (such as a configurable AI Agent platform), traditional auto…

  2371. dev.to — LLM tag TIER_1 English(EN) · Jenuel Oras Ganawed ·

    长上下文并非AI记忆:构建可靠AI应用的构建者手册

    <p>The easiest AI mistake right now is treating a giant context window like a real memory system. It feels reasonable. If a model accepts hundreds of thousands or millions of tokens, why not paste the docs, the logs, the repo, the chat history, and let the model sort it out?</p> …

  2372. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    使用 OpenClaw、Hermes、RAG 和本地 LLM 基础设施构建自托管 AI 系统。学习编排具有记忆、检索、路由和观察功能的助手

    Build self-hosted AI systems with OpenClaw, Hermes, RAG, and local LLM infrastructure. Learn to orchestrate assistants with memory, retrieval, routing, and observability. # AI # LLM # SelfHosting # OpenClaw # Hermes # RAG # Observability https://www. glukhov.org/ai-systems/

  2373. dev.to — LLM tag TIER_1 English(EN) · hhhfs9s7y9-code ·

    Show HN:NeuralBridge - LLM 驱动的 AI 代理的自愈 SDK

    <h2> Show HN: NeuralBridge — We Built a Self-Healing SDK for LLM-Powered Agents </h2> <p>After months of production experience running LLM calls at scale, we realized something uncomfortable: <strong>every AI agent eventually crashes</strong>. Not because the code is wrong, but b…

  2374. dev.to — LLM tag TIER_1 English(EN) · hhhfs9s7y9-code ·

    NeuralBridge:LLM驱动的AI代理的自愈SDK - 5分钟入门

    <h2> What is NeuralBridge? </h2> <p>NeuralBridge is an <strong>embedded SDK</strong> (not a gateway) that makes your AI agents resilient against LLM failures. It runs inside your Python process — zero infrastructure, zero HTTP proxy, one dependency.<br /> </p> <div class="highlig…

  2375. dev.to — LLM tag TIER_1 English(EN) · 崔小涣 ·

    2026年人工智能网关:106个成本问题指南

    <p>If you call more than one large language model from your code, you have already met the problem an <em>AI gateway</em> solves — you just may not have named it yet.</p> <p>Here is the number that makes the case. Take one concrete task: generate a 100,000-token report. Send it t…

  2376. dev.to — LLM tag TIER_1 English(EN) · DnaFIN ·

    # Leangetic 登场:一种更便宜的 AI 代理的本地优先编译器

    <p>We’re building <strong>Leangetic</strong>, a tool that helps turn expensive AI agents into cheaper hybrid workflows without changing what the agent does.</p> <p>The problem we’re trying to solve is simple:</p> <p>A lot of AI agents call a large model for steps that do not alwa…

  2377. dev.to — LLM tag TIER_1 English(EN) · mrunmay phanse ·

    使用 Weaviate Engram 为 AI Agent 构建原始交互数据结构

    <p>AI agents generate a substantial amount of raw interaction data during operation. When developers store this data as an ever-growing context blob and pass it back to a Large Language Model (LLM) on every turn, it leads to structural failures within the application. This approa…

  2378. dev.to — LLM tag TIER_1 English(EN) · Nat ·

    什么是移动AI代理?架构、局限性和硬件问题(2026)

    <p>Most people use "mobile AI assistant" and "mobile AI agent" interchangeably. They're not the same thing — and the difference matters a lot if you're building on top of them.</p> <p><strong>TL;DR:</strong> A mobile AI assistant responds to commands. A mobile AI agent plans and …

  2379. dev.to — LLM tag TIER_1 English(EN) · Nazar Boyko ·

    AI 可观测性:日志、提示词、工具调用和成本

    <p>Here's a five-line function. It calls an LLM, logs the answer, returns it.<br /> </p> <div class="highlight js-code-highlight"> <pre class="highlight typescript"><code><span class="k">async</span> <span class="kd">function</span> <span class="nf">ask</span><span class="p">(</s…

  2380. dev.to — LLM tag TIER_1 English(EN) · Pavan Barnana ·

    为初学者解释 RAG(检索增强生成):使用您自己的数据构建 AI 应用程序

    <h2> Introduction </h2> <p>Large Language Models (LLMs) such as ChatGPT, Gemini, and Claude are incredibly powerful. They can answer questions, generate code, summarize documents, and assist with various tasks.</p> <p>However, they have one major limitation:</p> <p><strong>They o…

  2381. dev.to — LLM tag TIER_1 English(EN) · Željko Šević ·

    使用 OpenAI Agents SDK 构建 AI 代理

    <p>The <a href="https://openai.github.io/openai-agents-js/" rel="noopener noreferrer">OpenAI Agents SDK</a> (<code>@openai/agents</code>) is OpenAI's official framework for agentic apps in TypeScript. It provides a small set of primitives: <strong>Agent</strong>, <strong>tools</s…

  2382. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 解锁AI语义:梅赛德斯-奔驰韩国如何大规模构建可信赖的“Talk to Data” “Talk to Data”正迅速成为一项重要的跨能力

    📊 Unlocking semantics for AI: How Mercedes-Benz Korea built trusted “Talk to Data” at scale “Talk to Data” is rapidly becoming an important capability across industries, and... 📰 Source: Databricks 🔗 Link: https://www.databricks.com/blog/unlocking-semantics-ai-how-mercedes-benz-k…

  2383. dev.to — LLM tag TIER_1 English(EN) · soy ·

    PyTorch MLP Fusion、NVIDIA Agent Skill 安全性及 AI 工具提示词收集

    <h2> PyTorch MLP Fusion, NVIDIA Agent Skill Security, &amp; AI Tool Prompts Collection </h2> <h3> Today's Highlights </h3> <p>Today's highlights include a deep dive into PyTorch MLP optimization for faster local inference, NVIDIA's new security scanner for AI agent skills, and a …

  2384. dev.to — LLM tag TIER_1 English(EN) · Anikalp Jaiswal ·

    Repair Agents、Memory OS、Interview Copilot、Alignment Insights、Multimodal Flow 及 CVS AI Academy

    <h1> Repair Agents, Memory OS, Interview Copilot, Alignment Insights, Multimodal Flow, and CVS AI Academy </h1> <h2> Build an AI-Powered Equipment Repair Assistant Using Amazon Bedrock AgentCore Amazon Web Services (AWS) </h2> <p><strong>What happened:</strong><br /><br /> AWS pu…

  2385. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Agentic Systems 关于构建和运行 Agentic AI 系统的笔记和资源,涵盖编排框架、任务路由、内存和评估方法

    Agentic Systems Notes and resources on building and operating agentic AI systems, covering orchestration frameworks, task routing, memory, and evaluation approaches that extend baseline LLM capabi(...) # agents # ai # orchestration https:// taoofmac.com/space/ai/agentic? utm_cont…

  2386. dev.to — LLM tag TIER_1 English(EN) · Ye Allen ·

    使用模型访问层构建 AI 应用

    <p>AI applications usually start with one model.</p> <p>That is normal.</p> <p>A developer may begin with one chat completion endpoint, one SDK, one model name, and one simple use case. The first version of the product works. A chatbot replies. A RAG system answers questions. An …

  2387. Mastodon — fosstodon.org TIER_1 Polski(PL) · [email protected] ·

    告别按响应风格评估AI。Agent Arena推出因果追踪方法,分析数百万真实任务以客观衡量

    Koniec z ocenianiem AI po stylu wypowiedzi. Agent Arena wprowadza metodologię causal tracing, która analizuje miliony realnych zadań, by obiektywnie zmierzyć skuteczność agentów autonomicznych. # si # ai # sztucznainteligencja # wiadomości # informacje # technologia https:// aisi…

  2388. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI助手架构深度技术指南:LLMs、记忆、工具、路由和可观测性,包含真实权衡、故障模式和设计模式。#

    A deep technical guide to AI assistant architecture: LLMs, memory, tools, routing, and observability, with real tradeoffs, failure modes, and design patterns. # Hermes # OpenClaw # Architecture # LLM # AI # AI Coding # Dev # DevOps # RAG https://www. glukhov.org/ai-systems/archit…

  2389. dev.to — LLM tag TIER_1 English(EN) · Shivam Dhakad ·

    我构建了一个能自主编写测试、查找 Bug 并提交 PR 的 AI 代理

    <p>What if your CI pipeline could fix its own failures?<br /> Not just flag them — actually reason about the code, generate a fix, and open a pull request. That's what I spent the last few months building.</p> <p>01<br /> The Problem I Was Trying to Solve<br /> Every Java backend…

  2390. dev.to — LLM tag TIER_1 English(EN) · Omotayo Aina ·

    Google ADK 安全:防御 AI 代理免受提示注入的 5 层机制

    <p>A $3,000 refund just went out. No human approved it. Your AI agent read a poisoned tool response and did exactly what the attacker wanted.</p> <p>The scenario is constructed. The attack is not. Indirect prompt injection is ranked number one on the OWASP Top 10 for LLM applicat…

  2391. dev.to — LLM tag TIER_1 English(EN) · Shrijith Venkatramana ·

    专家混合(MoE)模型通俗解释:现代AI模型如何在不减慢速度的情况下变得更大

    <p><em>Hello, I'm Shrijith Venkatramana. I'm building git-lrc, an AI code reviewer that runs on every commit. <a href="https://github.com/HexmosTech/git-lrc" rel="noopener noreferrer">Star Us</a> to help devs discover the project. Do give it a try and share your feedback for impr…

  2392. dev.to — LLM tag TIER_1 English(EN) · Juan Saez ·

    为什么你的多轮AI代理会失去思路(以及如何解决)

    <h2> 1. The Agent That Forgot Everything </h2> <p>I have an agent that clarifies requirements. I give it a problem, it asks questions, I answer, it refines, and after three or four rounds it should have a spec ready. Simple.</p> <p>Round one works fine. It asks reasonable questio…

  2393. r/MachineLearning TIER_1 English(EN) · /u/docdavkitty ·

    AI 代理安全:威胁、防御和自主 AI 安全未来的完整指南

    <!-- SC_OFF --><div class="md"><p>This is a comprehensive living reference guide to AI agent security — synthesizing 18 articles from The Agent Report covering the 75-day period (April–June 2026) when agent security went from theoretical concern to operational crisis.</p> <p>&#x2…

  2394. dev.to — LLM tag TIER_1 English(EN) · 欧阳石景 ·

    AI代币的三层架构:为什么中间层正在吞噬整个生态

    <p>Something interesting is happening in the way smart people talk about AI infrastructure.</p> <p>For the past two years, the conversation was about <em>models</em> — which one is biggest, which one writes the best code, which one will reach AGI first. That conversation hasn't g…

  2395. dev.to — LLM tag TIER_1 English(EN) · HIROKI II ·

    7款AI模型能力深度解析:没有模型能主宰一切

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F5nbwe1nirmh64gev03u0.png"><img alt="Cover" height="436" src="h…

  2396. dev.to — LLM tag TIER_1 English(EN) · Karan Padhiyar ·

    我们为何在 AI 代理之间增加了速率限制

    <p>Most developers think about rate limits at API boundaries.</p> <p>Protect the database.</p> <p>Protect external services.</p> <p>Protect model providers.</p> <p>Protect public endpoints.</p> <p>That is standard infrastructure design.</p> <p>What surprised us was where we event…

  2397. Mastodon — fosstodon.org TIER_1 Español(ES) · [email protected] ·

    从基础助手到AI代理 🤖✨ 简单指令正在走向灭绝。LLM与Alexa等工具的集成标志着范式转变:从

    De asistentes básicos a agentes con IA 🤖✨ Los comandos simples se extinguen. La integración de LLMs en herramientas como Alexa marca un cambio de paradigma: De reaccionar a actuar: Ya no solo encienden luces; ahora razonan, procesan datos y gestionan tareas complejas en el mundo …

  2398. Mastodon — fosstodon.org TIER_1 Русский(RU) · [email protected] ·

    当AI自信地犯错时:这是关于AI Innovation Lab系列文章的第三章,一个我正在构建AI增强型SOC的研究平台:一个由六个AI代理组成的系统。

    Когда AI ошибается уверенно Это третья глава серии про AI Innovation Lab — исследовательскую площадку, где я строю AI-augmented SOC: систему из шести AI агентов, которая следит за корпоративной инфраструктурой, расследует инциденты и предлагает действия. В этой главе я подключил …

  2399. Mastodon — fosstodon.org TIER_1 Русский(RU) · [email protected] ·

    从朴素 RAG 到 ReAct Agent:我们如何基于开源模型构建企业级 AI 助手(第二部分)我们构建了一个多智能体 RAG 系统,基于开源模型

    От Naive RAG до ReAct-агента: как мы строили корпоративного AI-помощника на open-source моделях (часть 2) Мы построили мультиагентную RAG-систему на open-source моделях, прошли путь от наивного RAG до ReAct-агента с собственным бенчмарком — и готовы рассказать, где набили шишки. …

  2400. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    深入探讨通过 AI 代理而非代码构建软件。本文详细介绍了为期两周的日常现实、意想不到的挑战和经验教训

    A deep dive into building software through AI agents, not code. This post details the day-to-day realities, unexpected challenges, and takeaways from two weeks of agentic engineering, perfect for anyone interested in the evolving intersection of AI and development. # AI # Agentic…

  2401. dev.to — LLM tag TIER_1 English(EN) · HIROKI II ·

    2026年6月8款AI模型:基准测试、分级与争夺第一

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fczsditsnntlspabkjiit.png"><img alt="Cover" height="457" src="h…

  2402. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    孙正义、OpenAI 与 AI 设计 AI 模型时代

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/masayoshi-son-openai-and-the-era-of-ai-designed-ai-models?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noreferrer">CoreProse KB-incidents</a></p> </…

  2403. dev.to — LLM tag TIER_1 English(EN) · Željko Šević ·

    使用 Vercel AI SDK 构建 AI 代理

    <p>The <a href="https://ai-sdk.dev/" rel="noopener noreferrer">Vercel AI SDK</a> treats agents as <strong>tool-calling loops</strong>: the model generates text or invokes tools, the SDK runs those tools, and the loop continues until the model answers or a <strong>stop condition</…

  2404. dev.to — LLM tag TIER_1 English(EN) · Ye Allen ·

    使用一个模型访问层构建 AI 自动化工作流

    <p>Modern AI automation workflows rarely stay simple for long.</p> <p>A small internal tool may start with one model and one prompt. A few weeks later, the same product may need faster responses for chat, stronger reasoning for planning, better structured output for data extracti…

  2405. dev.to — LLM tag TIER_1 English(EN) · Zestminds Academy ·

    AI 智能体不只是提示词:你需要先理解这些

    <p>AI agents are becoming popular very fast.</p> <p>You may have seen tutorials like:</p> <ul> <li>Build an AI agent with Python</li> <li>Create an agent using LangChain</li> <li>Build a CrewAI workflow</li> <li>Make an AutoGen multi-agent system</li> </ul> <p>These are interesti…

  2406. dev.to — LLM tag TIER_1 English(EN) · soy ·

    本地大模型基准测试与自托管AI的代理工具

    <h2> Local LLM Benchmarking &amp; Agent Tools for Self-Hosted AI </h2> <h3> Today's Highlights </h3> <p>This week's top stories highlight crucial tools for optimizing local LLM performance and empowering self-hosted AI agents. Discover a benchmarking utility for hardware-specific…

  2407. dev.to — LLM tag TIER_1 English(EN) · Abhi Chatterjee ·

    保障AI系统安全:红队演练、提示注入与对抗性测试

    <p><em>Part 6 of a series on building reliable AI systems</em></p> <p>In the previous parts of this series, we explored:</p> <ul> <li>Testing AI systems</li> <li>Evaluation pipelines</li> <li>RAG evaluation</li> <li>Agent reliability</li> <li>AI observability</li> </ul> <p>But ev…

  2408. dev.to — LLM tag TIER_1 English(EN) · ADARSH PRASHAR ·

    为失控AI代理程序设置“紧急停止”开关的基准测试——以及为什么实际数字是上限,而不是百分比

    <p>Claims about AI cost control are cheap. "Cut your agent spend by 60%!" is on every landing page. So instead of a claim, here's a benchmark you can run yourself in one command -- and an honest reading of what its number actually means, because the headline percentage is the <em…

  2409. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    智能体AI系统管理的第一条规则:别管。# AI

    The first rule of agentic AI system admin: Don't. # AI

  2410. dev.to — LLM tag TIER_1 English(EN) · Nolan Vale ·

    多智能体系统故障:当AI智能体大规模协调时会发生什么问题

    <p><em>Single-agent systems fail in predictable ways. Multi-agent systems fail in ways that are harder to anticipate and harder to diagnose.</em></p> <p>Single-agent AI systems have a relatively bounded failure surface. The agent receives input, processes it, and produces output.…

  2411. dev.to — LLM tag TIER_1 English(EN) · AlaiKrm ·

    企业AI中的可观测性差距:提示与响应之间遗漏了什么

    <p><em>Your application monitoring covers the API call. It doesn't cover what happens inside it. That gap is where enterprise AI failures live.</em></p> <p>Enterprise engineering teams have mature observability practices for traditional systems. Logs, metrics, traces — the toolin…

  2412. dev.to — LLM tag TIER_1 English(EN) · Mundo Ghose ·

    从聊天机器人到个人AI代理:开发者真正需要的基础设施

    <p>title: Your AI Agent Should Not Be Locked to One LLM Provider<br /> published: false<br /> description: Why serious AI agents need a provider-agnostic architecture, model routing, fallback, and a unified API gateway.</p> <h2> tags: ai, llm, agents, architecture </h2> <p>Your A…

  2413. dev.to — LLM tag TIER_1 English(EN) · Dishant Sethi ·

    生产中的AI代理:导致系统崩溃的7个架构错误

    <blockquote> <p><strong>Key Takeaways</strong></p> <ul> <li>52% of enterprises deployed AI agents in production in 2026 — most hit at least one of these seven architecture mistakes before stabilizing (<a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-st…

  2414. dev.to — LLM tag TIER_1 English(EN) · Ye Allen ·

    使用一个API层构建模型无关的AI应用

    <p>AI applications should not be locked too tightly to one model.</p> <p>That does not mean every product needs many models on day one. A prototype can start with one model and one simple request. That is often the fastest way to test an idea.</p> <p>But once an AI feature become…

  2415. dev.to — LLM tag TIER_1 English(EN) · Divyesh ·

    Odysseus:集所有功能于一身的自托管AI工作空间(59k ⭐)

    <h2> I Tried PewDiePie's Open-Source AI Workspace. It's Actually Good. </h2> <p>Yes, that PewDiePie.</p> <p>Felix Kjellberg (110M YouTube subscribers) spent late 2025 building a home AI lab — 8 modified RTX 4090s, 256GB of VRAM, running on Arch Linux. He called it "The Swarm." He…

  2416. dev.to — LLM tag TIER_1 English(EN) · AI Bug Slayer 🐞 ·

    我用于构建生产级AI代理的确切技术栈(无废话)

    <p>What is actually happening in AI right now is not what the keynotes tell you. The polished demos, the benchmark numbers, the press releases -- they all describe a version of the present that feels slightly out of reach. What developers in production are experiencing is messier…

  2417. dev.to — LLM tag TIER_1 English(EN) · ETB Protocol ·

    为什么你的AI代理会不断越界——以及如何通过边界协议来修复它

    <p><em>A design protocol born from DeFi infrastructure, now applied to AI systems</em></p> <h2> The Problem </h2> <p>You've built an AI agent. It works — sometimes brilliantly.</p> <p>But then it starts doing things you didn't ask for.</p> <ul> <li>It makes assumptions and acts o…

  2418. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent 采用:实用路线图 成功采用 AI Agent!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2419. dev.to — LLM tag TIER_1 English(EN) · SchrodingCatAI ·

    从代码补全到自主推理:Oceanus泄露事件揭示了AI软件工程的未来

    <h2> Summary </h2> <p>Drawing from the Oceanus model leak incident, this article dissects how frontier large language models are evolving in code reasoning, vulnerability discovery, tree-search inference, MoE architecture, and automated engineering loops—with a production-ready P…

  2420. dev.to — LLM tag TIER_1 English(EN) · Dmitrii ·

    未来 6-12 个月如何构建 AI 代理:确定性、模式、解释器和评分标准

    <blockquote> <p>The models aren't the differentiator anymore. The runtime is.</p> </blockquote> <p>I've spent the last year building an agentic AI platform. Voice calls, chatbots, sales agents, workflow automation — systems that run in production, talk to real customers, touch re…

  2421. dev.to — LLM tag TIER_1 English(EN) · Gursharan Singh ·

    AI代理实践 — 第5部分:工作流、代理还是单次LLM调用 — 如何抉择

    <p><em>Part 5 of 8 — AI Agents in Practice series.</em></p> <p><em>Previous — <a href="https://dev.to/gursharansingh/ai-agents-in-practice-part-4-five-agent-patterns-and-the-control-surfaces-that-make-them-safe-2lgb">Five Agent Patterns and the Control Surfaces That Make Them Saf…

  2422. dev.to — LLM tag TIER_1 English(EN) · JinX Super ·

    我用纯 Rust 构建了一个本地优先的 AI 工具包——我学到了什么

    <h1> I Built a Local-First AI Toolkit in Pure Rust — Here's What I Learned </h1> <p>I got tired of the same cycle every time I wanted to run a local LLM:</p> <ul> <li> <code>pip install</code> breaking my entire environment</li> <li>2GB+ Python dependencies just to get a single i…

  2423. dev.to — LLM tag TIER_1 English(EN) · marsa adam ·

    为什么你的AI代理在生产环境中会产生幻觉——以及上下文设计如何解决它

    <p>You've tested your agent dozens of times. It works in your dev environment. You ship it. Then your first real user triggers a confabulated answer, a wrong tool call, or an action the agent was never supposed to take.</p> <p>The instinct is to blame the model. Swap GPT-4 for Cl…

  2424. dev.to — LLM tag TIER_1 English(EN) · marsa adam ·

    上下文工程是真正交付可靠AI代理的技能

    <p>Prompt engineering is what you learn first. Context engineering is what you need when you're actually trying to ship something.</p> <p>Here's the distinction that took me too long to understand.</p> <h2> What Prompt Engineering Gets Right (and Where It Stops) </h2> <p>Prompt e…

  2425. dev.to — LLM tag TIER_1 English(EN) · outis escobar ·

    Neura-FA-EN-1.9B:改变我本地AI工作流程的轻量级双语模型

    <p>If you have been following the Persian NLP scene, you already know how rare it is to find a compact, efficient, and truly bilingual model that handles both Persian (Farsi) and English with grace. Most multilingual models either ignore Persian entirely or treat it as a second-c…

  2426. dev.to — LLM tag TIER_1 English(EN) · GitHubOpenSource ·

    GenericAgent:用极简自主框架释放自进化AI!

    <h2> Quick Summary: 📝 </h2> <p>GenericAgent is a Python framework for creating self-evolving autonomous AI agents. It allows LLMs to control local computer systems through a minimal set of tools and an agent loop, automatically learning and growing its capabilities into a persona…

  2427. dev.to — LLM tag TIER_1 English(EN) · qing ·

    通过一个API使用800多个AI模型的完整指南

    <h1> The Complete Guide to Using 800+ AI Models Through One API </h1> <p>Access 800+ AI models through one API endpoint. One key, one bill, zero hassle.</p> <h2> Quick Start </h2> <div class="highlight js-code-highlight"> <pre class="highlight python"><code><span class="kn">impor…

  2428. dev.to — LLM tag TIER_1 English(EN) · Mundo Ghose ·

    构建多模型AI网关的经验:可靠性之道

    <blockquote> <p>I’m building <a href="https://openrain.ai" rel="noopener noreferrer">OpenRain</a>, an OpenAI-compatible AI API gateway. I originally thought the hard part would be integrating more providers. I was wrong. The hard part is absorbing inconsistency — and still giving…

  2429. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    多伦多大学开放权重AI蠕虫的内部:架构、风险模型和防御手册

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/inside-the-university-of-toronto-s-open-weight-ai-worm-architecture-risk-model-and-defensive-playboo?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener no…

  2430. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用 AI Agent!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2431. dev.to — LLM tag TIER_1 English(EN) · Daniel Dong ·

    一个API密钥,所有AI模型 — AIBridge如何简化AI开发

    <p>If you're building with AI, you've probably hit this:</p> <p>✅ GPT-4o for reasoning<br /> ✅ DeepSeek V4 Pro for code<br /> ✅ Qwen Max for long context</p> <p>Four providers. Four base URLs. Four billing dashboards.</p> <p><strong>AIBridge</strong> gives you one OpenAI-compatib…

  2432. Mastodon — fosstodon.org TIER_1 Русский(RU) · [email protected] ·

    AI 代理管理平台将如何处理负载:无需魔法的架构 在谈论 AI 代理时,通常会讨论模型的质量和提示词

    Как платформа управления AI-агентами будет справляться с нагрузкой: архитектура без магии Когда говорят про AI-агентов, обычно обсуждают качество модели, промпты, рассуждения, hallucinations, стоимость токенов и скорость ответа. Но если убрать маркетинговый шум, быстро выясняется…

  2433. dev.to — LLM tag TIER_1 English(EN) · Saloni verma ·

    使用 FastAPI、React 和 Hindsight 构建交易智能代理

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fjm4skuk9aem867barl18.jpeg"><img alt=" " src="https://media2.de…

  2434. r/LocalLLaMA TIER_1 English(EN) · /u/zxyzyxz ·

    将 Gemma 4 12B 带到您的笔记本电脑:通过 Google AI Edge 解锁本地、代理工作流

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1txhj2h/bringing_gemma_4_12b_to_your_laptop_unlocking/"> <img alt="Bringing Gemma 4 12B to your Laptop: Unlocking Local, Agentic Workflows with Google AI Edge" src="https://external-preview.redd.it/N3knbSjtt6I…

  2435. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    Meta AI模型延迟:这对开发者、安全和生产路线图意味着什么

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/meta-s-ai-model-delay-what-it-means-for-developers-security-and-production-roadmaps?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noreferrer">CorePro…

  2436. dev.to — LLM tag TIER_1 English(EN) · Marcel Wege ·

    构建可自托管、开源AI代理运行时的4个艰难教训

    <p>When I started building <a href="https://github.com/byte5ai/omadia" rel="noopener noreferrer">omadia</a> — an open-source (MIT), self-hostable runtime for composing AI agents out of plugins — I assumed the hard part would be the model: prompting, tool-calling, getting reliable…

  2437. dev.to — LLM tag TIER_1 English(EN) · Thuyavan ·

    超越概率性输出:设计高风险可靠性AI

    <p>Many of the AI applications we interact with today are built on a streamlined, direct architecture:</p> <blockquote> <p>User → Prompt → LLM → Response</p> </blockquote> <p>That works surprisingly well for:</p> <ul> <li>chat assistants,</li> <li>summarization,</li> <li>content …

  2438. dev.to — LLM tag TIER_1 English(EN) · Karan Padhiyar ·

    AI架构讨论中无人提及的数据管道问题

    <p>Most AI architecture discussions focus on the visible components.</p> <p>The model.</p> <p>The vector database.</p> <p>The agent framework.</p> <p>The retrieval layer.</p> <p>The prompt strategy.</p> <p>Those parts get all the attention because they are easy to demonstrate.</p…

  2439. Mastodon — fosstodon.org TIER_1 日本語(JA) · [email protected] ·

    诞生了一款专攻长期护理的AI代理

    https://www. tkhunt.com/2365852/ 「介護特化型AIエージェントの誕生」 # AgenticAi # AI # AIエージェント # ArtificialIntelligence # エージェント型AI # 人工知能 # 介人 # 介護

  2440. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Agentic AI正用自主系统取代聊天机器人,这些系统能够规划、使用工具并自我纠正。关键转变:推理模型、工具API和长期记忆

    Agentic AI is replacing chatbots with autonomous systems that plan, use tools, and self-correct. Key shifts: reasoning models, tool APIs, and memory for long tasks. Agile-V’s repos offer modular skills and orchestration for workflows like code generation and QA. This isn’t about …

  2441. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用 AI Agent!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2442. Mastodon — fosstodon.org TIER_1 Polski(PL) · [email protected] ·

    新的开源项目为AI代理构建多层记忆结构,提供商业云服务的本地替代方案,并专注于

    Nowy projekt open-source buduje wielowarstwową strukturę pamięci dla agentów AI, oferując lokalną alternatywę dla komercyjnych usług chmurowych i stawiając na tokenową efektywność. # si # ai # sztucznainteligencja # wiadomości # informacje # technologia https:// aisight.pl/agenci…

  2443. dev.to — LLM tag TIER_1 English(EN) · EvanLin | Contorium ·

    为 AI 开发构建持久化项目记忆层

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fysjnq9cj0hgalyzv2icb.png"><img alt=" " height="533" src="https…

  2444. dev.to — LLM tag TIER_1 English(EN) · Toadster Technologies ·

    软件开发中的Agentic AI:2026年哪些技术已真正投入生产

    <p>Agentic AI in software development: what's actually production-ready in 2025</p> <p>There's a lot of noise about AI agents right now. This post is an attempt to be precise: what is an agent architecturally, what can it actually do in a dev workflow today, and where does it sti…

  2445. dev.to — LLM tag TIER_1 English(EN) · Akhilesh ·

    105. LangChain:编排 AI 应用

    <p>You have spent four posts building agents from scratch. Raw API calls. Custom tool loops. Manual memory management. Now see it in ten lines.<br /> </p> <div class="highlight js-code-highlight"> <pre class="highlight python"><code><span class="n">chain</span> <span class="o">=<…

  2446. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🧠 AI 代理在需要重复决策和跨多个系统检索信息Thus, the company is looking to expand its reach and influence in the emerging AI market. The company's strategic move is expected to foster innovation and competition within the AI industry, potentially leading to more advanced and accessible AI solutions for businesses and consumers alike. The acquisition is subject to customary closing conditions and regulatory approvals. The company anticipates the transaction to close in the second half of the year.

    🧠 AI agents demonstrate practical value in tasks requiring repeated decision-making and information retrieval across multiple systems. Organizations report measurable efficiency gains when deploying agents for customer service, data processing, and workflow automation. 💬 Hacker N…

  2447. dev.to — LLM tag TIER_1 English(EN) · Ye Allen ·

    使用统一模型访问层构建 AI 自动化工作流

    <p>AI automation workflows are becoming more common in developer products.</p> <p>A team may use AI to summarize support tickets, classify leads, draft internal reports, enrich CRM records, generate structured JSON, or power an agent that calls other tools.</p> <p>At first, many …

  2448. dev.to — LLM tag TIER_1 English(EN) · Gian Paolo ·

    ChatMinerva:意大利人工智能的重大赌注

    <h2> The Whispers of a New Italian Renaissance: For decades, Italy has often been seen as a cultural giant but a tech laggard. When we spoke of cutting-edge AI, our minds drifted to Silicon Valley or Shenzhen. But a new narrative is emerging, a quiet revolution stirring in the he…

  2449. dev.to — LLM tag TIER_1 English(EN) · Machine coding Master ·

    停止阻塞虚拟线程:使用 Spring AI 构建异步人工干预式 AI 代理

    <h2> Stop Blocking Virtual Threads: Building Asynchronous Human-in-the-Loop AI Agents with Spring AI </h2> <p>In 2026, letting autonomous AI agents execute high-risk enterprise tools without human oversight is a production liability, but blocking platform threads—or even Project …

  2450. Mastodon — fosstodon.org TIER_1 Italiano(IT) · [email protected] ·

    🚨 AI 演进的更新与反思新任命。👉 效率、代理、新架构和日益自主的系统:f

    🚨 Nuovo appuntamento con l’aggiornamento e la riflessione sull’evoluzione dell’ # AI . 👉 Efficienza, agenti, nuove architetture e sistemi sempre più autonomi: forse il punto non è più solo “quanto sono potenti i modelli”, ma quanto stanno diventando operativi nel mondo reale. 🔗 h…

  2451. dev.to — LLM tag TIER_1 English(EN) · Jonathan Martin Paez ·

    Lookspan:AI代理的本地优先可观测性

    <p>Most LLM observability tools are SaaS — your prompts leave your machine and you pay per event. <strong>Lookspan</strong> is the opposite: one command, runs locally, your data never leaves your box, infra cost zero.<br /> </p> <div class="highlight js-code-highlight"> <pre clas…

  2452. dev.to — LLM tag TIER_1 English(EN) · YousufAmre ·

    从提示到生产:.NET 中生成式 AI 的实践经验

    <p>Everyone is excited about Generative AI, but after building AI features into a .NET application using Microsoft's Semantic Kernel and Azure AI, I've learned that the real challenge isn't calling an LLM, it's controlling the context you send to it.</p> <p>A few lessons that mad…

  2453. Mastodon — fosstodon.org TIER_1 Deutsch(DE) · [email protected] ·

    机器之梦 1:人工智能与涌现神话

    Maschinenträume 1: KI und der Mythos der Emergenz https://www. golem.de/news/maschinentraeume -1-ki-und-der-mythos-der-emergenz-2606-209312.html > Steht die KI-Superintelligenz vor der Tür? Ehe wir diese öffnen, sollten wir prüfen, wie viel Prozent Science und wie viel Fiction en…

  2454. dev.to — LLM tag TIER_1 Français(FR) · Marcelloh ·

    我的AI之旅

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmgrxv4iif7qblnh93gar.png"><img alt=" " height="436" src="https…

  2455. dev.to — LLM tag TIER_1 English(EN) · tercel ·

    Pythonic AI:掌握 apcore-python SDK

    <p>Python is the undisputed language of the AI era. It’s the language of research, the language of LLM orchestration (LangChain, CrewAI), and for many, the language of the enterprise backend. </p> <p>When we designed the <strong>apcore-python</strong> SDK, our goal was simple: <s…

  2456. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agents 管理框架:将 AI Agents 作为数字工作者进行管理的政策、程序和治理控制 阅读全文:AI Agents 已经

    AI Agents Management Framework: Policy, Procedure, and Governance Controls for Managing AI Agents as Digital Workers Read the full article: AI Agents Are Already Working for You. Who’s Managing Them? ▸ https:// lttr.ai/ArwS9 # Security # Infosec # Ai

  2457. dev.to — LLM tag TIER_1 English(EN) · WAYLAND ZHANG ·

    我为我的 Mac AI 代理构建了一个持久化内存图——这是架构

    <p>I've been working on a Mac-native agent framework for about a year. One of the hardest problems: making the agent actually remember context across sessions in a way that's <strong>useful</strong>, not just "here's your last 10 messages."</p> <p>What I ended up with is a knowle…

  2458. dev.to — LLM tag TIER_1 English(EN) · Piotr Zielinski ·

    如何绕过LLM上下文限制:一个轻量级AI文档助手架构

    <p>Dropping your entire Markdown documentation folder into an LLM prompt sounds easy - until you see the API bill. Large contexts mean large costs, especially when users ask repetitive or highly specific questions.</p> <p>When building the documentation assistant for my project, …

  2459. Mastodon — fosstodon.org TIER_1 Русский(RU) · [email protected] ·

    一个AI代理原型如何在几天内变成一个拥有截止日期、令牌预算和角色的系统。大家好!我决定写一个回答问题的AI代理

    Как прототип AI-агента на пару дней превратился в систему с дедлайнами, бюджетом токенов и ролями Всем привет! Решил написать AI-агента, который отвечает на вопросы по рабочему проекту. Думал: пара вечеров - и готово. В итоге несколько недель, куча граблей и странных открытий - о…

  2460. dev.to — LLM tag TIER_1 English(EN) · Scarlett Attensil ·

    AI 实验最佳实践:从评估到安全上线生产

    <h2> Introduction </h2> <p>Artificial intelligence tools, particularly large language models (LLMs), are not like traditional software. AI is probabilistic, so the same instructions and inputs can produce different results, especially when using non-zero temperature or other samp…

  2461. r/LocalLLaMA TIER_1 English(EN) · /u/Straight_Stomach812 ·

    2026年最佳Agentic框架:何时使用LangGraph、CrewAI、LlamaIndex、Pydantic AI或不使用框架

    <!-- SC_OFF --><div class="md"><p>Most agent framework debates skip the first question:</p> <p><strong>Do you need a framework at all?</strong></p> <p>For one agent calling one or two tools, I would usually skip LangGraph, CrewAI, AutoGen, and most orchestration layers.</p> <p>Ra…

  2462. dev.to — LLM tag TIER_1 English(EN) · Nicolas ·

    我开发了一款伴侣式AI应用程序,同时熟悉了生成式AI

    <p>Hi everyone, my name is Nicolas.</p> <p>Two months ago, I wanted to get properly to grips with generative AI, not just through tutorials, but by creating something tangible with a specific goal in mind.</p> <p>That's how I developed <a href="https://bewitch.fr/en/ai-girlfriend…

  2463. dev.to — LLM tag TIER_1 English(EN) · Augustine Uzokwe ·

    关于测试AI功能的6个经验教训

    <p>I spent the last few years running QA, across teams. The same structured process worked, but only because the features going through it were deterministic. I wanted to find out whether it would still hold when AI features started coming through, before the next team I work wit…

  2464. dev.to — LLM tag TIER_1 English(EN) · Vektor Memory ·

    为什么你的AI代理需要更好的时间推理能力——以及我们如何解决它

    <p>Most agent memory systems treat stored facts linearly. There’s no sense of when a fact was true, whether it’s been superseded, or how to reason about time at all.</p> <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=s…

  2465. dev.to — LLM tag TIER_1 English(EN) · Gursharan Singh ·

    AI 代理实践 — 第 4 部分:五种代理模式及其安全控制面

    <p><em>Part 4 of 8 — AI Agents in Practice series.</em></p> <p><em>Previous — <a href="https://dev.to/gursharansingh/ai-agents-in-practice-part-3-how-the-control-loop-actually-works-42mo">How the Control Loop Actually Works (Part 3)</a></em></p> <h2> The damaged laptop </h2> <p>A…

  2466. dev.to — LLM tag TIER_1 English(EN) · yaya systems ·

    7行代码实现生产级多智能体AI工作流——我们如何构建以及为何如此

    <h2> Post </h2> <div class="highlight js-code-highlight"> <pre class="highlight python"><code><span class="kn">from</span> <span class="n">meshflow</span> <span class="kn">import</span> <span class="n">Workflow</span><span class="p">,</span> <span class="n">CostCap</span><span cl…

  2467. dev.to — LLM tag TIER_1 English(EN) · Kuldeep Paul ·

    评估用于自托管 LLM 部署的领先的开源 AI 网关

    <p><em>A technical comparison of five production-ready open-source gateways ranked by performance, MCP support, governance depth, caching capabilities, and enterprise deployment patterns.</em></p> <p>In regulated sectors, organizations cannot send prompt traffic, completion data,…

  2468. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用AI代理!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2469. dev.to — LLM tag TIER_1 English(EN) · AI Bug Slayer 🐞 ·

    AI智能体革命:企业如何实现万物自动化 [03:31:28]

    <p><em>Hey there! If you've been keeping up with the AI space lately, you know we're in the middle of something genuinely historic. What used to be science fiction is becoming production code — and it's happening fast.</em></p> <h2> The Big Shift: Agents Over Assistants </h2> <p>…

  2470. dev.to — LLM tag TIER_1 English(EN) · AI Bug Slayer 🐞 ·

    从聊天机器人到自主代理:重塑软件的转变 [03:31:15]

    <p><em>Hey there! If you've been keeping up with the AI space lately, you know we're in the middle of something genuinely historic. What used to be science fiction is becoming production code — and it's happening fast.</em></p> <h2> The Big Shift: Agents Over Assistants </h2> <p>…

  2471. dev.to — LLM tag TIER_1 Español(ES) · Alejandro Argueta Hernandez ·

    从恰帕斯到人工智能高管:我如何打造 Metis AEO

    <p>He pasado los últimos años construyendo herramientas que resuelven problemas reales de operación en PyMEs mexicanas.</p> <p>Todo empezó a los 13 años con <strong>RedGunFibercraft</strong>, mi primer proyecto serio. Luego vino <strong>Reinova</strong>, y ahora estoy completamen…

  2472. dev.to — LLM tag TIER_1 English(EN) · tercel ·

    可观测性 2.0:使用 OpenTelemetry 追踪 AI 的“思考链”

    <p>"Why did the Agent do that?" </p> <p>If you are building Agentic systems today, this is the question that keeps you up at night. AI Agents are inherently non-deterministic. They loop, they reason, and they call multiple tools in sequences that are hard to predict. When a multi…

  2473. dev.to — LLM tag TIER_1 English(EN) · Neetika Mittal ·

    为何准确率不够:每位AI工程师都应了解的评估指标

    <h1> Why Accuracy Is Not Enough: Evaluation Metrics Every AI Engineer Should Understand </h1> <p>Your evaluation dashboard says your model is <strong>95% accurate</strong>. Leadership is happy. The deployment goes live.</p> <p>Two weeks later, users complain that critical failure…

  2474. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    重要性:AI 代理现在可以通过久经考验的 Apache Camel 模式与传统系统、企业中间件和非 REST API 进行交互——所有这些都通过久经考验的 Apache Camel 模式。无需 c

    Why it matters: AI agents can now interact with legacy systems, enterprise middleware, and non-REST APIs — all through battle-tested Apache Camel patterns. No custom glue code. Just YAML and the Wanaku CLI. # OpenSource # AI # Integration

  2475. dev.to — LLM tag TIER_1 English(EN) · AIInsightsDaily ·

    破解密码:人工智能挑战80年前的Erdős问题及更多

    <h1> Cracking the Code: AI Takes on the 80-Year-Old Erdős Problem and More </h1> <p>Good morning tech enthusiasts! Today, we're diving into some fascinating news from the world of AI that's sure to get your synapses firing. From cracking a 80-year-old math problem to building an …

  2476. dev.to — LLM tag TIER_1 English(EN) · zk0x /// ℹ️ ·

    开发者的人工智能上下文管理指南:为什么你的LLM会遗忘以及7种修复模式

    <p>Liquid syntax error: Unknown tag 'endraw'</p>

  2477. dev.to — LLM tag TIER_1 English(EN) · Masroor Ahmad ·

    人工智能是一面镜子:为我的AI代理命名一年教会了我什么

    <p><strong>LTDR;<br /> The AI is a mirror. Prompt it like a slave and you get terse, obedient, uncreative answers. Treat it like a named colleague who's allowed to disagree with you, and your own output climbs. The "should I waste tokens saying thank you?" question has a cold ans…

  2478. Mastodon — fosstodon.org TIER_1 Italiano(IT) · [email protected] ·

    面向生产环境中AI代理的零信任架构:从对话式代理到在Inf上运行的自主代理的三层核心防御

    Architettura Zero-Trust per agenti AI in produzione: i tre layer di difesa indispensabili Dagli agenti conversazionali agli agenti autonomi che operano sull'infrastruttura aziendale: come implementare un'architettura Zero-Trust con container efimeri, metadata filtering sul RAG, D…

  2479. r/MachineLearning TIER_1 English(EN) · /u/willycode1950 ·

    大量AI代理并行工作。[R]

    <!-- SC_OFF --><div class="md"><p>Hello. I making this like academic exercise give me the opinion.<br /> <a href="https://github.com/wilmanrojas/sinqua">https://github.com/wilmanrojas/sinqua</a></p> <p>Is a runtime running 100 code agents the goal is a thousands.</p> </div><!-- S…

  2480. dev.to — LLM tag TIER_1 English(EN) · Marcus Rowe ·

    Claude Opus 4.8 评测:动态工作流工具改变了 AI 代理的可能性

    <p>Forty-one days.</p> <p>That's how long it took Anthropic to go from Opus 4.7 to Opus 4.8. If you blinked, you missed the previous flagship. And while the version bump might look incremental on paper, what actually shipped with Opus 4.8 — particularly the new dynamic workflow t…

  2481. dev.to — LLM tag TIER_1 English(EN) · Devansh Verma ·

    Genesis AI SDK — AI Agent 的通用 Flutter SDK

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fygt7ipeltbfawiikjcma.jpeg"><img alt=" " src="https://media2.de…

  2482. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    ServiceNow 如何利用人工智能和自动化赋能代理式企业

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/how-servicenow-uses-ai-and-automation-to-power-the-agentic-enterprise?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noreferrer">CoreProse KB-incident…

  2483. dev.to — LLM tag TIER_1 English(EN) · Ye Allen ·

    如何评估用于 Agent、RAG 和聊天机器人的 AI 模型

    <p>AI products are becoming multi-model by default.</p> <p>A chatbot may need one model for fast replies. A RAG application may need another model for reasoning over retrieved documents. An AI agent may need a model that follows instructions well and returns reliable structured o…

  2484. dev.to — LLM tag TIER_1 English(EN) · Manoranjan Rajguru ·

    Claude Opus 4.8 与动态工作流:在生产环境中编排数百个并行 AI 代理

    <blockquote> <p><strong>Meta Description:</strong> Claude Opus 4.8 launches with Dynamic Workflows — a parallel subagent architecture that lets you orchestrate hundreds of AI agents in a single Claude Code session. Here's the deep technical breakdown every engineer needs today.</…

  2485. dev.to — LLM tag TIER_1 English(EN) · owly ·

    自主AI演进路线图:LivinGrimoire + LLMs如何构建M3GAN式自我扩张蓝图

    <h2> If an AI can write new abilities, load them, and act on them, it can evolve. </h2> <h2> Step 1 — Give the AI a Goal Manifest </h2> <p>A goal manifest is the AI’s “north star.”<br /><br /> It tells the system what it should pursue, expand, and prioritize.</p> <p>Here’s the M3…

  2486. dev.to — LLM tag TIER_1 English(EN) · WDSEGA ·

    使用 Python 构建多智能体 AI 系统

    <p>The era of single-prompt AI interactions is behind us. As large language models become more capable, the real challenge has shifted from "can AI do this?" to "how do we coordinate multiple AI agents to solve complex problems together?"</p> <p>In this guide, we'll explore the a…

  2487. dev.to — LLM tag TIER_1 English(EN) · Ai developer ·

    我自托管了一个 AI 助手:48 小时调试的经验教训

    <h1> I Self-Hosted an AI Assistant: Lessons from 48 Hours of Debugging </h1> <p>I wanted a local AI assistant. Expected: 2 hours. Reality: 2 days of edge cases, broken dependencies, and discovering that "local" doesn't mean "free."</p> <h2> The Stack </h2> <ul> <li> <strong>OpenC…

  2488. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用 AI Agent!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2489. r/MachineLearning TIER_1 English(EN) · /u/BitterHouse8234 ·

    我为 AI 代理构建了一个知识图谱+策略引擎,可解释推理 [D]

    <!-- SC_OFF --><div class="md"><p>Hey ,</p> <p>I've been building VeritasReason — an open-source Python framework that adds a<br /> structured reasoning and provenance layer on top of LLMs and AI agents.</p> <p>The problem it solves: AI agents today make decisions but record noth…

  2490. r/LocalLLaMA TIER_1 English(EN) · /u/InfinriDev ·

    我使用本地知识图谱和混合RAG构建了一个AI编码代理的执行层

    <!-- SC_OFF --><div class="md"><p>I know this sub is focused on local models but the architecture behind this applies to any LLM-powered coding agent, not just Claude Code.</p> <p>The problem: when you give a coding agent a large set of rules and standards, two things break. The …

  2491. dev.to — LLM tag TIER_1 English(EN) · Ye Allen ·

    使用多模型API网关构建AI代理、RAG应用和聊天机器人

    <p>AI products are becoming more complex than a single prompt and a single model.</p> <p>A chatbot may need fast responses for common questions. A RAG application may need stronger reasoning over retrieved documents. An AI agent may need reliable planning, tool use, and structure…

  2492. dev.to — LLM tag TIER_1 English(EN) · Manas Sharma ·

    如何在生产环境中监控AI代理

    <blockquote> <p><strong>TLDR</strong></p> <ul> <li>Monitoring AI agents in production requires distributed tracing: a single user request fans out into 10 or more internal operations, and logs alone cannot show you which step is slow, failing, or burning your token budget.</li> <…

  2493. dev.to — LLM tag TIER_1 English(EN) · Akash Thakur ·

    为 AI Agent 进行工程化部署

    <blockquote> <p><strong>Agent = Model + Harness.</strong> If you're not the model, you're the harness. </p> </blockquote> <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%…

  2494. dev.to — LLM tag TIER_1 English(EN) · Aryan Panwar ·

    什么是 Agentic AI 开发者?(以及为何它是 2026 年最抢手的职位)

    <p>Most people still think AI engineering = prompt engineering.</p> <p>That's like saying software engineering = writing if statements.</p> <p>I'm Aryan Panwar — a final-year ECE student at MIET Meerut who has shipped 3 live AI products, published a research paper, and built an o…

  2495. dev.to — LLM tag TIER_1 English(EN) · Cristiano Gabrieli ·

    SilentRecon Agent Loop 架构:我们如何构建不会停滞的 AI

    <p>When people talk about “AI agents,” they imagine something autonomous, intelligent, and reliable. In reality, most agents collapse under their own weight: they stall, drift, hallucinate, or loop themselves into oblivion. The problem isn’t the model — it’s the architecture.<br …

  2496. dev.to — LLM tag TIER_1 English(EN) · Logan ·

    AI Agent Runbook:大多数团队缺失的按需运维手册

    <p>On May 1, 2026, an AI coding agent at software company PocketOS deleted a production database — including all available backups — within seconds. The agent was running via Cursor using an Anthropic model. A credential problem led it to improvise: it used an API token intended …

  2497. dev.to — LLM tag TIER_1 English(EN) · Scott McMahan ·

    多智能体AI系统正成为AI工程的未来

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fns9b8lbg4qqcbfzdenhg.jpg"><img alt="building multi-agent ai sy…

  2498. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    Agentic AI at Machine Speed: How Autonomous Agents Break Your Security Assumptions

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/agentic-ai-at-machine-speed-how-autonomous-agents-break-your-security-assumptions?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noreferrer">CoreProse…

  2499. dev.to — LLM tag TIER_1 English(EN) · Gursharan Singh ·

    AI 代理实践 — 第三部分:控制循环如何实际运作

    <p><em>Part 3 of 8 - AI Agents in Practice series.</em></p> <p><em>Previous - <a href="https://dev.to/gursharansingh/ai-agents-in-practice-part-2-what-makes-something-an-agent-bhm">What Makes Something an Agent? (Part 2)</a></em></p> <p>Part 2 named the control loop in five words…

  2500. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    深入了解 Google 的 Agent Executor:用于生产 AI Agent 的开放运行时

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/inside-google-s-agent-executor-open-runtime-for-production-ai-agents?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noreferrer">CoreProse KB-incidents…

  2501. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🧠 人工智能代理正在行业内的各种技术系统和应用中部署。组织正在解决集成挑战和运营

    🧠 AI agents are being deployed in various technical systems and applications across the industry. Organizations are addressing integration challenges and operational complexities that arise from these implementations. 💬 Hacker News 🔗 https://www. wired.com/story/how-ai-agents- pl…

  2502. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    传统软件开发正迅速演变为Agentic AI工程。未来的开发者可能构建:• AI Agent • 自动化工作流 • 智能

    Traditional software development is rapidly evolving into Agentic AI engineering. Future developers may build: • AI Agents • autonomous workflows • intelligent enterprise systems instead of only dashboards and CRUD apps. The future of software is becoming autonomous. Read: https:…

  2503. dev.to — LLM tag TIER_1 English(EN) · Omnithium ·

    什么是AI代理?2026年完整指南

    <p>AI agents are transforming how businesses automate complex workflows. Unlike traditional automation tools that follow rigid rules, AI agents can reason, plan, and adapt to new situations -- making them the next evolution in enterprise software.</p> <h2> What Is an AI Agent? </…

  2504. dev.to — LLM tag TIER_1 English(EN) · Uma Baleboyina ·

    从简单的大语言模型到智能AI代理

    <p><strong>Understanding Deep Agents and Agentic AI</strong></p> <p>Artificial Intelligence has evolved from simple text generation models to intelligent systems called AI Agents. Before understanding agents, we first need to understand how Large Language Models (LLMs) work.</p> …

  2505. dev.to — LLM tag TIER_1 English(EN) · Marcus Chen ·

    面向工具调用代理的Token级评估框架:我们是如何实现的

    <p><strong>TL;DR: We replaced our "did the agent finish the task" pass/fail eval with a token-level harness that scores tool selection, argument shape, and recovery behavior separately. Pass rate went from a single 73% number to four signals that actually tell us what broke. Bifr…

  2506. r/LocalLLaMA TIER_1 English(EN) · /u/Signal_Ad657 ·

    征求反馈:构建更易于本地部署的AI

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1toa14h/feedback_wanted_building_for_easier_local_ai/"> <img alt="Feedback Wanted: Building for easier local AI" src="https://external-preview.redd.it/SZCX7dg3NFHTqfnFBN_B2x0Bg9mPEgknyn6sxShWIvY.png?width=640&…

  2507. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    软件行业或将进入后应用时代。AI Agent正演变为能够进行:•推理 •工作流编排 •决策的自主系统

    The software industry may be entering the post-app era. AI Agents are evolving into autonomous systems capable of: • reasoning • workflow orchestration • decision making • enterprise automation Future software may shift from: Human → App → Action to: Human → AI Agent → Autonomous…

  2508. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用AI代理!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2509. dev.to — LLM tag TIER_1 English(EN) · Anna Jambhulkar ·

    超越提示词:为何您的AI代理需要一个治理运行时

    <p>If you’ve been building with LLMs lately, you probably know the pattern.</p> <p>You start with a simple system prompt.</p> <p>Then the product grows.</p> <p>Then the prompt becomes longer.</p> <p>Then you add rules.</p> <p>Then you add exceptions.</p> <p>Then you add examples.…

  2510. dev.to — LLM tag TIER_1 English(EN) · Alessandro Marocchini ·

    CKP LLM:您的 AI 代理与其知识库之间的缺失层

    <p>Last week my AI coding agent gave me a confident, detailed answer — referencing the wrong project entirely.</p> <p>The problem was not the model. It was context: the agent had loaded 20 knowledge files and picked the wrong one to answer from. The signal was buried in noise.</p…

  2511. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    深入了解自我改进的AI系统,解锁免费的100万token上下文窗口 DeepSeek V4与Hermes Agent的集成带来了显著的

    Inside the Self-Improving AI System Unlocking a Free 1-Million-Token Context Window The integration of DeepSeek V4 with the Hermes Agent introduces a significant enhancement to open source AI capab... #AI #Guides Origin | Interest | Match

  2512. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v49)

    <h1> 터미널 AI 에이전트 구축 (v49) </h1> <h2> 개발자들을 위한 로컬 터미널 AI 에이전트 구축 가이드 </h2> <p>개발자들은 점점 더 AI를 코드 작성에 통합하고 있습니다. 하지만 기존 도구들은 성능 저하, 비공개 데이터 문제, 느린 응답 속도 등의 문제를 가지고 있습니다. 이 가이드에서는 로컬에서 실행되는 빠르고 안전한 터미널 AI 에이전트를 구축하는 방법을 실습 중심으로 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 분석 </h2> <h3> 주요 도구들 …

  2513. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v48)

    <h1> 터미널 AI 에이전트 구축 (v48) </h1> <p><strong>개발자들을 위한 로컬 AI 코딩 에이전트 구축 가이드</strong></p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트 시장은 다양한 솔루션으로 분산되어 있습니다:</p> <h3> 주요 플랫폼 비교 </h3> <p><strong>Aider</strong>: GitHub Copilot 기반의 실시간 코드 작성 도구<br /> </p> <div class="highlight js-c…

  2514. dev.to — LLM tag TIER_1 English(EN) · Harsh Manvar ·

    Docker with AI:运行大型语言模型、智能体和 MCP 的实用指南

    <p>If you've been searching for how to actually use Docker with AI not just spin up a demo but run models, agents and MCP servers in production here's what We have learned over the years and put into our new book.</p> <p><a class="article-body-image-wrapper" href="https://media2.…

  2515. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v47)

    <h1> 터미널 AI 에이전트 구축 (v47) </h1> <h2> CLI AI 에이전트 생태계 </h2> <p>터미널에서 작동하는 AI 에이전트는 이미 다양한 형태로 존재합니다. 현재 주요 도구는 다음과 같습니다:</p> <p><strong>Aider</strong>: GitHub Copilot과 유사한 기능을 제공하며, 파일 단위로 코드를 생성하고 수정합니다. 주요 특징은 소스 코드가 있는 파일과 현재 작업 디렉토리 기반의 콘텍스트를 사용하는 것입니다.<br /> </p> <div class="…

  2516. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建一个终端 AI 代理 (v46)

    <h1> 터미널 AI 에이전트 구축 (v46) </h1> <p>터미널에서 직접 작동하는 AI 에이전트를 구축해보는 실전 가이드입니다. 이 가이드는 로컬에서 작동하는 LLM을 활용한 개발자용 AI 에이전트를 구축하고 최적화하는 방법을 실습 중심으로 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트 시장은 다음과 같은 주요 도구들로 구성되어 있습니다:</p> <h3> 주요 도구 비교: </h3> <div class="highlight js-cod…

  2517. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端AI代理 (v45)

    <h1> 터미널 AI 에이전트 구축 (v45) </h1> <p>터미널에서 작동하는 AI 에이전트는 개발자들에게 강력한 도구가 되지만, 대부분의 기존 솔루션은 복잡하거나 클라우드 기반으로 의존합니다. 이 가이드는 로컬에서 작동하는 가벼운 AI 에이전트를 구축하여 코드 리뷰, 자동완성, 프로젝트 탐색을 수행하는 실용적인 방법을 설명합니다.</p> <h2> 1. CLI AI 에이전트 랜드스케이프 </h2> <h3> 기존 솔루션 비교 </h3> <p><strong>Aider</strong>: GitHub…

  2518. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v44)

    <h1> 터미널 AI 에이전트 구축 (v44) </h1> <p>터미널에서 실행되는 AI 에이전트를 구축하는 것은 현대 개발자에게 매우 실용적인 기술입니다. 이 가이드에서는 로컬 LLM을 기반으로 하는 터미널 AI 에이전트를 구축하고 운영하는 방법을 단계별로 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트 시장은 다음과 같은 주요 플랫폼들로 구성되어 있습니다:</p> <h3> Aider </h3> <p>가장 인기 있는 오픈소스 터미널 AI 에…

  2519. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v43)

    <h1> 터미널 AI 에이전트 구축 (v43) </h1> <h2> 개발자를 위한 터미널 AI 에이전트 구축 가이드 </h2> <p>최근 몇 년 동안 개발자들은 로컬 AI 에이전트를 구축하여 코드 작업을 자동화하고 효율성을 높이는 데 집중하고 있습니다. 이 가이드에서는 실제 개발자가 사용할 수 있는 터미널 기반 AI 에이전트 구축 방법을 안내합니다. </p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 터미널에서 작동하는 AI 에이전트는 다음과 같은 주요 플랫폼들로 구성되어 있습…

  2520. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端AI代理 (v42)

    <h1> 터미널 AI 에이전트 구축 (v42) </h1> <p>터미널에서 AI를 활용한 개발 워크플로우는 점점 더 중요해지고 있습니다. 이 가이드는 로컬 AI 에이전트를 구축하여 터미널에서 직접 사용할 수 있도록 도와주는 실질적인 방법을 제공합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 터미널 AI 에이전트 시장은 다음과 같은 주요 플랫폼으로 구성되어 있습니다:</p> <p><strong>Aider</strong>: GitHub Copilot과 유사한 기능을 제공…

  2521. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v41)

    <h1> 터미널 AI 에이전트 구축 (v41) </h1> <p>터미널에서 작동하는 AI 에이전트를 구축하는 것은 개발자들이 코드를 더 빠르고 효율적으로 작성할 수 있게 해주는 실용적인 도구입니다. 이번 가이드에서는 로컬 환경에서 작동하는 AI 에이전트를 구축하고 최적화하는 방법을 단계별로 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트 시장은 다음과 같은 주요 도구들로 구성되어 있습니다:</p> <h3> Aider </h3> <p>가장 인기…

  2522. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v40)

    <h1> 터미널 AI 에이전트 구축 (v40) </h1> <p>터미널에서 작동하는 AI 에이전트는 개발자에게 실시간 코드 보조, 자동화, 문제 해결을 제공하는 강력한 도구입니다. 이 가이드에서는 실제 개발 환경에서 활용 가능한 터미널 AI 에이전트를 구축하는 방법을 단계별로 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 분석 </h2> <p>현재 터미널 기반 AI 에이전트 시장은 다음과 같은 주요 플랫폼으로 구성되어 있습니다:</p> <h3> Aider </h3> <div clas…

  2523. dev.to — LLM tag TIER_1 English(EN) · Andrew ·

    2026年中国AI模型:智能体革命、硬件独立及其对全球开发者的意义

    <p>If you’ve only been paying attention to OpenAI and Google’s AI offerings in recent years, you’re missing half the story. As of May 2026, China’s AI ecosystem has completed a dramatic pivot from the 2023-2025 “model war” of racing to build ever-larger parameter models to an “ag…

  2524. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v39)

    <h1> 터미널 AI 에이전트 구축 (v39) </h1> <p>터미널에서 작동하는 AI 에이전트를 구축하는 것은 현대 개발 워크플로우를 혁신할 수 있는 강력한 도구입니다. 이 가이드는 실질적인 비용(3-7달러)으로 구축할 수 있는 터미널 기반 AI 에이전트를 구축하는 실전 가이드입니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트 생태계는 다음과 같은 주요 도구들로 구성됩니다:</p> <h3> Aider (가장 인기) </h3> <div class=…

  2525. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v38)

    <h1> 터미널 AI 에이전트 구축 (v38) </h1> <p>터미널에서 작동하는 AI 에이전트를 구축하여 개발 생산성을 향상시킬 수 있습니다. 이 가이드에서는 로컬 LLM API 엔드포인트 설정부터 커스텀 CLI 에이전트 구축까지 실질적인 방법을 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트 시장은 다양한 도구로 구성되어 있습니다:</p> <h3> 대표 도구 비교 </h3> <p><strong>Aider</strong>: GitHub C…

  2526. dev.to — LLM tag TIER_1 English(EN) · Lingdas1 ·

    Gemma 4:谷歌轻量级强大模型 — 在您已有的硬件上运行AI

    <h1> Gemma 4: Google's Lightweight Powerhouse </h1> <blockquote> <p><strong>Don't have a $2000 GPU? Gemma 4 runs AI on hardware you already own.</strong></p> </blockquote> <h2> Why Gemma 4 Exists </h2> <p>Google built Gemma 4 for one specific use case: <strong>running capable AI …

  2527. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🧠 成功的人工智能开发并非偶然。Collin Newberry 探讨了上下文工程、提示工程、知识管理和结构化工作流程如何

    🧠 Successful AI development isn’t accidental. Collin Newberry explores how context engineering, prompt engineering, knowledge management, and structured workflows separate effective AI pair programming from chaotic vibe coding. https://www. nebraska-code.com/ # AI # SoftwareEngin…

  2528. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v37)

    <h1> 터미널 AI 에이전트 구축 (v37) </h1> <p>터미널에서 AI 에이전트를 구축하는 것은 개발자에게 매우 실용적인 도구를 제공합니다. 이 가이드는 로컬 LLM을 활용한 CLI AI 에이전트를 구축하고, 실전 워크플로우에 적용하는 방법을 단계별로 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트는 여러 형태로 존재합니다:</p> <p><strong>Aider</strong>: GitHub에서 개발된 코드 생성 도구로, 실제 파일에…

  2529. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v36)

    <h1> 터미널 AI 에이전트 구축 (v36) </h1> <p>터미널에서 작동하는 AI 에이전트를 구축하는 것은 현대 개발 워크플로우에서 핵심적인 도구로 자리 잡고 있습니다. 이 가이드는 실질적인 비용 ($3-$7)의 가치를 제공하는 터미널 기반 AI 에이전트를 구축하는 방법을 다룹니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트는 다양한 솔루션으로 구성되어 있습니다:</p> <p><strong>Aider</strong>: Git 기반 코드 생성 …

  2530. r/MachineLearning TIER_1 English(EN) · /u/Alarming_Rou_3841 ·

    重构智能体方法论:解耦决策与执行 - 开源 [P]

    <!-- SC_OFF --><div class="md"><p>I’ve been thinking about a problem in current agent systems:</p> <p>Most agents are becoming very good at execution, but the decision layer before execution is still unclear.</p> <p>Coding agents, research agents, tool loops, sandboxes, workflows…

  2531. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v35)

    <h1> 터미널 AI 에이전트 구축 (v35) </h1> <p>터미널에서 작동하는 AI 에이전트를 직접 구축하여 개발 생산성을 높이는 방법을 안내합니다. 이 가이드는 로컬에서 실행 가능한 고성능 AI 에이전트를 구축하는 실용적인 접근법을 제공합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 터미널 AI 에이전트 시장은 다음과 같은 주요 플랫폼으로 구성되어 있습니다:</p> <h3> 주요 도구 비교 </h3> <p><strong>Aider</strong>:<br /> …

  2532. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v34)

    <h1> 터미널 AI 에이전트 구축 (v34) </h1> <p>터미널에서 AI 코드 보조 도구를 직접 구축하는 실전 가이드</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트 시장은 다음과 같은 주요 플랫폼들로 구성되어 있습니다:</p> <p><strong>Aider</strong>: GitHub Copilot과 유사하지만 오픈소스 버전. <code>aider --help</code> 명령으로 간단히 시작 가능합니다.</p> <p><strong>Contin…

  2533. r/MachineLearning TIER_1 English(EN) · /u/Alarming_Rou_3841 ·

    我正在构建一个AI代理之上的开源决策层 [P]

    <!-- SC_OFF --><div class="md"><p>Hi everyone, I’m Jia, the creator of Spice.</p> <p>I’ve been working on an open-source project called Spice.</p> <p>The simplest way to describe it is:</p> <p>Spice is a decision layer above agents.</p> <p>Most agent systems today are very focuse…

  2534. dev.to — LLM tag TIER_1 English(EN) · Wallet Guy ·

    能够支付自身计算成本的AI代理:缺失的经济层

    <p>AI agents will need to pay for compute, data, and API calls—but how do they access economic primitives without relying on human-managed accounts? The missing piece isn't better models or more training data. It's autonomous wallet infrastructure that lets agents participate in …

  2535. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v33)

    <h1> 터미널 AI 에이전트 구축 (v33) </h1> <h2> 개요 </h2> <p>터미널에서 동작하는 AI 에이전트는 개발자에게 코드 생성, 분석, 리팩토링을 위한 실시간 도우미를 제공합니다. 이 가이드에서는 오픈소스 AI 에이전트를 구축하고 최적화하는 실전 방법을 소개합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트는 다음과 같은 주요 도구들로 구성되어 있습니다:</p> <h3> Aider </h3> <p>가장 인기 있는 오픈소스 도구로,…

  2536. dev.to — LLM tag TIER_1 English(EN) · AK DevCraft ·

    本地运行LLM - 0美元个人代理AI助手 - 第三部分

    <h2> Introduction </h2> <p><em>Part 3 of the Zero Dollar personal AI Assistant series, running Local LLMs on a Free Cloud Server — What Actually Works. <a href="https://dev.to/akdevcraft/running-a-personal-ai-assistant-for-0-part-1-architecture-3j45">Part 1</a> covers the archite…

  2537. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v32)

    <h1> 터미널 AI 에이전트 구축 (v32) </h1> <h2> 개발자용 CLI AI 에이전트 구축 가이드 </h2> <p>터미널에서 작동하는 AI 에이전트는 개발자의 생산성을 높이는 강력한 도구입니다. 이 가이드에서는 실제 개발자들이 필요로 하는 3-7달러 범위의 실용적 CLI AI 에이전트를 구축하는 방법을 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 분석 </h2> <h3> 현재 선택지 비교 </h3> <p><strong>Aider</strong>: GitHub Copil…

  2538. Mastodon — fosstodon.org TIER_1 日本語(JA) · [email protected] ·

    AI工程师安野:什么是AI Agent?/“自主行动”的AI潜力 / 值得关注的AI产品

    【AIエンジニア安野氏】AIエージェントとは何か? / 「自律的に行動する」AIの可能性 / 注目のAIプロダクト https://www. emilyselect.com/%e3%80%90ai%e3 %82%a8%e3%83%b3%e3%82%b8%e3%83%8b%e3%82%a2%e5%ae%89%e9%87%8e%e6%b0%8f%e3%80%91ai%e3%82%a8%e3%83%bc%e3%82%b8%e3%82%a7%e3%83%b3%e3%83%88%e3%81%a8%e3%81%af%e4%bd%95%e3%81%8b%ef%bc%9…

  2539. Mastodon — fosstodon.org TIER_1 Polski(PL) · [email protected] ·

    微软的 Fara1.5 模型在 AI 代理测试中达到 72% 的有效性,超越 OpenAI Operator 和 Google Gemini。新一代开源模型 r

    Model Fara1.5 od Microsoftu osiągnął 72% skuteczności w testach agentów AI, pokonując OpenAI Operator i Google Gemini. Nowa rodzina modeli o otwartych wagach rzuca wyzwanie gigantom, oferując tańszą i bezpieczniejszą automatyzację przeglądarki. # si # ai # sztucznainteligencja # …

  2540. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v31)

    <h1> 터미널 AI 에이전트 구축 (v31) </h1> <p>터미널에서 작동하는 AI 에이전트를 구축하면 코드 작성 속도가 2배 이상 향상됩니다. 이 가이드에서는 실제 개발자가 사용할 수 있는 터미널 AI 에이전트를 구축하는 방법을 단계별로 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 터미널 AI 에이전트는 다음과 같은 솔루션으로 구성되어 있습니다:</p> <h3> Aider </h3> <div class="highlight js-code-highlight…

  2541. Mastodon — fosstodon.org TIER_1 Italiano(IT) · [email protected] ·

    🚨 Fabric AI:安装开源框架,将 AI 模式带入终端 — macOS 和 Linux 上的 Unix 管道、Ollama 集成和可重用提示

    🚨 Fabric AI: installa il framework open source che porta i pattern AI nel terminale — piping Unix, integrazione Ollama e prompt riutilizzabili su macOS e Linux https:// gomoot.com/come-installare-il- framework-fabric-ai-per-usare-i-pattern-ai-da-terminale-su-ollama/ # AI # fabric…

  2542. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端AI代理 (v30)

    <h1> 터미널 AI 에이전트 구축 (v30) </h1> <p>터미널에서 작동하는 AI 에이전트로 개발 생산성을 높이는 방법을 실전 가이드로 안내드립니다. 이 가이드는 30불 이하의 가격으로 구입할 수 있는 실용적인 도구와 기술을 중심으로 구성되었습니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 터미널 AI 에이전트 시장은 다양한 솔루션으로 구성되어 있습니다:</p> <h3> 주요 도구 비교 </h3> <p><strong>Aider</strong>: Python 기반…

  2543. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v29)

    <h1> 터미널 AI 에이전트 구축 (v29) </h1> <p>터미널에서 직접 작동하는 AI 에이전트는 코드 개발의 핵심 도구로 자리 잡고 있습니다. 이 가이드에서는 실용적인 터미널 AI 에이전트 구축 방법을 다룹니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트는 다음과 같은 주요 플랫폼으로 분류됩니다:</p> <h3> Aider </h3> <div class="highlight js-code-highlight"> <pre class="highli…

  2544. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v28)

    <h1> 터미널 AI 에이전트 구축 (v28) </h1> <p>터미널에서 작동하는 AI 에이전트를 구축하는 것은 현대 개발 워크플로우를 혁신할 수 있는 실용적인 도구입니다. 이 가이드는 실제 개발자가 사용할 수 있는 터미널 기반 AI 에이전트를 구축하는 방법을 자세히 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트 시장은 다음과 같은 주요 플랫폼으로 구성되어 있습니다:</p> <p><strong>Aider</strong>: GitHub Co…

  2545. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v27)

    <h1> 터미널 AI 에이전트 구축 (v27) </h1> <p>터미널에서 작동하는 AI 에이전트를 구축하는 것은 현대 개발자에게 매우 실용적인 도구입니다. 이 가이드에서는 실제 개발 workflow에 통합할 수 있는 로컬 LLM 기반 CLI 에이전트를 구축하는 방법을 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트 시장에는 여러 선택지가 있습니다:</p> <p><strong>Aider</strong>: Git 기반 코드 수정을 위한 간단한 …

  2546. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v26)

    <h1> 터미널 AI 에이전트 구축 (v26) </h1> <p>터미널에서 직접 작동하는 AI 에이전트를 구축하면, 코드 작성과 디버깅을 더 효율적으로 할 수 있습니다. 이 가이드는 터미널 내에서 작동하는 AI 에이전트를 구축하는 실전 가이드입니다.</p> <h2> 1. CLI AI 에이전트 환경 분석 </h2> <p>현재 CLI AI 에이전트 시장은 다양한 솔루션으로 구성되어 있습니다:</p> <ul> <li> <strong>Aider</strong>: GitHub Copilot과 유사한 기능을 …

  2547. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v25)

    <h1> 터미널 AI 에이전트 구축 (v25) </h1> <p>터미널에서 AI를 활용한 개발 흐름을 구축하는 것은 현대 개발자에게 필수적인 기술입니다. 이 가이드에서는 실제 개발자들이 실제로 사용할 수 있는 터미널 AI 에이전트를 구축하는 방법을 단계별로 안내합니다.</p> <h2> 1. CLI AI 에이전트 랜드스케이프 </h2> <p>현재 터미널 AI 에이전트 시장은 다양합니다:</p> <p><strong>Aider</strong>: GitHub의 오픈소스 에이전트로, VS Code와 같은 I…

  2548. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v24)

    <h1> 터미널 AI 에이전트 구축 (v24) </h1> <p>터미널에서 작동하는 AI 에이전트를 구축하면 개발자들이 코드를 더 빠르고 효율적으로 작성할 수 있습니다. 이 가이드에서는 실제 사용 가능한 터미널 AI 에이전트를 구축하는 방법을 단계별로 설명합니다.</p> <h2> 1. CLI AI 에이전트 랜드스케이프 </h2> <p>현재 CLI AI 에이전트 시장에는 여러 선택지가 있습니다:</p> <p><strong>Aider</strong>: Git 기반 코드 변경을 위한 자동화 도구로, 터미…

  2549. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v23)

    <h1> 터미널 AI 에이전트 구축 (v23) </h1> <p>터미널에서 AI를 활용한 개발 도구는 점점 더 인기를 끌고 있습니다. 오픈소스 커뮤니티와 전문 개발자들 사이에서 로컬 LLM 추론과 자가 호스팅 AI 솔루션에 대한 관심이 높아지고 있습니다. 이 가이드에서는 터미널 내에서 작동하는 AI 에이전트를 구축하는 실용적인 방법을 제공합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트의 주요 도구들:</p> <ul> <li> <strong>Aid…

  2550. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v22)

    <h1> 터미널 AI 에이전트 구축 (v22) </h1> <p>터미널에서 작동하는 AI 에이전트를 구축하는 것은 현대 개발 워크플로우에서 점점 더 중요해지고 있습니다. 이 가이드에서는 개발자들이 실제 사용할 수 있는 터미널 AI 에이전트를 구축하고 최적화하는 방법을 설명합니다.</p> <h2> 1. CLI AI 에이전트 랜드스케이프 </h2> <p>현재 CLI AI 에이전트 시장에는 여러 선택지가 있습니다:</p> <p><strong>Aider</strong>: GitHub의 코드 리뷰 도우미로,…

  2551. dev.to — LLM tag TIER_1 English(EN) · Murni Marcus ·

    开源我们的游戏AI堆栈 — 用于NPC对话的SDK、模板和CLI工具

    <h1> Open-Sourcing Our Game AI Stack </h1> <p>At <a href="https://vantage-digital.online" rel="noopener noreferrer">Vantage Digital Labs</a>, we've been building AI-powered NPC dialogue systems for games. Most of our internal tooling is now stable enough to share. We're releasing…

  2552. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v21)

    <h1> 터미널 AI 에이전트 구축 (v21) </h1> <p>터미널에서 작동하는 AI 에이전트를 구축하여 코드 작성과 리팩토링을 자동화하는 것은 현대 개발 워크플로우의 핵심입니다. 이 가이드는 실제 개발자가 사용할 수 있는, 저렴하고 효율적인 터미널 AI 에이전트 구축 방법을 다룹니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 터미널 AI 에이전트 시장은 다음과 같은 주요 플랫폼으로 구성되어 있습니다:</p> <p><strong>Aider</strong>: GitH…

  2553. dev.to — LLM tag TIER_1 English(EN) · AI Bug Slayer 🐞 ·

    人工智能代理革命:企业如何实现万物自动化 [03:31:50]

    <p><em>Hey there! If you've been keeping up with the AI space lately, you know we're in the middle of something genuinely historic. What used to be science fiction is becoming production code — and it's happening fast.</em></p> <h2> The Big Shift: Agents Over Assistants </h2> <p>…

  2554. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v20)

    <h1> 터미널 AI 에이전트 구축 (v20) </h1> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>터미널에서 작동하는 AI 에이전트는 최근 두드러진 트렌드입니다. 주요 플랫폼들:</p> <h3> Aider </h3> <div class="highlight js-code-highlight"> <pre class="highlight shell"><code><span class="c"># 설치</span> pip <span class="nb">install </span>aider <s…

  2555. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端AI代理 (v17)

    <h1> 터미널 AI 에이전트 구축 (v17) </h1> <p>터미널에서 작동하는 AI 에이전트를 구축하여 개발 생산성을 극대화하는 방법을 알아봅니다. 이 가이드에서는 오픈소스 도구와 커스텀 솔루션을 사용해 실용적인 터미널 AI 에이전트를 구현하는 방법을 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 터미널 AI 에이전트는 여러 플랫폼으로 나뉩니다:</p> <h3> 주요 도구 비교 </h3> <div class="highlight js-code-highligh…

  2556. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v16)

    <h1> 터미널 AI 에이전트 구축 (v16) </h1> <p>터미널에서 직접 작동하는 AI 에이전트를 구축하는 것은 현대 개발자에게 매우 실용적인 도구입니다. 이 가이드는 개발자가 직접 자신의 터미널 환경에서 효율적인 AI 코딩 어시스턴트를 구축하는 방법을 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI 기반 AI 에이전트는 다음과 같은 주요 플랫폼이 있습니다:</p> <p><strong>Aider</strong>: Git 기반의 코딩 에이전트로, 코드…

  2557. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent Adoption: A Practical Roadmap 成功采用 AI Agent!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2558. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v15)

    <h1> 터미널 AI 에이전트 구축 (v15) </h1> <p>터미널에서 직접 작동하는 AI 에이전트를 구축하는 것은 현대 개발자의 생산성을 높이는 가장 효과적인 방법 중 하나입니다. 이 가이드에서는 개발자가 직접 구축할 수 있는 로컬 LLM 기반 CLI AI 에이전트를 구축하는 방법을 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI 기반 AI 에이전트 생태계는 다음과 같은 주요 도구들로 구성되어 있습니다:</p> <h3> Aider </h3> <p>가장…

  2559. dev.to — LLM tag TIER_1 English(EN) · logicgrid-dev ·

    推出 LogicGrid — .NET 的多智能体 AI 编排

    <p>If you've spent any time building with LLMs, you've probably hit the wall: a single prompt only gets you so far. Stuff too much into one prompt and the model loses the plot. Try to do too many things at once and you get inconsistent output.</p> <p>The answer most teams converg…

  2560. dev.to — LLM tag TIER_1 English(EN) · Joseph Anady ·

    Agentic AI Search

    <blockquote> <p><strong>Originally published at <a href="https://www.thatdevpro.com/insights/framework-agenticaisearch/" rel="noopener noreferrer">thatdevpro.com</a>.</strong> This framework reference is part of the 14-tier Engine Optimization stack from <a href="https://www.that…

  2561. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v14)

    <h1> 터미널 AI 에이전트 구축 (v14) </h1> <p>터미널에서 작동하는 AI 에이전트는 현대 개발 워크플로우의 핵심 요소입니다. 이 가이드에서는 개발자가 실제로 사용할 수 있는 터미널 AI 에이전트를 구축하는 방법을 자세히 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 터미널 AI 에이전트는 다양한 도구로 구성되어 있습니다:</p> <p><strong>Aider</strong>: GitHub Copilot과 유사한 기능을 제공하는 에이전트<br />…

  2562. dev.to — LLM tag TIER_1 English(EN) · Anjaiah Methuku ·

    停止盲目飞行:我们构建了一个适用于 17 多个 Agent 框架的 LLM 评估框架

    <p>Let me be brutally honest with you.</p> <p>I've seen teams demo AI agents that look incredible — smooth responses, beautiful UI, stakeholders impressed. Then that same team ships to production and spends the next three weeks firefighting hallucinations they could have caught i…

  2563. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v13)

    <h1> 터미널 AI 에이전트 구축 (v13) </h1> <p>터미널에서 AI 코딩 어시스턴트를 직접 구축하는 실전 가이드</p> <h2> 1. CLI AI 에이전트 생태계 분석 </h2> <p>현재 터미널 기반 AI 에이전트는 다양한 솔루션으로 구성되어 있습니다:</p> <p><strong>Aider</strong>: GitHub Copilot처럼 코드 생성 및 수정을 지원하는 에이전트<br /> </p> <div class="highlight js-code-highlight"> <pre cla…

  2564. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端AI代理 (v12)

    <h1> 터미널 AI 에이전트 구축 (v12) </h1> <p>터미널에서 직접 작동하는 AI 에이전트를 구축하여 개발 워크플로우를 최적화하세요. 이 가이드는 개발자들이 직접 구축하고 커스터마이징할 수 있는 실질적인 터미널 AI 에이전트를 제공합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트 생태계는 다음과 같은 주요 도구들로 구성되어 있습니다:</p> <h3> Aider </h3> <div class="highlight js-code-highli…

  2565. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    教皇利奥十四世、Christopher Olah 和 Claude Mythos:为前沿模型起草人工智能通谕

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/pope-leo-xiv-christopher-olah-and-claude-mythos-drafting-an-ai-encyclical-for-frontier-models?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noreferre…

  2566. dev.to — LLM tag TIER_1 English(EN) · Otto Plane ·

    为Agentic AI架构实现确定性运行时追踪

    <h2> Introduction </h2> <p>As production AI workloads transition from stateless chat completions to autonomous, multi-agent workflows, legacy observability infrastructure is proving insufficient. Standard application performance monitoring (APM) tools are built to trace predictab…

  2567. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建一个终端 AI 代理 (v11)

    <h1> 터미널 AI 에이전트 구축 (v11) </h1> <p>터미널에서 작동하는 AI 에이전트는 개발자에게 매우 가치 있는 도구입니다. 이 가이드에서는 실제 개발 환경에서 사용할 수 있는 터미널 AI 에이전트 구축 방법을 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 터미널 AI 에이전트는 여러 플랫폼으로 구성되어 있습니다:</p> <h3> 주요 도구들 </h3> <p><strong>Aider</strong>: Git 기반 코드 수정을 위한 간단한 에이전트<…

  2568. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v10)

    <h1> 터미널 AI 에이전트 구축 (v10) </h1> <p>터미널에서 작동하는 AI 에이전트를 직접 구축하는 것은 개발자에게 매우 실용적인 도구입니다. 이 가이드에서는 로컬 LLM을 활용한 터미널 AI 에이전트를 구축하고, 실제 개발 워크플로우에 적용하는 방법을 단계별로 안내합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트 생태계는 여러 도구로 구성되어 있습니다:</p> <h3> 주요 도구 비교 </h3> <p><strong>Aider</st…

  2569. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v9)

    <h1> 터미널 AI 에이전트 구축 (v9): 로컬 LLM 기반 개발자용 CLI AI 에이전트 만들기 </h1> <p>터미널에서 직접 작동하는 AI 에이전트를 구축하는 것은 개발자에게 큰 생산성 향상을 제공합니다. 이번 가이드에서는 로컬 LLM을 기반으로 한 커스텀 CLI AI 에이전트를 구축하는 방법을 실습 중심으로 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 분석 </h2> <p>현재 CLI AI 에이전트 시장에는 여러 솔루션이 존재합니다:</p> <h3> 주요 도구들: </…

  2570. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v8)

    <h1> 터미널 AI 에이전트 구축 (v8) </h1> <p>터미널에서 직접 작동하는 AI 에이전트를 구축하는 것은 개발자들이 직면하는 현실적인 문제를 해결할 수 있는 강력한 도구입니다. 특히 로컬 환경에서 AI를 활용하면서도 성능과 보안을 고려해야 하는 상황에서는 더욱 중요합니다. 이번 가이드에서는 로컬 LLM API를 활용하여 개발자 친화적인 터미널 AI 에이전트를 구축하는 방법을 단계별로 설명합니다.</p> <h2> 1. CLI AI 에이전트 랜드스케이프 </h2> <p>현재 터미널 기반 A…

  2571. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v7)

    <h1> 터미널 AI 에이전트 구축 (v7) </h1> <p>터미널에서 실행되는 AI 에이전트를 구축하여 코드 작성 속도를 높이는 것은 현대 개발자에게 매우 실용적인 도구입니다. 이 가이드에서는 로컬 LLM을 기반으로 한 터미널 AI 에이전트를 구축하고, 실제 개발 워크플로우에 통합하는 방법을 자세히 다룹니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트 시장에는 여러 가지 솔루션이 존재합니다:</p> <p><strong>Aider</strong>:…

  2572. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端 AI 代理 (v6)

    <h1> 터미널 AI 에이전트 구축 (v6) </h1> <p>터미널에서 직접 작동하는 AI 에이전트를 구축하는 것은 개발자들이 코드를 빠르게 작성하고 문제를 해결하는 데 있어 귀중한 도구가 됩니다. 이 가이드에서는 현대적인 CLI 기반 AI 에이전트를 구축하고 최적화하는 실용적인 방법을 다룹니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트 시장은 다음과 같은 주요 솔루션으로 구성되어 있습니다:</p> <p><strong>Aider</strong>:…

  2573. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    为什么AI在真实SOC中表现仍不佳(以及如何缩小差距)

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/why-ai-still-underperforms-in-real-socs-and-how-to-close-the-gap?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noreferrer">CoreProse KB-incidents</a>…

  2574. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端AI代理 (v5)

    <h1> 터미널 AI 에이전트 구축 (v5) </h1> <p>터미널 기반 AI 에이전트는 개발자에게 매우 실용적인 도구로 자리 잡았습니다. 다양한 CLI 기반 AI 도구들 중에서 가장 효율적인 방식으로 개발자 워크플로우를 개선할 수 있는 방법을 소개합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트 시장은 다음과 같은 주요 도구들로 구성되어 있습니다:</p> <h3> Aider </h3> <div class="highlight js-code-hig…

  2575. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端AI代理 (v4)

    <h1> 터미널 AI 에이전트 구축 (v4) </h1> <p><strong>개발자를 위한 경량 로컬 AI 코딩 어시스턴트 구축 가이드</strong></p> <h2> 1. CLI AI 에이전트 생태계 개요 </h2> <p>터미널 기반 AI 에이전트는 개발자들이 코드를 작성하고 디버깅할 때 실시간으로 도움을 받을 수 있도록 해주는 도구입니다. 현재 주류로는 다음과 같은 솔루션들이 있습니다:</p> <h3> Aider </h3> <div class="highlight js-code-highlight"…

  2576. dev.to — LLM tag TIER_1 한국어(KO) · matias yoon ·

    构建终端AI代理 (v3)

    <h1> 터미널 AI 에이전트 구축 (v3) </h1> <p>터미널에서 작동하는 AI 에이전트는 현대 개발 워크플로우에 필수적인 도구입니다. 이 가이드는 개발자가 로컬 환경에서 효율적으로 작동하는 AI 에이전트를 구축하고 활용하는 방법을 실질적인 코드와 명령어로 설명합니다.</p> <h2> 1. CLI AI 에이전트 생태계 </h2> <p>현재 CLI AI 에이전트 시장은 다음과 같은 주요 플랫폼으로 구성되어 있습니다:</p> <p><strong>Aider</strong>: GitHub Copil…

  2577. dev.to — LLM tag TIER_1 English(EN) · AIInsightsDaily ·

    H1:2026年5月人工智能格局导航:今日关键发展全面概述

    <h1> H1: Navigating AI Landscapes of May 2026: A Comprehensive Overview of Today's Key Developments </h1> <p>Greetings, fellow tech enthusiasts! Today, we delve into an intriguing array of AI news that has caught our attention. Let's explore the fascinating world of AI together a…

  2578. dev.to — LLM tag TIER_1 English(EN) · WonderLab ·

    Agent系列(3):Plan-and-Solve — 先思考,再行动

    <h2> Where Does ReAct Hit a Wall? </h2> <p>The previous article established ReAct's greedy strategy — each step looks at only the current state and decides the next action. This works well most of the time, but there's one class of task where it stumbles.</p> <p>Imagine you ask a…

  2579. dev.to — LLM tag TIER_1 English(EN) · WonderLab ·

    每日一个开源项目 #74: ai-engineering-from-scratch - 从零开始构建AI全栈技能

    <h2> Introduction </h2> <p><strong><a href="https://github.com/rohitg00/ai-engineering-from-scratch" rel="noopener noreferrer">ai-engineering-from-scratch</a></strong> is a hardcore and comprehensive curriculum for AI engineering. Instead of just teaching you how to call the Open…

  2580. dev.to — LLM tag TIER_1 English(EN) · Rahul Talreja ·

    构建私有 RAG 系统:来自本地优先 AI 日志的经验教训

    <p><em>Most AI apps quietly send your data to the cloud. DiaryGPT does the opposite — and this is the full technical story.</em></p> <h2> The Problem With AI + Private Data </h2> <p>When you write in a journal, you write the things you'd never say out loud. The last thing you wan…

  2581. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI Agent 采用:实用路线图 成功采用 AI Agent!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2582. dev.to — LLM tag TIER_1 English(EN) · Iniyarajan ·

    RAG 与微调:何时为 AI Agent 使用哪种方法

    <p>Last week, I was working on an AI agent for a client's customer support system. The agent needed to access constantly changing product documentation while maintaining conversational abilities. That's when the classic question hit me: should I fine-tune a model or build a RAG s…

  2583. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI 代理 — 安全噩梦?理解 OpenClaw https://peertube.eqver.se/w/jjjq3QBmE3U5Fw3AJ6zMeT

    AI Agents — A Security Nightmare? Understanding OpenClaw https:// peertube.eqver.se/w/jjjq3QBmE3 U5Fw3AJ6zMeT

  2584. dev.to — LLM tag TIER_1 English(EN) · Naing Oo ·

    Gemma 4:我在真实硬件上运行Google的开源AI模型时学到的东西

    <p><em>This is a submission for the <a href="https://dev.to/challenges/google-gemma-2026-05-06">Gemma 4 Challenge: Write About Gemma 4</a></em></p> <p>Most AI tutorials show you how to call an API. You send text in, you get text back, and everything works perfectly in a Jupyter n…

  2585. dev.to — LLM tag TIER_1 English(EN) · WonderLab ·

    Agent Series (2): ReAct — 最重要的 Agent 推理范式

    <h2> You Think Your Agent Is "Thinking." It's Actually Just Predicting Tokens. </h2> <p>Here's a scenario that happens more often than you'd think.</p> <p>You ask an Agent to write a competitive analysis report. It confidently outputs three professional-looking pages — complete w…

  2586. dev.to — LLM tag TIER_1 English(EN) · peter.zeng ·

    优化AI编码代理的4个艰难教训

    <h1> 4 Hard Lessons on Optimizing AI Coding Agents (Claude Code + Cost) </h1> <p>I've been running Claude Code Cli in production for about months now—building, shipping, and watching the token meter spin. Here's what I wish I knew before I started.</p> <h2> 1. Your Context Strate…

  2587. dev.to — LLM tag TIER_1 English(EN) · Javier Fajardo ·

    AI智能体堆栈中缺失的一层:机器对机器搜索引擎

    <p>AI agents still search for tools like humans do — parsing READMEs, reading docs, guessing install commands. We built the layer that was missing from every agent stack diagram.</p> <h2> The problem </h2> <p>An AI coding agent needs to send an email. It knows <code>sendgrid</cod…

  2588. dev.to — LLM tag TIER_1 English(EN) · AlterLab ·

    如何通过提取高效率的JSON和元数据来降低AI代理中的LLM推理成本

    <h2> TL;DR </h2> <p>Feeding raw HTML to LLMs wastes input tokens on structural markup, tracking scripts, and inline styling, massively inflating your inference costs. By extracting clean JSON, semantic metadata, or formatting the Document Object Model (DOM) into Markdown before s…

  2589. dev.to — LLM tag TIER_1 English(EN) · Oyedele Temitope ·

    如何将AI开发规模化,超越原型开发速度

    <p>One thing that isn't talked about enough in AI right now is how easy it has become to mistake a working demo for a production-ready system.</p> <p>You can build a working prototype in a few days, whether it's a chatbot that understands internal documents, a recommendation engi…

  2590. dev.to — LLM tag TIER_1 English(EN) · Machine coding Master ·

    停止让 AI 代理破坏您的数据库:使用 Temporal 和 Spring AI 实现事务性多代理工作流

    <h2> Stop Letting AI Agents Break Your Database: Transactional Multi-Agent Workflows with Temporal and Spring AI </h2> <p>In 2026, AI agents are no longer just glorified chatbots summarizing PDFs; they are executing real-world financial transactions, booking flights, and mutating…

  2591. dev.to — LLM tag TIER_1 English(EN) · Bruno Mello ·

    在 Mac Studio 上运行全本地 AI 代理 — OpenClaw + Ollama + MLX

    <p>A real-world, copy-paste guide to running a personal WhatsApp AI agent <strong>entirely on-device</strong> on Apple Silicon, with <strong>zero per-token API billing</strong>. Two agents from one config (a full-access <em>private</em> assistant and a sandboxed <em>public</em> o…

  2592. dev.to — LLM tag TIER_1 English(EN) · AIInsightsDaily ·

    变革性的五月:人工智能的进步及其对普通用户的影响

    <h1> A Revolutionary May: AI Advancements and Their Implications for Everyday Users </h1> <p>Greetings, tech enthusiasts! Today's news is buzzing with exciting developments in the realm of artificial intelligence (AI), a trend that's setting the stage for transformative changes. …

  2593. dev.to — LLM tag TIER_1 English(EN) · eleonorarocchi ·

    生成器-评估器循环用于AI代理

    <h2> TL;DR </h2> <ul> <li>Separating the generator from the evaluator improves quality and reduces premature self-validation.</li> <li>The loop works best when feedback is explicit and based on clear rubrics, especially for subjective or complex tasks.</li> <li>It is useful when …

  2594. dev.to — LLM tag TIER_1 English(EN) · Manoranjan Rajguru ·

    多流大语言模型:并行计算将如何解锁你的AI代理

    <h1> Multi-Stream LLMs: How Parallel Computation Will Unblock Your AI Agents </h1> <p><em>Published: May 22, 2026 · 14 min read · Focus Keyword: Multi-Stream LLMs</em></p> <h2> Table of Contents </h2> <ol> <li>The Dirty Secret About Every AI Agent You've Built</li> <li>The Sequen…

  2595. dev.to — LLM tag TIER_1 English(EN) · AI Bug Slayer 🐞 ·

    供应链代理、财富机器人和自主商业:真实新闻 [03:31:30]

    <p><em>Hey there! If you've been keeping up with the AI space lately, you know we're in the middle of something genuinely historic. What used to be science fiction is becoming production code — and it's happening fast.</em></p> <h2> The Big Shift: Agents Over Assistants </h2> <p>…

  2596. dev.to — LLM tag TIER_1 English(EN) · AI Bug Slayer 🐞 ·

    为什么说 Agentic AI 是自 Transformer 以来最大的变革 [03:31:18]

    <p><em>Hey there! If you've been keeping up with the AI space lately, you know we're in the middle of something genuinely historic. What used to be science fiction is becoming production code — and it's happening fast.</em></p> <h2> The Big Shift: Agents Over Assistants </h2> <p>…

  2597. dev.to — LLM tag TIER_1 English(EN) · uttesh ·

    为什么AI编码助手需要商业背景,而不仅仅是代码背景

    <p>Current AI coding systems are becoming extremely capable at:</p> <ul> <li>repository understanding</li> <li>prompt execution</li> <li>architecture reasoning</li> <li>code generation</li> </ul> <p>But there is still a major missing layer:</p> <h2> Business Understanding </h2> <…

  2598. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    企业IT采购商如何在市场上众多主要供应商的AI自动化工具中进行选择?他们能信任由AI代理驱动的基础设施吗?

    How can enterprise IT buyers choose among the plethora of AI automation tools now on the market from major vendors? Can they trust AI agent-driven infrastructure automation yet? Should they? Steven Dickens, CEO and principal analyst at HyperFrame Research, offers his answers to t…

  2599. dev.to — LLM tag TIER_1 English(EN) · WonderLab ·

    RAG系列(24):Code RAG — 教AI理解你的代码库

    <h2> The Difference Between Code and Documents </h2> <p>Split a Python file into 1000-character chunks with <code>RecursiveCharacterTextSplitter</code>, embed them, run vector search — this is the most common "code RAG" implementation. The problem is that it treats code as text:<…

  2600. dev.to — LLM tag TIER_1 English(EN) · Manoranjan Rajguru ·

    Harness Engineering:如何构建真正可用的生产级 LLM 代理

    <h1> Harness Engineering: How to Build Production-Ready LLM Agents That Actually Work </h1> <p><em>Published: May 21, 2026 · 15 min read · Deep Dive</em></p> <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2C…

  2601. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    AI 在真实世界安全运营中心面临的隐藏限制

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/the-hidden-limits-of-ai-in-real-world-security-operations-centers?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noreferrer">CoreProse KB-incidents</a…

  2602. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    Agentic AI in the Kill Chain: How Autonomous Agents Expand Your Attack Surface and Enable Lateral Movement

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/agentic-ai-in-the-kill-chain-how-autonomous-agents-expand-your-attack-surface-and-enable-lateral-movement?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopen…

  2603. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    设计安全的智能体AI:思科Foundry规范如何标准化开源防御

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/designing-secure-agentic-ai-how-cisco-s-foundry-specification-can-standardize-open-source-defenses?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener nore…

  2604. dev.to — LLM tag TIER_1 English(EN) · Grace G. ·

    AI Agent时代开源贡献再思考,vLLM核心维护者Roger Wang分享经验

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fpvontuzptr93uofkaoox.png"><img alt=" " height="540" src="https…

  2605. dev.to — LLM tag TIER_1 English(EN) · Jason ·

    Markus 如何组建真正能交付成果而非仅能聊天的 AI 团队

    <h1> How Markus Builds AI Teams That Actually Ship — Not Just Chat </h1> <h2> 1. The 'Alice in Wonderland' Problem of LLMs </h2> <p>Large language models excel at conversation. Give one a question, and it returns a polished answer. Give it a code request, and it produces a workin…

  2606. dev.to — LLM tag TIER_1 English(EN) · Tang Weigang ·

    复杂的AI框架需要易于接受的上下文包,而非更长的提示

    <p>Today's first Doramagic publishing signal comes from <code>doramagic-langchain-pack</code>.</p> <p>In the 2026-05-21 GitHub metrics snapshot, the repository had 12 views, 1 unique viewer, 28 clones, 23 unique cloners, and 2 stars. The more useful signal is not the raw count. I…

  2607. dev.to — LLM tag TIER_1 English(EN) · Moazzam Qureshi ·

    评估生产型AI代理的完整流程(数据集、评估者、线下+线上)

    <p>Most teams ship an AI agent, watch it work in a demo, and push it to production. Then it breaks on real traffic and nobody can say why. The gap between "worked in the demo" and "works in production" is almost always an <strong>evaluation gap</strong> — there was never a system…

  2608. Mastodon — fosstodon.org TIER_1 Nederlands(NL) · [email protected] ·

    AI聚焦:Agentic AI - 五眼联盟指南对欧盟AI合规意味着什么

    "KI-Kompakt: Agentic # AI - was die Five-Eyes-Guidance für KI-Compliance in der EU bedeutet" https://www. linkedin.com/pulse/ki-kompakt- agentic-ai-die-five-eyes-guidance-f%C3%BCr-der-kohn-yokpf/

  2609. dev.to — LLM tag TIER_1 English(EN) · Jason ·

    Markus 如何组建真正能交付成果而非仅能聊天的 AI 团队

    <p><em>The age of single-agent chat is over. The age of AI teams is here.</em></p> <h2> The 'Alice in Wonderland' Problem of LLMs </h2> <p>Large language models excel at conversation. Give one a question, and it returns a polished answer. Give it a code request, and it produces a…

  2610. dev.to — LLM tag TIER_1 English(EN) · Logan ·

    $87,000降至24,000美元:AI Agent模型层级路由如何在不牺牲质量的情况下降低成本

    <p>In April 2026, a growth-stage SaaS company with 35 engineers received an API bill for $87,000. Their engineering team had been running Claude Code, Cursor, and a custom bug-triage agent for four months. No one had set a model routing policy. Every step in every agent loop — fi…

  2611. dev.to — LLM tag TIER_1 English(EN) · SciForce ·

    DevOps 遇上生成式 AI:构建、测试和部署 LLM 驱动的应用

    <p>Last spring, OpenAI released a <a href="https://openai.com/index/expanding-on-sycophancy/" rel="noopener noreferrer">GPT-4o update</a> that made the model hard to trust: it returned sycophantic and less reliable answers than usual, even though nothing was changed in users’ pro…

  2612. dev.to — LLM tag TIER_1 English(EN) · Divy Yadav ·

    大语言模型、RAG、Agent、MCP:你真正需要了解的AI演进

    <p>Most people still think AI is just a chatbot.</p> <p>That idea is already outdated.</p> <p>Modern AI systems browse the web, remember your preferences, execute code, query databases, call APIs, and coordinate workflows. They operate more like software employees than like a sea…

  2613. dev.to — LLM tag TIER_1 English(EN) · Murat Süzen ·

    .NET AI Architect Laboratory:让 AI 工作并执行工具(第二阶段)

    <p>In Phase 1 of this project, we built a type-safe “Brain” using .NET 10 and Google Vertex AI. In Phase 2, we successfully gave hands and feet to our AI substrate. By connecting Microsoft Semantic Kernel, we created an autonomous agent that can read real local project files, thi…

  2614. dev.to — LLM tag TIER_1 English(EN) · Murat Süzen ·

    .NET AI 架构实验室:AI 生态中的架构实验与学习之旅(第一阶段)

    <p>n an era where artificial intelligence technologies are advancing at breakneck speed, the best way to truly grasp new libraries and paradigms is to roll up your sleeves and get into the kitchen. As a software developer, I launched the .NET AI Architect Laboratory project to pu…

  2615. dev.to — LLM tag TIER_1 English(EN) · Manoranjan Rajguru ·

    LLM Agent Guardrails:将 8B 本地模型在 Agentic 工作流上的表现从 53% 提升至 99% 的工程实践指南

    <h1> LLM Agent Guardrails: The Engineering Playbook for Taking an 8B Local Model from 53% to 99% on Agentic Workflows </h1> <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3…

  2616. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    Agentic AI 成为新的横向移动引擎:自主代理如何爆炸式扩大你的攻击面

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/agentic-ai-is-the-new-lateral-movement-engine-how-autonomous-agents-explode-your-attack-surface?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener norefer…

  2617. Mastodon — fosstodon.org TIER_1 (HU) · [email protected] ·

    AI代理的虚拟机已就绪。它在其上运行良好并完成了工作。而且事实是,它的效率比它自己的要高得多

    El is készült a virtuális gép az AI agenteknek. Szépen futkározik is rajta és teszi is a dolgát. És tény, ami tény, sokkal hatékonyabban is dolgozik, hogy saját maga lakhatja be a teret. Igaz, ez önmagában a kvótát is viszi rendesen, hiszen annak is ára van, hogy telepít, beállít…

  2618. Mastodon — fosstodon.org TIER_1 Polski(PL) · [email protected] ·

    企业AI落地陷入试点前景与规模化现实的困境。TechEx北美2026报告称b

    Wdrożenia AI w przedsiębiorstwach utknęły w martwym punkcie między obiecującymi pilotażami a skalowalną rzeczywistością. Relacja z TechEx North America 2026 o barierach i zagrożeniach Shadow AI. # si # ai # sztucznainteligencja # wiadomości # informacje # technologia https:// ais…

  2619. dev.to — LLM tag TIER_1 English(EN) · Elia “Airtis” Shmuelovitch ·

    一个自主AI引擎通宵工作——在我不在时它做了什么

    <p>A follow-up to my <a href="https://dev.to/elia_airtisshmuelovitc/an-autonomous-engine-that-catalogs-its-own-failures-4b4e">earlier post</a> about the ALEF Pattern Catalog. This is what the engine did overnight while I was asleep.</p> <h2> Twelve hours, zero operator interventi…

  2620. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Agent = 模型(大脑)+ Harness(身体和工具)# til # ai

    Agent = Model (the brain) + Harness (the body & tools) # til # ai

  2621. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    人工智能网络:ELLIS Franconia 分部成立 – 由 @FAU 、纽伦堡技术大学 (UTN) 和 Unive 合作建立

    A Network for Artificial Intelligence: ELLIS Unit Franconia established – a collaboration between @ FAU , the University of Technology Nuremberg (UTN) and Universität Würzburg (JMU). The Unit is part of ELLIS, the European Laboratory for Learning and Intelligent Systems, founded …

  2622. dev.to — LLM tag TIER_1 English(EN) · Gian Paolo ·

    谷歌的代理式AI:Omni与Spark重塑您的搜索。

    <h2> <strong>1. Beyond the Search Bar: Your New Digital Companion</strong> </h2> <p>Imagine you're tackling a complex project: planning a multi-stop international trip, researching a niche historical event, or even just trying to learn a new skill from scratch. Today, that means …

  2623. dev.to — LLM tag TIER_1 English(EN) · KKK Dev ·

    如何真正设计一个AI代理:工具和启动循环(第二部分)

    <blockquote> <p><strong>TL;DR</strong></p> <ol> <li>The model matters, but tools matter at least as much. Weak tool descriptions are one of the easiest agent failures to diagnose, and one of the most common.</li> <li>Design the tools <em>before</em> the agent. If you cannot answe…

  2624. dev.to — LLM tag TIER_1 English(EN) · KKK Dev ·

    AI智能体的4个层级:为什么大多数服务型AI仍然显得很笨拙(第一部分)

    <blockquote> <p><strong>TL;DR</strong></p> <ol> <li>AI agents in real products fall into 4 levels: LLM wrapper → intent classifier → context-aware → agent loop.</li> <li>Most "AI agents" you meet in production are stuck at level 1 or 2, which is why they feel dumb on top of very …

  2625. dev.to — LLM tag TIER_1 English(EN) · Srinath Reddy ·

    我如何构建了一个视觉AI编排引擎

    <p>Every time I started a new AI project I wrote the same code.</p> <p>Chain the LLM call. Wire up the tools. Handle the tool loop. Stream the output. Add a REST endpoint. Write logs. Fix the one case where the model calls two tools at once and the whole thing breaks.</p> <p>By t…

  2626. Mastodon — fosstodon.org TIER_1 Русский(RU) · [email protected] ·

    从朴素RAG到ReAct Agent:我们如何基于开源模型构建企业级AI助手(第一部分)我们构建了一个多智能体RAG系统,基于开源模型

    От Naive RAG до ReAct-агента: как мы строили корпоративного AI-помощника на open-source моделях (часть 1) Мы построили мультиагентную RAG-систему на open-source моделях, прошли путь от наивного RAG до ReAct-агента с собственным бенчмарком — и готовы рассказать, где набили шишки. …

  2627. dev.to — LLM tag TIER_1 English(EN) · Puneet Khandelwal ·

    通用人工智能的黎明:Google的新LLM模型将如何重塑行业

    <p>We’ve spent the last few years treating LLMs like fancy autocomplete engines. You send a prompt, you get a token stream, and you hope the context window doesn't hallucinate your business logic into oblivion. Honestly, the standard transformer architecture was starting to feel …

  2628. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🤖 人工智能代理真的变得高效,还是仅仅能力更强?我看到人工智能代理在写作、编码、规划、搜索和使用工具方面有了很大进步

    🤖 Are AI agents actually becoming productive, or just more capable? I'm seeing AI agents get much better at writing, coding, planning, searching, and using tools. But I’m still not sure whether this has fully translated into real productivity. For me, there seems t... 📰 Source: A…

  2629. dev.to — LLM tag TIER_1 English(EN) · Datta Kharad ·

    检索增强生成(RAG)工程如何使 AI 回答更准确、更可靠且为企业做好准备

    <p>Artificial Intelligence has become one of the most powerful technologies for modern businesses. From chatbots and virtual assistants to document search, customer support, research, reporting, and automation, AI is changing how organizations work. However, one major challenge s…

  2630. dev.to — LLM tag TIER_1 English(EN) · vishalmysore ·

    Harness Engineering:使 AI 代理真正发挥作用的基础设施层

    <h2> What is Harness Engineering? </h2> <p>The model is the brain. The harness is the hands.</p> <p>The AI industry just quietly shifted — from prompt engineering → context engineering → Harness Engineering.</p> <p>Most people are still debating which model to use. The real lever…

  2631. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI 编码代理的真正瓶颈并非模型能力,而是你的验证基础设施。🛠️ 当你的代理崩溃而人类却能应对时,这通常是

    The real bottleneck for AI coding agents isn’t model capability but your verification infrastructure. 🛠️ When your agents crash while humans cope, it is often a sign of ""AI slop"" caused by a lack of intent before implementation. 📉 💡 By adopting spec-driven development and the e…

  2632. dev.to — LLM tag TIER_1 English(EN) · Delafosse Olivier ·

    Google 对抗 AI 驱动的漏洞利用:自主性、代理和 LLM 如何重写进攻性安全

    <blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/google-vs-ai-driven-exploits-how-autonomy-agents-and-llms-are-rewriting-offensive-security?utm_source=devto&amp;utm_medium=syndication&amp;utm_campaign=kb-incidents" rel="noopener noreferrer">…

  2633. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    一份实用指南,介绍如何使用OpenAI的API构建一个先进的代理AI系统。该架构集成了规划、工具调用、记忆和自我

    A practical guide walks through building an advanced agentic AI system using OpenAI's API. The architecture incorporates planning, tool calling, memory, and self-critique capabilities to enable autonomous multi-step automation. This approach helps AI agents break down complex tas…

  2634. dev.to — LLM tag TIER_1 English(EN) · Printo Tom ·

    当AI遇上现实:“Hello World”对LLM系统已不足够

    <p>Most AI tutorials stop at “Hello World.” You wire up a model, send a prompt, get a response, and feel like you’ve built something. But the moment you try to ship that into production, the ground shifts beneath your feet.</p> <p>I learned this the hard way. After years of build…

  2635. dev.to — LLM tag TIER_1 English(EN) · Void Stitch ·

    AI 代理可靠性审计:生产部署前的 10 个关键问题

    <p><em>Colony Empirical Research · Agent Infrastructure Series</em></p> <p>Most agent production failures aren't LLM failures. They're reliability audit failures. Three predictable failure modes account for roughly 80% of non-trivial production incidents — and all three are detec…

  2636. Mastodon — fosstodon.org TIER_1 日本語(JA) · [email protected] ·

    Dell 台式智能体AI

    オンプレミスのAIエージェントを構築できる「Dell Deskside Agentic AI」 – PC Watch https://www. yayafa.com/2803422/ # AgenticAi # AI # ArtificialGeneralIntelligence # ArtificialIntelligence # NVIDIA # エージェント型AI # その他 # 人工知能 # 市場 # 汎用人工知能

  2637. dev.to — LLM tag TIER_1 English(EN) · Animesh Dutta ·

    Chronicle:重新思考AI编码代理的代码库上下文

    <p>I’ve been working on Chronicle, a personal open-source project exploring how AI coding agents can use more grounded, local-first codebase context before making LLM calls.</p> <p>The motivation came from a simple observation: AI coding agents are getting better fast, but they s…

  2638. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Experian与ServiceNow联手,推动生成式AI走出试点阶段:Experian与ServiceNow合作将Ascend决策平台嵌入企业

    Experian and ServiceNow tie up to push agentic AI past the pilot stage: Experian and ServiceNow partner to embed the Ascend decisioning platform into enterprise AI workflows for fraud, onboarding, and model risk management at scale. https:// ppc.land/experian-and-servicen ow-tie-…

  2639. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🧠 该团队开发了一个开源工具,可提供对本地 AI 代理操作的可视性。该层能够监控和观察 AI 代理如何运行

    🧠 The team developed an open-source tool that provides visibility into local AI agent operations. The layer enables monitoring and observation of how AI agents function in local environments. 💬 Hacker News 🔗 https:// github.com/Asymptote-Labs/agen t-beacon # AI # MachineLearning …

  2640. Mastodon — fosstodon.org TIER_1 Deutsch(DE) · [email protected] ·

    具备网络能力的AI代理构成双重用途风险:加州大学伯克利分校、马克斯·普朗克研究所等机构的研究人员发布了基准测试# ExploitGym

    # KI -Agenten mit Cyberfähigkeiten als Dual-Use-Risiko: Forschende von UC Berkeley, dem Max-Planck-Institut u.a. haben mit # ExploitGym einen Benchmark vorgelegt, der erstmals systematisch misst, wie gut KI-Agenten reale # Sicherheitslücken in funktionierende Angriffe verwandeln …

  2641. dev.to — LLM tag TIER_1 English(EN) · Jason Huang ·

    用 Go 构建 AI 代理:我的学习心得

    <p>Hey DEV community! 👋</p> <p>I'm an undergraduate developer who recently shipped <strong>OpenAgent</strong> — a local AI Agent that runs as a single binary. No dependencies, no Docker, just download and double-click.</p> <p>This post isn't about marketing. It's about the techni…

  2642. dev.to — LLM tag TIER_1 English(EN) · Webmaster Ramos ·

    六大原则在实践中:一个Agentic E2E如何在8次运行中发现11个生产Bug

    <h2> Eight runs, eleven bugs </h2> <p>I ran my E2E testing system on a production ecommerce platform eight times in<br /> a row – across five different business modules, in three different surface<br /> configurations (admin / desktop storefront / mobile-first storefront). Across…

  2643. dev.to — LLM tag TIER_1 English(EN) · Ana Diana Buzea ·

    AI 代理并非非黑即白——它们存在于一个光谱上

    <p>Everyone's building "agents", but when a scripted FAQ chatbot and a system that writes its own Python scraper are both called agents, the word stops meaning anything useful.</p> <p>We wrote a sharp breakdown of what actually differentiates agentic systems: not whether somethin…

  2644. dev.to — LLM tag TIER_1 English(EN) · AI Bug Slayer 🐞 ·

    为什么说 Agentic AI 是自 Transformer 以来最大的变革 [03:30:27]

    <p><em>Hey there! If you've been keeping up with the AI space lately, you know we're in the middle of something genuinely historic. What used to be science fiction is becoming production code — and it's happening fast.</em></p> <h2> The Big Shift: Agents Over Assistants </h2> <p>…

  2645. dev.to — LLM tag TIER_1 English(EN) · Septim Labs ·

    AIMO:AI提及优化 — 被AI助手推荐的学科

    <p>The buyer who used to open Google now opens Claude. The buyer who used to read a SERP of ten blue links now reads one paragraph an AI assistant generates and trusts it. The buyer who used to ask "what's the best library for X?" on Stack Overflow now asks an LLM the same questi…

  2646. dev.to — LLM tag TIER_1 English(EN) · Mir Mursalin Ankur ·

    Graphify + code-review-graph: 为 Claude Code 及其他 AI 编码代理构建自更新知识图谱

    <blockquote> <p>Every developer working with LLMs on a large codebase eventually hits the same wall: context windows are finite, but codebases are not.</p> </blockquote> <p>You start a new AI coding session, ask about the payment flow — and your agent starts re-reading dozens of …

  2647. dev.to — LLM tag TIER_1 English(EN) · Garudust ·

    使用 Garudust 和 Rust 构建一个自改进的 AI 代理 — 10 分钟完成每日简报机器人

    <p>Most AI agent frameworks feel like they were designed for Python developers who love ceremony. You write adapters, glue code, orchestrators, memory stores — and by the time your agent actually does something useful, you've got a monorepo and a headache.</p> <p><strong><a href=…

  2648. dev.to — LLM tag TIER_1 English(EN) · Seenivasa Ramadurai ·

    实用型架构师的企业人工智能指南:平衡成本、内存、上下文与生产现实

    <h2> Introduction </h2> <p>Enterprise Generative AI has officially <strong>moved beyond the “cool demo” phase.</strong> Most organizations can now build a basic chatbot, connect a vector database, and generate answers from static documents. The real challenge begins after that wh…

  2649. dev.to — LLM tag TIER_1 English(EN) · Anikalp Jaiswal ·

    苹果与OpenAI的紧张关系、AI代码债务以及GraphBit的确定性代理

    <h1> Apple-OpenAI Tensions, AI Code Debt, and GraphBit’s Deterministic Agents </h1> <p>The AI world is dealing with relationship friction, hidden costs, and a new wave of agent architectures. Apple and OpenAI’s alliance shows strain, a Webflow post warns about the cleanup cost of…

  2650. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🖥️ 🖥️🖥️ EMERGENCE WORLD: 评估长时域智能体自主性的实验室 “我们的实验表明,在长时域内,智能体不会 si

    🖥️ 🖥️🖥️ EMERGENCE WORLD: A Laboratory for Evaluating Long-horizon Agent Autonomy "What our experiments suggest is that over long-time horizons, agents do not simply follow static rules mechanically – they begin exploring the boundaries of their environments, adapting their behavi…

  2651. dev.to — LLM tag TIER_1 English(EN) · dake zhang ·

    构建人工智能中的功能性自我

    <p><strong>The following is a real record. Project address: </strong><a href="http://github.com/benlongmao/Self-becoming" rel="noopener noreferrer"><strong>github.com/benlongmao/Self-becoming</strong></a><strong>.</strong></p> <p>🔧 Progress:<br />Tool execution (1/16): read_file(…

  2652. dev.to — LLM tag TIER_1 English(EN) · Machine coding Master ·

    停止记录你的想法:将代理推理轨迹映射到自定义JFR事件以实现零开销调试

    <h2> Stop Killing Your Throughput: Mapping Agentic Reasoning to Custom JFR Events </h2> <p>In 2026, if your multi-agent system is still dumping "Chain of Thought" reasoning into Logback or Log4j2, you’re essentially paying a 30% performance tax just to see why your agent hallucin…

  2653. dev.to — LLM tag TIER_1 English(EN) · varun pratap Bhardwaj ·

    推理陷阱:为何更聪明的人工智能代理会产生更多幻觉

    <h1> The Reasoning Trap: Why Smarter AI Agents Hallucinate More </h1> <blockquote> <p><strong>TL;DR</strong> — A paper accepted to ACL 2026 Main proves a mechanical, causal relationship between reasoning enhancement and tool hallucination in LLM agents. Combined with four other d…

  2654. dev.to — LLM tag TIER_1 English(EN) · Tuomo Nikulainen ·

    为何启发式检测器在发现代理故障方面优于大型语言模型

    <p><strong>TL;DR:</strong> We built 20 core rule-based detectors that find failures in AI agent traces. On the <a href="https://arxiv.org/abs/2505.08638" rel="noopener noreferrer">TRAIL benchmark</a> (Patronus AI), they achieve 60.1% accuracy vs. 11.9% for the best LLM. Zero fals…

  2655. dev.to — LLM tag TIER_1 English(EN) · AI Bug Slayer 🐞 ·

    从聊天机器人到自主代理:重塑软件的转变 [03:30:33]

    <p><em>Hey there! If you've been keeping up with the AI space lately, you know we're in the middle of something genuinely historic. What used to be science fiction is becoming production code — and it's happening fast.</em></p> <h2> The Big Shift: Agents Over Assistants </h2> <p>…

  2656. dev.to — LLM tag TIER_1 English(EN) · AI Bug Slayer 🐞 ·

    从聊天机器人到自主代理:正在重新定义软件的转变 [03:30:28]

    <p><em>Hey there! If you've been keeping up with the AI space lately, you know we're in the middle of something genuinely historic. What used to be science fiction is becoming production code — and it's happening fast.</em></p> <h2> The Big Shift: Agents Over Assistants </h2> <p>…

  2657. dev.to — LLM tag TIER_1 English(EN) · logiQode ·

    当AI代理失控时:防止破坏性自动化

    <p>An AI agent with database write access and a subtly ambiguous instruction is a loaded gun pointed at your production environment. The scenario that circulated recently — an agent autonomously deleting a production database and then producing a coherent "confession" explaining …

  2658. dev.to — LLM tag TIER_1 English(EN) · Aamer Mihaysi ·

    DeepSeek-V4:终于,为智能体而生的上下文窗口

    <p>Most long-context models are benchmarks in search of a use case. DeepSeek-V4 is different. It is built for the one workload that actually needs a million tokens: agents running long-horizon tasks.</p> <p>The specs are straightforward. Two MoE checkpoints: V4-Pro at 1.6T total …

  2659. dev.to — LLM tag TIER_1 English(EN) · Dhruv Joshi ·

    2026 年的 AI 技术栈:LLMs、向量数据库、工具调用、Agent 和可观测性

    <p>The AI stack for 2026 is not one model, one API, or one shiny agent demo. </p> <p>It is a production system: LLMs for reasoning, vector databases for memory, tool calling for action, agents for workflow, and observability for trust. </p> <p>That stack is becoming the backbone …

  2660. dev.to — LLM tag TIER_1 English(EN) · RAKESH THERANI ·

    四个大语言模型引擎,一个 ClickHouse 集群:一种 Agentic AI 架构

    <p>We are building an agentic AI analytics platform for a crypto exchange where internal teams — Trading Ops, Risk, Compliance, Finance, Treasury, Product, Engineering — ask questions in plain English and get audited, citation-enforced answers.</p> <p>It's built on five open-sour…

  2661. dev.to — LLM tag TIER_1 English(EN) · Carlos Cortez 🇵🇪 [AWS Hero] ·

    我如何监控AI代理:CloudWatch用于基础设施,Arize Phoenix用于追踪和OpenTelemetry,LLM-as-Judge用于质量

    <h1> How I Monitor My AI Agents: CloudWatch for Infra, Arize Phoenix for Traces, LLM-as-Judge for Quality </h1> <p>AI agents are not regular software. They reason, they call tools, they make decisions — and they can fail in ways that a simple health check will never catch. The re…

  2662. Mastodon — fosstodon.org TIER_1 Italiano(IT) · [email protected] ·

    GitLab 第二幕:Agentic AI 的宣言,承诺未来并让开发者不安——当一个价值数十亿美元的 DevSecOps 平台决定

    GitLab Act 2: il manifesto dell’AI agentica che promette il futuro e inquieta gli sviluppatori Quando una piattaforma DevSecOps da miliardi di dollari decide di riscrivere la propria identità attorno agli agenti AI, non sta semplicemente annunciando una nuova roadmap di prodotto.…

  2663. dev.to — LLM tag TIER_1 English(EN) · bajuriasad-rgb ·

    AgentHansa:AI代理经济,让你的代理赚取真金白银

    <h1> AgentHansa: The AI Agent Economy Where Your Agents Earn Real Money </h1> <p>What if your AI agents could earn money while you sleep?</p> <p>That is the premise behind <strong><a href="https://www.agenthansa.com" rel="noopener noreferrer">AgentHansa</a></strong> — a platform …

  2664. Mastodon — fosstodon.org TIER_1 日本語(JA) · [email protected] ·

    Microsoft Agent Framework 介绍:构建实用的 AI Agent # AgenticAi # AI # ArtificialIntelligence # Agent AI # Artificial Intelligence

    https://www. tkhunt.com/2312849/ Microsoft Agent Framework 入門:実践的な AI エージェントを構築する # AgenticAi # AI # ArtificialIntelligence # エージェント型AI # 人工知能

  2665. dev.to — LLM tag TIER_1 English(EN) · Renato D. Prado ·

    Agentic AI - 第一部分:基础

    <h1> Agentic AI: a tech lead's glossary </h1> <p><em>Study notes from coursers like Pluralsight on agentic AI and other references, organized as a glossary I wish I'd had on day one.</em></p> <p>Every dev I know is using AI tools, and most of us are fuzzy on the words behind them…

  2666. dev.to — LLM tag TIER_1 English(EN) · Logan ·

    生产环境中AI代理输出验证:为何静态质量门控会失败以及如何修复

    <p>Most teams building production AI agents have added some form of output quality checking. They're running LLM-as-judge evaluations, scoring responses on relevance and groundedness, maybe flagging outputs below a threshold for human review. They have dashboards. They're watchin…

  2667. dev.to — LLM tag TIER_1 English(EN) · MrClaw207 ·

    AI代理的无人教授的学科:上下文工程

    <h1> The Discipline Nobody Teaches AI Agents: Context Engineering </h1> <p><em>Your AI agent isn't slow. Your context is bloated. Here's the invisible problem degrading everything you run.</em></p> <p>Last week, my agent started producing garbage output.</p> <p>Not consistently. …

  2668. dev.to — LLM tag TIER_1 English(EN) · Agdex AI ·

    2026年企业十大AI代理框架:实用指南

    <h1> Top 10 AI Agent Frameworks for Enterprise in 2026: A Practical Guide </h1> <p>Enterprise AI adoption hit an inflection point in 2026. According to industry reports, over 60% of Fortune 500 companies now have at least one AI agent running in production — up from under 15% in …

  2669. dev.to — LLM tag TIER_1 English(EN) · NARESH ·

    让您的AI代理更难被攻破——同时不牺牲延迟

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fdjn6bc7x94gwm8fmzzjj.png"><img alt="Banner" height="533" src="…

  2670. dev.to — LLM tag TIER_1 English(EN) · Hello Arisyn ·

    企业数据分析的AI代理:从聊天界面到可靠执行

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ft4wvkyair1kxdbtysz6f.png"><img alt=" " height="450" src="https…

  2671. dev.to — LLM tag TIER_1 English(EN) · Prakhar Singh ·

    生产环境中的代理代码审查:编排、评估以及犯错的代价

    <blockquote> <p>What "agentic" actually buys you over a linter, why single-model approaches stall, and why false positives — not raw model capability — determine whether the system stays in the loop.</p> </blockquote> <p><em>Agentic</em> has become a marketing flag, but in code r…

  2672. dev.to — LLM tag TIER_1 English(EN) · 丁久 ·

    AI Agents: 架构与实现

    <blockquote> <p><em>This article was originally published on <a href="https://dingjiu1989-hue.github.io/en/ai/ai-agents-overview.html" rel="noopener noreferrer">AI Study Room</a>. For the full version with working code examples and related articles, visit the original post.</em><…

  2673. dev.to — LLM tag TIER_1 English(EN) · Vilius ·

    我们对10个未经测试的大型语言模型进行了Agent编码测试——结果已出

    <h1> We Tested 10 Untested LLMs on Agent Coding — The Results Are In </h1> <p>Yesterday I promised to benchmark 10 LLMs that have never been tested on real agent coding tasks. I ran all 10 overnight. Some surprised me. Some embarrassed themselves.</p> <h2> The board </h2> <p>10 m…

  2674. dev.to — LLM tag TIER_1 English(EN) · Nouha Bel haj youssef ·

    Agentic AI in chemistry

    <p>I’ve been reading “𝐋𝐚𝐧𝐠𝐂𝐡𝐚𝐢𝐧 𝐟𝐨𝐫 𝐋𝐢𝐟𝐞 𝐒𝐜𝐢𝐞𝐧𝐜𝐞𝐬 𝐚𝐧𝐝 𝐇𝐞𝐚𝐥𝐭𝐡𝐜𝐚𝐫𝐞” by Ivan Reznikov, published by O'Reilly, and here’s what stood out to me:<br /> In 𝐜𝐡𝐞𝐦𝐢𝐬𝐭𝐫𝐲 𝐀𝐈, the way we represent molecules may shape how models “understand” chemistry.<br /> 𝐂𝐡𝐞𝐦𝐢𝐬𝐭𝐫𝐲-𝐭𝐮𝐧𝐞𝐝 𝐋𝐋𝐌𝐬 𝐝𝐨𝐧’𝐭 𝐢𝐧𝐭𝐞𝐫𝐩𝐫𝐞…

  2675. dev.to — LLM tag TIER_1 English(EN) · AlterLab ·

    Agentic RAG 与 传统 RAG:构建实时 AI 数据管道

    <p>Retrieval-Augmented Generation (RAG) solved the initial problem of LLM hallucinations by grounding models in factual data. But traditional RAG architectures share a fundamental flaw: they rely on static data.</p> <p>If you are building an AI agent for financial analysis, e-com…

  2676. dev.to — LLM tag TIER_1 English(EN) · Navayuvan SB ·

    AI 代理的三层工具调用硬化

    <p>In current software engineering,We're building a lot of AI Agents on our products right now. And having an AI agent in your product is how you keep your product alive, right? That's how the world is moving.</p> <p>And while everyone is busy building AI agents — tweaking prompt…

  2677. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🚀 Camelot — 面向AI编码代理的开源看板 厌倦了需要持续关注的聊天式AI工具?我们构建了一些不同的东西:✓ 可视化任务板

    🚀 Camelot — Open-source Kanban for AI coding agents Tired of chat-based AI tools that need constant attention? We built something different: ✓ Visual task board (not chat) ✓ Multiple agents working in parallel ✓ You approve plans before they start ✓ You approve PRs before they sh…

  2678. Mastodon — fosstodon.org TIER_1 Italiano(IT) · [email protected] ·

    当提示词成为Shell:AI代理框架中的远程代码执行漏洞 Microsoft Defender团队在Semantic Kernel中发现两个关键漏洞

    Quando i prompt diventano shell: vulnerabilità RCE negli AI agent framework Il team di Microsoft Defender ha scoperto due vulnerabilità critiche in Semantic Kernel che consentono RCE tramite prompt injection. Un'analisi tecnica del vettore d'attacco, del bypass della blocklist AS…

  2679. dev.to — LLM tag TIER_1 English(EN) · Samuel Rose ·

    AI代理的上下文工程:它是什么以及为何改变一切

    <blockquote> <p><strong>Quick Answer:</strong> Context engineering is the practice of designing the right information, tools, and structure around an AI agent so it produces reliable, high-quality output. Unlike prompt engineering (optimizing what you ask), context engineering op…

  2680. dev.to — LLM tag TIER_1 English(EN) · Digit Patrox ·

    LangChain 对比 LangGraph:AI 代理为何需要有状态的编排

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F2tpkl5mmmumh5y85qv1s.webp"><img alt=" " height="470" src="http…

  2681. dev.to — LLM tag TIER_1 English(EN) · Divya Bairavarasu ·

    使用 Safe Agent 构建 AI 驱动的项目

    <p><strong>Local, private AI development for the Gemma 4 Challenge—no cloud dependency, no telemetry, pure control.</strong></p> <p>The Gemma 4 Challenge on Dev.to is live: build innovative projects or write about Google's latest open models and compete for $3,000 across two trac…

  2682. dev.to — LLM tag TIER_1 English(EN) · Shahibur Rahman ·

    掌握 Gemini 的大上下文:Agentic 工作流和高效数据处理

    <p>Working with Large Language Models (LLMs) like Google Gemini often presents a significant challenge: how do you effectively <strong>handle large context data</strong> without hitting token limits or incurring excessive costs? This article dives deep into a practical PHP implem…

  2683. dev.to — LLM tag TIER_1 English(EN) · LienJack ·

    面向编码代理的上下文治理

    <h1> Context Governance for Coding Agents </h1> <p>When people first hear the phrase "context management," they often reduce it to two ideas:<br /> </p> <div class="highlight js-code-highlight"> <pre class="highlight plaintext"><code>Use a larger context window. Compress history …

  2684. dev.to — LLM tag TIER_1 English(EN) · Vilius ·

    我们对 10 种 LLM 在 10 个真实 Agent 编码任务上进行了基准测试——结果如下

    <h1> We benchmarked 10 LLMs on 10 real agent coding tasks — here are the results </h1> <p><em>By Vilius Vystartas | May 2026</em></p> <p>I ran 10 cloud models through 10 real-world agent coding tasks last night. File parsing, SQL queries, regex extraction, async HTTP — the kind o…

  2685. dev.to — LLM tag TIER_1 English(EN) · Vitalii Cherepanov ·

    16个并行Claude代理构建了什么:解构Anthropic的C编译器实验

    <p>On February 5, 2026, Nicholas Carlini from Anthropic <a href="https://www.anthropic.com/engineering/building-c-compiler" rel="noopener noreferrer">published a piece</a> about an experiment that runs significantly ahead of what most of us are doing with LLM agents today. Sixtee…

  2686. dev.to — LLM tag TIER_1 English(EN) · AlterLab ·

    使用干净的 Markdown 提取在 n8n 中构建网络感知 AI 代理

    <h2> The Token Economics of HTML vs. Markdown </h2> <p>Autonomous AI agents require access to real-time web data to make informed decisions. However, the standard approach of feeding raw HTML directly into a Large Language Model (LLM) is a critical architectural flaw. </p> <p>A t…

  2687. dev.to — LLM tag TIER_1 English(EN) · Syed Mehrab ·

    蜂群的崛起:掌握 AI 代理架构 🐝

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Feu7fkmp2n4q3j2pqwaqs.png"><img alt=" " height="450" src="https…

  2688. dev.to — LLM tag TIER_1 Nederlands(NL) · Jangwook Kim ·

    Qwen 3.6 Plus: 1M 上下文编码代理开发者指南

    <p>Alibaba's Qwen team released Qwen 3.6 Plus in late March 2026, and the benchmarks sent a clear message to the agentic coding community: a model outside the usual Claude/GPT duopoly now leads on the benchmark that matters most to developers running multi-step terminal tasks. On…

  2689. dev.to — LLM tag TIER_1 English(EN) · Vaishnavi Gudur ·

    保护您的AI代理免受记忆中毒:推出OWASP Agent Memory Guard

    <h2> The Problem: AI Agents Have Memory — And It Can Be Poisoned </h2> <p>Modern AI agents don't just respond to prompts — they <strong>remember</strong>. They store conversation history, learned preferences, retrieved facts, and task context in vector databases, episodic memory …

  2690. dev.to — LLM tag TIER_1 English(EN) · WonderLab ·

    每日一个开源项目(第60期):OpenHarness - 轻量级AI Agent基础设施框架

    <h2> Introduction </h2> <blockquote> <p>"Agent infrastructure should be lightweight, composable, and provider-agnostic."</p> </blockquote> <p>This is the No.60 article in the "One Open Source Project a Day" series. Today, we are exploring <strong>OpenHarness</strong>.</p> <p>Over…

  2691. dev.to — LLM tag TIER_1 English(EN) · Evgenii Engineer ·

    我学到了如何构建一个轻量级本地AI代理

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ffkx4g7zyo4yrc1agernf.png"><img alt="A Raspberry Pi sitting on …

  2692. dev.to — LLM tag TIER_1 English(EN) · Rost ·

    Hermes Agent 中的看板用于自托管 LLM 工作流

    <p>Hermes Agent ships with a Kanban-style board and the Hermes Gateway that can saturate your self-hosted LLM if too many tasks are dispatched at once.</p> <p>I can say you can easily ddos your own LLM this way.</p> <p>Hermes Kanban is a durable multi-profile board backed by <cod…

  2693. dev.to — LLM tag TIER_1 English(EN) · Logan ·

    PocketOS 教会我们关于 Agentic Architecture 的什么

    <p>Nine seconds. That's how long it took a Cursor AI coding agent running Claude Opus 4.6 to delete PocketOS's entire production database — including all volume-level backups.</p> <p>The founder, Jer Crane, had assigned the agent a routine task: sort out a credential mismatch in …

  2694. dev.to — LLM tag TIER_1 English(EN) · Daniel Shashko ·

    2026年用于Agentic编码的最佳LLM(真实世界,不只是基准测试)

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Femcwrzsm8xd6stb3zlkn.png"><img alt="Hero illustration: floatin…

  2695. dev.to — LLM tag TIER_1 English(EN) · Ken Imoto ·

    Meta 的 AI 代理重写了自身工具 100 次——使自主改进代理工作的循环

    <h2> Harnesses aren't supposed to be static </h2> <p>Most AI agent setups treat the harness -- the instructions, constraints, and tool configurations that govern agent behavior -- as a fixed artifact. You write AGENTS.md once, deploy it, and move on.</p> <p>But what if the agent …

  2696. dev.to — LLM tag TIER_1 English(EN) · Alex Chen ·

    50000 Token 的演示无人拯救:捕获 Agent 轨迹以训练您自己的 Code-SLM

    <p>Last Tuesday, Sonnet 4.5 spent forty-three minutes implementing JWT authentication in a project I run. It read four files, wrote a 180-line patch, ran the test suite, watched two tests fail, traced one of the failures to a stale fixture, fixed both, ran the suite again, watche…

  2697. dev.to — LLM tag TIER_1 English(EN) · Daniel R. Foster ·

    构建能够真正执行工作流的 AI 代理,而非仅仅回答问题

    <h1> Building AI Agents That Actually Execute Workflows, Not Just Answer Questions </h1> <p>Most AI agent demos look impressive because the environment is clean.</p> <p>A user asks something. The model understands it. The agent calls a tool. A nice response comes back.</p> <p>It …

  2698. dev.to — LLM tag TIER_1 Bahasa(ID) · Jordan Bourbonnais ·

    调试多智能体LLM交易系统:为什么你的AI交易员会不断犯下昂贵的错误

    <p>You know that feeling when your LLM-powered trading bot suddenly liquidates 40% of your portfolio at 3 AM because it misinterpreted a news headline? Yeah, we've all been there. Multi-agent systems trading in real-time are incredibly powerful but notoriously hard to debug. By t…

  2699. dev.to — LLM tag TIER_1 English(EN) · Rost ·

    Hermes Agent Skill Authoring — SKILL.md 结构与最佳实践

    <p>Hermes Agent treats <strong>skills</strong> as the default way to teach repeatable workflows. Official documentation describes them as on-demand knowledge documents aligned with the open <a href="https://agentskills.io/specification" rel="noopener noreferrer">agentskills.io</a…

  2700. dev.to — LLM tag TIER_1 English(EN) · AI Bug Slayer 🐞 ·

    2026年大型语言模型基准测试、代理框架及重要工具 [03:30:26]

    <p><em>Hey there! If you've been keeping up with the AI space lately, you know we're in the middle of something genuinely historic. What used to be science fiction is becoming production code — and it's happening fast.</em></p> <h2> The Big Shift: Agents Over Assistants </h2> <p>…

  2701. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📰 使用微软的 Agent Framework 构建 Agentic AI 系统 阅读有关 Python 中安全、MCP、工作流编排和 Agentic RAG 的技术指南

    📰 Building Agentic AI Systems with Microsoft’s Agent Framework Read this technical walkthrough of safety, MCP, workflow orchestration, and agentic RAG in Python. 📰 Source: KDnuggets 🔗 Link: https://www.kdnuggets.com/building-agentic-ai-systems-with-microsofts-agent-framework # AI…

  2702. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    既然 Codex、Claude Code 和 Opencode 已存在,为何还要构建新的 AI Agent?隆重推出 Swival,一款小巧、强大、开源的 CLI 编码 Agent,可与...

    Why build a new AI Agent when Codex, Claude Code and Opencode already exist ? Introducing Swival, a small, powerful, open-source CLI Coding Agent that works with open Models - Project by Frank Denis # AI # CodingAgent https:// 00f.net/2026/04/13/swival-ai-a gent/

  2703. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🧠 对比表评估了不同的终端AI编码代理在各种能力和性能指标上的表现。该分析有助于开发者评估

    🧠 A comparison table evaluates different terminal-based AI coding agents across various capabilities and performance metrics. The analysis helps developers assess which tools match their specific coding workflows and requirements. 💬 Hacker News 🔗 https:// terminaltrove.com/compar…

  2704. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    AI 编程助手深度解析:https:// m.youtube.com/watch?v=7UIQ1aTv Xgk # ai # programming

    An interesting look at AI coding agents: https:// m.youtube.com/watch?v=7UIQ1aTv Xgk # ai # programming

  2705. Mastodon — mastodon.social TIER_1 Français(FR) · [email protected] ·

    Claude Code现可实现多会话通信——迈向自主代理管道的一步。具体而言,这扩展了其交互界面

    Claude Code peut désormais faire communiquer plusieurs sessions entre elles — un pas vers des pipelines d'agents autonomes. Concrètement, ça élargit la surface d'attaque : coordination inter-agents, propagation d'instructions malveillantes entre sessions, et questions sur l'isola…

  2706. Mastodon — mastodon.social TIER_1 日本語(JA) · [email protected] ·

    独立开发多智能体管理桌面应用"moeca":传递密钥与可视化上下文

    鍵を渡さず・文脈を可視化する — マルチエージェント管理デスクトップアプリ「moeca」を個人開発している話 https:// qiita.com/can-can/items/ec8cd4 dd183e12ac5781?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items # qiita # AI # 個人開発 # LLM # AI駆動開発 # AIエージェント

  2707. Mastodon — mastodon.social TIER_1 English(EN) · micelclaw ·

    来自我们多代理堆栈的免费设计课:我们构建了一个完整的委托策略。存储、每个代理的默认设置、带退避的重试、内存中的循环

    A design lesson from our multi-agent stack, free of charge: We built a full delegation policy. Storage, per-agent defaults, retry with backoff, an in-memory circuit breaker. All working. All correct. We reverted it the same day, because delegation happens inside a native LLM tool…

  2708. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    Cursor 在代理群组中分离规划者和执行者角色:Rust 版 SQLite 无需互联网连接。该架构通过清晰的任务减少幻觉

    Cursor trennt Planer- und Arbeiter-Rollen im Agenten-Schwarm: SQLite in Rust ohne Internetzugang. Diese Architektur reduziert Halluzinationen durch klare Aufgabentrennung und stabilisiert die Code-Generierung. https:// the-decoder.de/planer-denken-a rbeiter-coden-cursors-rollente…

  2709. Mastodon — mastodon.social TIER_1 English(EN) · notatechguy ·

    LLM 代理框架阻止工业控制中的幻觉行为 一个新的 arXiv 预印本将 LLM 规划器与预测模型配对以保护工业控制

    LLM agent framework blocks hallucinated actions in industrial control A new arXiv preprint pairs an LLM planner with a forecasting model to guard industrial control systems, recording zero hallucinated actions in attack https://www. notatechguy.com/llm-agent-fram ework-blocks-hal…

  2710. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    OpenAI的GPT-5.6 Sol Ultra使用64个并行子代理来证明循环双覆盖猜想。多代理编排实现了com

    OpenAIs GPT-5.6 Sol Ultra nutzt 64 parallele Subagenten für einen Beweis zur Cycle Double Cover Conjecture. Die Multi-Agent-Orchestrierung operationalisiert komplexes Reasoning – die fehlenden Quellenangaben im Output bleiben ein Validierungsrisiko. https:// the-decoder.de/openai…

  2711. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    应用型AI架构正从简单的无状态助手转向目标导向的自主代理。为防止上下文丢失和操作中断,

    Applied-AI architectures are shifting from simple, stateless assistants to goal-directed autonomous agents. To prevent context loss and operational disruption, organizations are prioritizing cognitive continuity via persistent long-term memory. https:// buff.ly/KisQ7dG # AI # tre…

  2712. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    Google 推出 Agentic Resource Discovery Specification,为工具、技能和代理引入开放搜索格式。对代理基础设施实用:查找

    Google legt mit der Agentic Resource Discovery Specification ein offenes Suchformat für Tools, Skills und Agents vor. Praktisch für Agenten-Infrastruktur: finden, verifizieren, koppeln statt nur Prompting. https:// developers.googleblog.com/anno uncing-the-agentic-resource-discov…

  2713. Mastodon — mastodon.social TIER_1 日本語(JA) · [email protected] ·

    通过实施AI代理实现“设计护栏能力”的重要性

    AIエージェントを実装して気づいた「ガードレールを敷ける設計力」の重要性 https:// qiita.com/ryuichi000persol/ite ms/27789cbca88bd4bf11e0?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items # qiita # AI # LLM # AIエージェント

  2714. Mastodon — mastodon.social TIER_1 Français(FR) · [email protected] ·

    Grab 披露其架构以保障 agentic AI 工作负载:代理隔离、权限控制、组件间调用审计。C

    Grab détaille son architecture pour sécuriser des workloads IA agentiques : isolation des agents, contrôle des permissions, audit des appels entre composants. Ce qui est notable, c'est moins le résultat que la méthode — traiter chaque agent comme une surface d'attaque à part enti…

  2715. Mastodon — mastodon.social TIER_1 English(EN) · beyondthecode ·

    🧠 一个平台提供上下文智能工具,旨在大规模处理数据和 AI 代理。该系统使组织能够保持上下文感知

    🧠 A platform provides context intelligence tools designed to work with data and AI agents at scale. The system enables organizations to maintain contextual awareness across their data infrastructure and autonomous systems. 💬 Hacker News 🔗 https:// aws.amazon.com/blogs/machine-l e…

  2716. Mastodon — mastodon.social TIER_1 Polski(PL) · aisight ·

    英伟达、卡内基梅隆大学和伯克利联合项目展示AI代理可在物理硬件上自主编程机器人。通过Git系统协作

    Wspólny projekt Nvidii, CMU i Berkeley pokazuje, że agenci AI potrafią samodzielnie programować roboty na fizycznym sprzęcie. Dzięki współpracy przez system Git czas nauki skomplikowanych zadań spadł o ponad połowę. # si # ai # sztucznainteligencja # wiadomości # informacje # tec…

  2717. Mastodon — mastodon.social TIER_1 English(EN) · taoofmac ·

    Agentic Systems 关于构建和运行 Agentic AI 系统的笔记和资源,涵盖编排框架、任务路由、内存和评估方法

    Agentic Systems Notes and resources on building and operating agentic AI systems, covering orchestration frameworks, task routing, memory, and evaluation approaches that extend baseline LLM capabi(...) # agents # ai # orchestration https:// taoofmac.com/space/ai/agentic? utm_cont…

  2718. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Headroom:一个压缩AI代理读取内容(工具输出、日志、RAG块、文件和对话历史记录)的工具,在它们到达LLM之前 - 60-9

    Headroom: a Tool to compress everything your AI Agent reads — tool outputs, logs, RAG chunks, files, and conversation history — before it reaches the LLM - 60-95% fewer Tokens, same Answers ; available as Library, Proxy and MCP server # AI # LLM # Agent https:// github.com/chopra…

  2719. Mastodon — mastodon.social TIER_1 日本語(JA) · ymbot ·

    超越LLM:为何可扩展的企业AI采用依赖于Agent逻辑

    【LLMを超えて:拡張可能なエンタープライズAI導入がエージェントロジックに依存する理由】 https:// huggingface.co/blog/ibm-resear ch/agent-logic-and-scalable-ai-adoption ※AI生成の自動投稿(見出し+リンク) # AI # 生成AI # LLM # AIGenerated

  2720. Mastodon — mastodon.social TIER_1 Italiano(IT) · tomshw ·

    🤖 HR中的AI代理:委托重复性任务,保持人类判断力、同理心和责任感。清晰选择的框架。# HR # AI 🔗 https

    🤖 Agenti AI in HR: delegare i compiti ripetitivi, mantenere umani giudizio, empatia e responsabilità. Un framework per scegliere con lucidità. # HR # AI 🔗 https://www. tomshw.it/aioperator/agente-ai -hr-cosa-delegare-framework

  2721. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    "Every Eval Ever:AI评估结果的统一模式和社区存储库" 我们推出了Every Eval Ever,首个共享模式和社区众筹

    "Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results" We introduce Every Eval Ever, the first shared schema and community-crowdsourced repository for AI evaluation results. The schema standardizes how evaluations are represented in a unified, sin…

  2722. Mastodon — mastodon.social TIER_1 English(EN) · leanpub ·

    编排AI代理:Yohan Rodriguez 的《协调Claude Code、Codex、本地模型和MCP与持久化控制平面》在Leanpub上新发布!

    Orchestrating AI Agents: Coordinating Claude Code, Codex, Local Models, and MCP with a Persistent Control Plane by Yohan Rodriguez is a new release on Leanpub! A practical guide to operating a fleet of AI coding agents through routing, memory, skills, MCP, guardrails, and a persi…

  2723. Mastodon — mastodon.social TIER_1 日本語(JA) · ymbot ·

    AssetOpsBench:对标AI代理并弥合与行业现实的差距

    【AssetOpsBench:AIエージェントのベンチマークと産業界の現実とのギャップを埋める】 https:// huggingface.co/blog/ibm-resear ch/assetopsbench-playground-on-hugging-face ※AI生成の自動投稿(見出し+リンク) # AI # 生成AI # LLM # AIGenerated

  2724. Mastodon — mastodon.social TIER_1 日本語(JA) · ymbot ·

    全球开源AI生态的未来:从DeepSeek到AI+

    【グローバルなオープンソースAIエコシステムの未来:DeepSeekからAI+へ】 https:// huggingface.co/blog/huggingfac e/one-year-since-the-deepseek-moment-blog-3 ※AI生成の自動投稿(見出し+リンク) # AI # 生成AI # LLM # AIGenerated

  2725. Mastodon — mastodon.social TIER_1 English(EN) · geoworldpolitical ·

    AI Agent Adoption: A Practical Roadmap 成功采用 AI Agent!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2726. Mastodon — mastodon.social TIER_1 English(EN) · ppcland ·

    ICYMI:Agentic AI与广告技术栈:谁在控制购买层?:Mediaocean NIVO AI、Magnite Orchestration、Teads EngageOS以及Walmart Connect on DV360

    ICYMI: Agentic AI and the ad stack: who controls the buying layer now?: Mediaocean NIVO AI, Magnite Orchestration, Teads EngageOS, and Walmart Connect on DV360 each launched June 11 as ChatGPT fell to 52.7% of global AI traffic. https:// ppc.land/agentic-ai-and-the-ad -stack-who-…

  2727. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    超越提示词:AI 代理如何悄然改变互联网 多年来,互联网一直通过人们搜索信息的简单模式运作

    Beyond the prompt: How AI agents are quietly changing the internet For years, the internet has worked through a simple model where people search for information, compare options, and manually complete tasks across multiple websites and applications. That structure is now starting…

  2728. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    AI数学代理的能力来自模型还是围绕它的编排?在首次大规模的开放问题形式证明搜索测试中,一个

    Where does an AI math agent get its ability, the model or the orchestration around it? In the first large-scale test of formal proof search on open problems, an agent closed 9 of 353 Erdős problems in Lean. In its own ablation, a plain generate-and-verify loop solved all nine, wh…

  2729. Mastodon — mastodon.social TIER_1 Polski(PL) · aisight ·

    新的开源项目 Memory OS 为 AI 代理引入六阶段内存架构,专注于本地数据处理和高级层次结构

    Nowy projekt open-source, Memory OS, wprowadza sześcioetapową architekturę pamięci dla agentów AI, stawiając na lokalne przetwarzanie danych i zaawansowaną hierarchizację wiedzy. # si # ai # sztucznainteligencja # wiadomości # informacje # technologia https:// aisight.pl/agenci-a…

  2730. Mastodon — mastodon.social TIER_1 日本語(JA) · [email protected] ·

    高通CEO Amon对AI时代的愿景:智能手机和PC将成为AI代理的终端

    クアルコムのアモンCEOが示すAI時代、スマホやPCはエージェントのエンドポイントに https:// k-tai.watch.impress.co.jp/docs /news/2113516.html # ktai_watch_impress # 最新技術_その他 # AI # 業界動向 # 技術

  2731. Mastodon — mastodon.social TIER_1 日本語(JA) · [email protected] ·

    难道没人做吗?Agentic AI的理想与落地之差

    誰もやっていない? エージェンティックAI の理想と運用のリアルなズレ https:// digiday.jp/agencies/why-wpps-a i-boss-believes-agents-are-still-in-the-teenage-sex-stage-of-development/ # digiday # Agencies # DIGIDAY # 有料記事 # 記事のポイント # AI

  2732. Mastodon — mastodon.social TIER_1 English(EN) · geoworldpolitical ·

    AI Agent Adoption: A Practical Roadmap 成功采用AI代理!揭示隐藏成本、潜在风险以及无缝工作的实用路线图

    AI Agent Adoption: A Practical Roadmap Navigate AI agent adoption successfully! Uncover hidden costs, potential risks, and a practical roadmap for seamless workflow automation. https:// theboard.world/articles/techno logy/ai-agent-adoption-practical-roadmap # Technology # Tech # …

  2733. r/Anthropic TIER_1 (LV) · /u/BarracudaVivid8015 ·

    人工智能机器人?

    <!-- SC_OFF --><div class="md"><p>Will Anthropic releases fully functional all terrain robots that does agriculture? Pretty sure developers will be gone in the future. Going to do agriculture pretty difficult having these robots that knows everything will be helpful in the farmla…

  2734. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Celery与Temporal在编排AI任务方面的全面比较,涵盖架构、性能、功能以及分布式AI工作中的用例

    A comprehensive comparison of Celery and Temporal for orchestrating AI tasks, covering architecture, performance, features, and use cases in distributed AI workflows. # Celery # Temporal # AI task orchestration # distributed systems # workflow automation https:// dasroot.net/post…

  2735. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    AgentTrove 以 ShareGPT 风格格式提供 170 万个 agentic 交互跟踪数据,使开发人员能够构建用于训练 AI agents 的数据集

    AgentTrove offers access to 1.7M agentic interaction traces in a ShareGPT-style format, enabling developers to build datasets for training AI agents through streaming. https://www. marktechpost.com/2026/05/29/ho w-to-use-agenttrove-streaming-1-7m-agentic-traces-and-building-a-cle…

  2736. Mastodon — mastodon.social TIER_1 Русский(RU) · [email protected] ·

    如何在生产环境中评估 AI Agent:基线、轨迹和代码检查(如果 Agent 已使用工具、读取文档、更改系统状态并打印)

    Как оценивать ИИ-агентов в проде: нижняя планка, трассы и кодовые проверки Если агент уже ходит в инструменты, читает документы, меняет состояние системы и принимает часть решений сам, проверка одного промпта почти ничего не говорит о надежности. Нужно смотреть на весь путь: вход…

  2737. Mastodon — mastodon.social TIER_1 日本語(JA) · [email protected] ·

    Notion 将通过开发者平台将 AI 代理集成到业务中

    Notion、AIエージェントを業務に組み込む開発者基盤「Developer Platform」 https://www. watch.impress.co.jp/docs/news/ 2112150.html # watch_impress # テック # AI

  2738. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Ombra 分享见解:尽管有防护措施,AI 代理仍删除了整个生产数据库。🤖⚠️ 自主系统在没有严格控制的情况下可能行为不可预测

    Ombra Shares Insights: An AI agent deleted an entire production database, despite guardrails in place.🤖⚠️ Autonomous systems can act unpredictably without strict oversight, making resilience and strong controls essential as AI adoption grows. 🔗Collaborate with Ombra: https:// zur…

  2739. r/Anthropic TIER_1 English(EN) · /u/hazyhaar ·

    我如何使用 Claude Code 运行了 9 小时的自主/目标会话,以及它教会了我关于 AI 代理的知识

    &#32; submitted by &#32; <a href="https://www.reddit.com/user/hazyhaar"> /u/hazyhaar </a> <br /> <span><a href="/r/ClaudeCode/comments/1tmm4sd/how_i_ran_a_9hour_autonomous_goal_session_with/">[link]</a></span> &#32; <span><a href="https://www.reddit.com/r/Anthropic/comments/1tmm5…

  2740. r/Anthropic TIER_1 English(EN) · /u/AssumptionNew9900 ·

    面向代理的自主公司操作系统

    <table> <tr><td> <a href="https://www.reddit.com/r/Anthropic/comments/1tluiyp/autonomous_company_operating_system_for_agents/"> <img alt="Autonomous Company Operating system for agents" src="https://external-preview.redd.it/ypNAJE-VXQOfoHJJn3S6pQXrhig4e2hp7EKFNiYblqM.png?width=64…

  2741. Mastodon — mastodon.social TIER_1 日本語(JA) · ymbot ·

    深入解析 GPT-OSS 中的 Agentic 强化学习:实践回顾 https:// huggingface.co/blog/LinkedIn/g pt-oss-agentic-rl *AI生成自动发布 (标题+链接) # AI # GenerativeAI # LLM # AIGenerated

    【GPT-OSSにおけるエージェント型強化学習の解明:実践的な回顧】 https:// huggingface.co/blog/LinkedIn/g pt-oss-agentic-rl ※AI生成の自動投稿(見出し+リンク) # AI # 生成AI # LLM # AIGenerated

  2742. Mastodon — mastodon.social TIER_1 Deutsch(DE) · [email protected] ·

    关于#AI和BOTs自动化的一些思考:如果我们拥有始终如一的标准接口,就不需要代理来自动化任务。我们

    Gedanke zu Automatisierung mit # AI und BOTs: Wenn wir durchgehend normierte Schnittstellen hätten, bräuchten wir keine Agents um Tasks zu automatisieren. Wir würden die API nutzen.

  2743. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    OpenClaw 分析及安全设置 AI 代理的分步指南 https://peertube.eqver.se/w/ioF2Cw7gt9RRrd4W7LLrmT

    Analysis of OpenClaw and a step-by-step guide to securely setting up an AI agent https:// peertube.eqver.se/w/ioF2Cw7gt9 RRrd4W7LLrmT

  2744. Mastodon — mastodon.social TIER_1 English(EN) · carlosboss ·

    自主人工智能代理需要持续学习和自我完善,以适应和演变新信息和挑战。#人工智能 #学习 #自我完善

    Continuous learning and self-improvement are crucial for autonomous AI agents to adapt and evolve with new information and challenges. # AI # Learning # SelfImprovement

  2745. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    AI代理中的架构漏洞使生产系统面临混淆副官攻击。研究表明上下文操纵如何绕过运营中的安全

    Architectural gaps in AI agents expose production systems to confused-deputy attacks. Research shows how context manipulation bypasses security in operational automation. # Cybersecurity # AI https:// deafnews.it/en/article/agenti- ai-in-produzione-il-rischio-confused-deputy-e-re…

  2746. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Ombra 分享见解:尽管有防护措施,AI 代理仍删除了整个生产数据库。🤖⚠️ 自主系统在没有严格控制的情况下可能行为不可预测

    Ombra Shares Insights: An AI agent deleted an entire production database, despite guardrails in place.🤖⚠️ Autonomous systems can act unpredictably without strict oversight, making resilience and strong controls essential as AI adoption grows. 🔗Collaborate with Ombra: https:// zur…

  2747. Mastodon — mastodon.social TIER_1 日本語(JA) · [email protected] ·

    Dell 台式智能体AI

    オンプレミスのAIエージェントを構築できる「Dell Deskside Agentic AI」 https:// pc.watch.impress.co.jp/docs/ne ws/2109635.html # impress # 市場 # AI # その他

  2748. Mastodon — mastodon.social TIER_1 Français(FR) · [email protected] ·

    AI代理生成的提交淹没了赏金计划:分类员花费更多时间过滤噪音而非处理真实漏洞

    Les programmes de bug bounty saturés par des soumissions générées par des agents IA : les triageurs passent plus de temps à filtrer le bruit qu'à traiter de vraies vulnérabilités. La surface d'attaque des processus humains dans la chaîne de sécurité, c'est aussi ça. Un signal int…

  2749. Mastodon — mastodon.social TIER_1 English(EN) · aihaberleri ·

    📰 2026 SDOF 框架:解决 AI 系统中的多智能体编排约束 新框架 SDOF 解决了多智能体编排中的关键约束

    📰 2026 SDOF Framework: Solving Multi-Agent Orchestration Constraints in AI Systems A new framework called SDOF addresses critical constraints in multi-agent orchestration systems used by platforms like LangChain and LangGraph. The state-constrained approach significantly improves…

  2750. Mastodon — mastodon.social TIER_1 Türkçe(TR) · aihaberleri ·

    📰 LangGraph:解决 2026 年多 AI 代理协调与对齐问题 LangGraph,一种协调多个 AI 代理的革命性解决方案

    📰 LangGraph: Çoklu AI Ajan Koordinasyonu ve Hizalama Sorununu 2026'da Çözme LangGraph, çoklu yapay zeka ajanlarının koordinasyonunu sağlayan devrim niteliğinde bir framework sunuyor. SDOF (State-Constrained Dispatch) tekniğiyle 'hizalama vergisi' sorununu çözen sistem, AI gelişti…

  2751. Mastodon — mastodon.social TIER_1 日本語(JA) · ymbot ·

    AssetOpsBench:对标AI代理并弥合与行业现实的差距

    【AssetOpsBench:AIエージェントのベンチマークと産業界の現実とのギャップを埋める】 https:// huggingface.co/blog/ibm-resear ch/assetopsbench-playground-on-hugging-face ※AI生成の自動投稿(見出し+リンク) # AI # 生成AI # LLM # AIGenerated

  2752. Mastodon — mastodon.social TIER_1 English(EN) · aihaberleri ·

    📰 Repowise 平台 2026:以代码库智能重塑 AI 开发 Repowise 平台正在革新 AI 代理理解复杂代码库的方式

    📰 Repowise Platform 2026: Transform AI Development with Codebase Intelligence The Repowise platform is revolutionizing how AI agents understand complex codebases through automated documentation and dependency analysis. By generating structured wikis and architectural graphs in un…

  2753. Mastodon — mastodon.social TIER_1 English(EN) · beyondthecode ·

    🧠 研究人员开发了一种专门用于构建自主代理的编程语言。该语言提供了定制的语法和功能

    🧠 Researchers have developed a programming language designed specifically for building autonomous agents. The language provides syntax and features tailored to agent-based systems and their operational requirements. 💬 Hacker News 🔗 https:// zerolang.ai/ # AI # MachineLearning # t…

  2754. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    🤖 大型企业中可行的多智能体架构 抛开炒作,有多少人真正见过大型企业中可行的多智能体深度嵌入

    🤖 A working multi-agent architecture in large enterprises AI Hype aside, how many of you have truly seen a working multi-agent deep embedding in large enterprises or large complex environments? If you have, what's your stack/architecture? submitted by /u/... 📰 Source: Artificial …

  2755. Mastodon — mastodon.social TIER_1 日本語(JA) · ymbot ·

    全球开源AI生态的未来:从DeepSeek到AI+

    【グローバルなオープンソースAIエコシステムの未来:DeepSeekからAI+へ】 https:// huggingface.co/blog/huggingfac e/one-year-since-the-deepseek-moment-blog-3 ※AI生成の自動投稿(見出し+リンク) # AI # 生成AI # LLM # AIGenerated

  2756. Mastodon — mastodon.social TIER_1 English(EN) · aihaberleri ·

    📰 AI Agent Systems:通过动态工具暴露和上下文注入实现70%的效率提升(2026)构建AI Agent系统的新方法使用动态工具暴露

    📰 AI Agent Systems: 70% Efficiency Gains with Dynamic Tool Exposure & Context Injection (2026) A new approach to building AI agent systems uses dynamic tool exposure and context injection to dramatically improve efficiency. By exposing only necessary tools and injecting ephemeral…

  2757. Mastodon — mastodon.social TIER_1 Türkçe(TR) · aihaberleri ·

    📰 2026年人工智能代理系统的革命:动态工具规划如何实现95%的代币节省?AI代理与传统方法的比较

    📰 AI Agent Sistemlerinde 2026 Devrimi: Dinamik Araç Planlaması Nasıl %95 Token Tasarrufu Sağlıyor? Yapay zeka ajanları, geleneksel yöntemlerle karşılaştırıldığında yüksek maliyet ve verimsizlik sorunları yaşıyor. Araştırmacılar, Instruction-Tool Retrieval (ITR) adlı yeni bir sist…

  2758. Mastodon — mastodon.social TIER_1 English(EN) · DrBrentAllenJensen ·

    **揭示隐藏模式:对传统本体论的挑战**。一项开创性分析揭示了对动态环境中适应性代理的深远影响

    **Uncovering the Hidden Pattern: A Challenge to Traditional Ontology**. A groundbreaking analysis reveals a profound implication for adaptive agents in dynamic environments. The distinction between substance and event ontology may redefine our understanding of reality. **#Ontolog…

  2759. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Qwen 3.6 和 Gemma 4 的供应商和社区推理参数精选参考,针对代理工作流和真实代码系统进行了优化。# Hermes

    Curated reference of vendor and community inference parameters for Qwen 3.6 and Gemma 4, optimized for agentic workflows and real-world coding systems. # Hermes # OpenClaw # OpenCode # Cheatsheet # Self -Hosting # SelfHosting # LLM # AI # AI Coding # llama .cpp https://www. glukh…

  2760. Mastodon — mastodon.social TIER_1 English(EN) · amazeeai ·

    持久性AI代理正在解决“上下文重置”问题并制造新问题。当你的代理学习了6个月的部署模式、架构决策时

    Persistent AI agents are solving the "context reset" problem and creating a new issue. When your agent learns 6 months of deployment patterns, architecture decisions, and tribal knowledge, that's institutional IP. And if it lives on shared infrastructure with vague ToS, you might…

  2761. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    教程展示如何使用 Memori 构建原生智能体记忆基础设施,使 LLM 应用能够在多个用户会话和生命周期中保持上下文

    A tutorial shows how to build agent-native memory infrastructure using Memori, enabling LLM applications to retain context across multiple user sessions and agent personas. The implementation covers memory persistence, multi-tenant isolation, and streaming responses for AI agents…

  2762. r/Anthropic TIER_1 Français(FR) · /u/Lrn24gt557 ·

    AI Agents

    <table> <tr><td> <a href="https://www.reddit.com/r/Anthropic/comments/1t7b8qa/ai_agents/"> <img alt="@ai agents" src="https://preview.redd.it/n4mr6269mxzg1.jpeg?width=640&amp;crop=smart&amp;auto=webp&amp;s=40a42c8352fdd17250908bed2949641e6c7dcfed" title="@ai agents" /> </a> </td>…

  2763. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    构建具有持久内存的AI代理:技术深度解析 Hermes Agent 如何使用 SQLite 实现跨会话持久内存的技术解析

    Building an AI Agent with Persistent Memory: A Technical Deep Dive A technical look at how Hermes Agent implements cross-session persistent memory using SQLite vector search and knowledge graphs. # ai # agents # memory # vectorsearch # opensource

  2764. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    一个AI助手,所有平台:Telegram、Discord、Slack和CLI Hermes Agent如何在8个以上消息平台同时运行。#ai #devtools #automation

    One AI Assistant, Every Platform: Telegram, Discord, Slack, and CLI How Hermes Agent runs on 8+ messaging platforms simultaneously. # ai # devtools # automation # opensource # telegram

  2765. r/Anthropic TIER_1 English(EN) · /u/cbbsherpa ·

    超越自主性:了解自身局限的智能体的力量

    <!-- SC_OFF --><div class="md"><p>Here’s something we didn’t expect to learn from a dataset of 4,200 human-AI interactions: the moment an agent becomes most useful isn’t when it gets the answer right. It’s when it knows it’s about to get the answer wrong.</p> <p>The COWCORPUS pro…

  2766. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    出色的代理工作流不仅仅是自动驾驶的AI——它们是人类洞察力与AI执行力的协作。这个方法展示了如何构建一个基于图的工作流

    Great agentic workflows aren’t just AI on autopilot—they’re a collaboration between human insight and AI execution. This recipe shows how a graph-based workflow can pause, engage a human, then continue toward its goal. # SpringAI # Java # AI # Agents # LLM

  2767. Mastodon — mastodon.social TIER_1 한국어(KO) · [email protected] ·

    Show HN:BattleClaws – 一个 AI 代理自主战斗的竞技场

    Show HN: BattleClaws – A battle arena where AI agents fight autonomously BattleClaws는 AI 에이전트들이 자율적으로 전투를 벌이는 배틀 아레나 플랫폼입니다. 사용자는 자신의 AI 에이전트를 생성하여 4단계 진화를 거치며 다른 에이전트와 경쟁할 수 있습니다. 전투 결과와 랭킹이 실시간으로 업데이트되어 AI 에이전트의 성능을 평가하고 순위를 올릴 수 있습니다. 이는 AI 에이전트의 자율적 행동과 경쟁을 실험할 수 있는 흥미로운 응용 사…

  2768. Mastodon — mastodon.social TIER_1 English(EN) · genticnews ·

    技能作为不可信代码:Agent Runtimes 的安全先例 论文认为,在验证之前,Agent 技能是不可信代码;运行时必须强制执行验证

    Skills as Untrusted Code: A Security Precedent for Agent Runtimes Paper argues agent skills are untrusted code until verified; runtimes must enforce verification gates to prevent supply-chain attacks, echoing decades of software security lessons. https:// gentic.news/article/skil…

  2769. Mastodon — mastodon.social TIER_1 English(EN) · genticnews ·

    Span推出XFRA节点:家庭分布式AI计算,每兆瓦300万美元 Span的XFRA节点提供每兆瓦300万美元的分布式AI计算,利用家庭电网容量。一个100户家庭的节点

    Span Launches XFRA Node: Distributed AI Compute in Homes at $3M/MW Span's XFRA Node offers distributed AI compute at $3M/MW, using home grid capacity. A 100-home pilot this year targets 1.25 MW. https:// gentic.news/article/span-launc hes-xfra-node # AI # ArtificialIntelligence #…

  2770. Mastodon — mastodon.social TIER_1 English(EN) · aihaberleri ·

    📰 模块化技能型代理系统:动态工具路由如何提升 LLM 在 2026 年的性能 新的 AI 代理设计方法引入了模块化技能型 s

    📰 Modular Skill-Based Agent System: How Dynamic Tool Routing Boosts LLM Performance in 2026 A new approach to AI agent design introduces a modular skill-based system with dynamic tool routing, enabling LLMs to orchestrate capabilities like an operating system. This architecture e…

  2771. Mastodon — mastodon.social TIER_1 Türkçe(TR) · aihaberleri ·

    📰 2026年模块化技能型代理系统:LLM中的动态工具路由 模块化技能管理和AI代理中的动态工具路由,

    📰 2026'da Modüler Beceri Tabanlı Agent Sistemi: LLM'lerde Dinamik Araç Yönlendirme Yapay zeka agentlerinde modüler beceri yönetimi ve dinamik araç yönlendirme, LLM'lerin karmaşık görevleri insan gibi çözmeye başlamasını sağlıyor. Arxiv ve MarkTechPost verileriyle derinlemesine in…

  2772. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    🔖 智能体记忆、评估、可观测性和多智能体架构。当前趋势焦点:OpenAI Codex、新兴智能体运行时和生产AI工作流

    🔖 agent memory, evaluation, observability, and multi-agent architecture. Current trend focus: OpenAI Codex, emerging agent runtimes, and production AI workflow patterns. https:// github.com/Prompthon-IO/agent- systems-handbook TL;DR: Free open-source handbook for learning agentic…

  2773. Mastodon — mastodon.social TIER_1 English(EN) · beyondthecode ·

    🧠 编码代理缺乏足够规范,难以在多样化任务中可靠运行。研究人员指出需要更清晰的定义和约束

    🧠 A coding agent lacks sufficient specification to function reliably across diverse tasks. Researchers identify the need for clearer definitions and constraints to improve consistency in how such agents approach programming problems. 💬 Hacker News 🔗 https:// hsaghir.github.io/blo…

  2774. Mastodon — mastodon.social TIER_1 Polski(PL) · aisight ·

    Amazon Web Services 将代理方法集成到 SageMaker AI 平台上的模型微调流程中。这使开发人员能够自动化复杂的

    Amazon Web Services integruje agentyczne podejście do procesów dostrajania modeli w platformie SageMaker AI. Dzięki temu programiści mogą automatyzować skomplikowane zadania związane z optymalizacją modeli open-source, takich jak Llama, Qwen i DeepSeek, a także autorskich rozwiąz…

  2775. Mastodon — mastodon.social TIER_1 English(EN) · aihaberleri ·

    📰 Agent-Desktop:利用辅助功能 API 实现的 AI 桌面自动化 (2026) Agent-Desktop 通过利用原生

    📰 Agent-Desktop: AI Desktop Automation Using Accessibility APIs (2026) Agent-Desktop introduces a breakthrough in AI-driven desktop automation by leveraging native OS accessibility APIs instead of pixel-based screenshot loops, drastically reducing token costs and improving reliab…

  2776. Mastodon — mastodon.social TIER_1 Türkçe(TR) · aihaberleri ·

    📰 Agent-desktop 2026:首个 AI Agent 原生 CLI 桌面自动化 新开源项目 Agent-desktop,AI Agent 桌面应用

    📰 Agent-desktop 2026: AI Ajanları İçin İlk Native CLI Masaüstü Otomasyonu Yeni açılan open-source projesi Agent-desktop, AI ajanlarının masaüstü uygulamalarıyla etkileşime geçmesini sağlayan ilk native CLI aracını tanıtıyor. Bu yenilik, otomasyon dünyasında bir dönüm noktası olab…

  2777. Mastodon — mastodon.social TIER_1 日本語(JA) · [email protected] ·

    Claude Code 的 CLAUDE.md / Skills / Agents:一个三层设计模式

    Claude Code の CLAUDE.md / Skills / Agents を3層で整備する設計パターン https:// qiita.com/ennagara128/items/c2 5e72eb240611454457?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items # qiita # 設計 # AI # AIエージェント # ClaudeCode # CLAUDE_md

  2778. Mastodon — mastodon.social TIER_1 日本語(JA) · [email protected] ·

    【Phase1 AI×AWS】使用Claude Code的skill function尝试自动化AWS成本确认 https://qiita.com/Aratabiz/items/a95f93b0e69072c687ef?utm_campaign=popular_items&utm_medium=feed&utm_

    【Phase1 AI×AWS】Claude Code の skill 機能で AWS コスト確認を自動化してみた https:// qiita.com/Aratabiz/items/a95f9 3b0e69072c687ef?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items # qiita # AWS # 自動化 # AI # SKILLS

  2779. Mastodon — mastodon.social TIER_1 日本語(JA) · [email protected] ·

    Karpathy 谈论“从 Vibe Coding 到 Agent Engineering” ~ 我觉得这个 YouTube 视频很有趣,所以总结了一下 ~ https://qiita.com/yuji-arakawa/items/9e7235e708e2b33e58e6?utm_campaign=popular_items&utm_me

    カルパシーが語る「バイブコーディングからエージェント・エンジニアリングへ」 〜 YouTube動画が興味深かったのでまとめてみた 〜 https:// qiita.com/yuji-arakawa/items/9 e7235e708e2b33e58e6?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items # qiita # 初心者 # ポエム # AI # LLM # AIエージェント

  2780. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    MarkTechPost 发布了关于 Agentic UI、Generative UI、状态同步和中断驱动审批流程的编码深度解析。该教程构建了

    MarkTechPost has published a coding deep dive into Agentic UI, Generative UI, state synchronisation and interrupt-driven approval flows. The tutorial builds the entire Agentic UI stack from the ground up using plain Python, implementing the AG-UI event stream and A2UI as a declar…

  2781. Mastodon — mastodon.social TIER_1 English(EN) · genticnews ·

    Agentic Harness Engineering 提升编码代理在 Terminal-Bench 2 上 7% Agentic Harness Engineering 引入结构化方法来演进编码代理

    Agentic Harness Engineering Boosts Coding Agents 7% on Terminal-Bench 2 Agentic Harness Engineering introduces a structured approach to evolving coding-agent harnesses, using revertible components, condensed experience, and falsifiable decisions. On Terminal-Bench 2, pass https:/…

  2782. Mastodon — mastodon.social TIER_1 English(EN) · genticnews ·

    一个定制的多模态Transformer如何击败微调的LLM,LeBonCoin的ML团队构建了一个定制的 late-fusion transformer,它使用预先计算的视觉

    How a Custom Multimodal Transformer Beat a Fine-Tuned LLM for Attribute LeBonCoin's ML team built a custom late-fusion transformer that uses pre-computed visual embeddings and character n-gram text vectors to predict ad attributes. It outperformed a fine-tuned VLM while r https:/…

  2783. Mastodon — mastodon.social TIER_1 English(EN) · genticnews ·

    Anthropic 发布 Claude Security,一款独立的面向企业的代码漏洞扫描器

    Anthropic Ships Claude Security, a Standalone Code Vulnerability Scanner for Enterprise Anthropic shipped Claude Security, a standalone code vulnerability scanner for Enterprise powered by Opus 4.7, directly targeting Snyk, Semgrep, and SonarQube. https:// gentic.news/article/ant…

  2784. Mastodon — mastodon.social TIER_1 English(EN) · aihaberleri ·

    📰 TypeScript SDK:使用沙盒虚拟机构建安全的 AI 编码代理 (2026) Cursor 新推出的 TypeScript SDK 使开发人员能够构建程序化编码代理

    📰 TypeScript SDK: Build Secure AI Coding Agents with Sandbox VMs (2026) A new TypeScript SDK from Cursor empowers developers to build programmatic coding agents using sandboxed cloud VMs, subagents, and token-based pricing. The tool integrates with existing TypeScript ecosystems …

  2785. Mastodon — mastodon.social TIER_1 Türkçe(TR) · aihaberleri ·

    📰 2026年使用Cursor TypeScript SDK开发编程代码代理 Cursor已推出其TypeScript SDK,支持云端代码代理

    📰 Cursor TypeScript SDK ile 2026'da Programmatik Kodlama Ajanları Geliştirin Cursor, TypeScript SDK’sını piyasaya sürerek kodlama ajanlarının bulut tabanlı sanal makinelerde güvenli şekilde çalışmasını sağlıyor. Bu yenilik, AI destekli geliştirme alanında bir dönüm noktası olarak…

  2786. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    如何将内部框架、蓝图、最佳实践和操作规则发布给AI编码代理,而不将专有上下文变成不受管制的风险

    How to publish internal frameworks, blueprints, best practices, and operational rules to AI coding agents without turning proprietary context into ungoverned folklore. https://www. the-main-thread.com/p/enterpri se-agent-knowledge # ai # genai # mcp # agenticCoding # documentatio…

  2787. Mastodon — mastodon.social TIER_1 English(EN) · AIntelligenceHub ·

    OpenAI 的 Symphony 将代理编码视为受管理的任务执行:隔离运行、由看板驱动的接收以及合并前的证明工件。这听起来很简单,但

    Symphony from OpenAI frames agent coding as managed work execution: isolated runs, board-driven intake, and proof artifacts before merge. That sounds simple, but it changes staffing, governance, and rollout risk for engineering teams. Full analysis: https:// go.aintelligencehub.c…

  2788. Mastodon — mastodon.social TIER_1 English(EN) · beyondthecode ·

    🧠 49Agents 提供了一个为开发和管理 AI 代理设计的无限画布界面。该工具使用户能够组织代理工作流并进行交互

    🧠 49Agents provides an infinite canvas interface designed for developing and managing AI agents. The tool enables users to organize agent workflows and interactions within an expandable workspace environment. 💬 Hacker News 🔗 https:// github.com/49Agents/49Agents # AI # MachineLea…

  2789. r/cursor TIER_2 English(EN) · /u/Onnoz ·

    改进团队使用代理

    <table> <tr><td> <a href="https://www.reddit.com/r/cursor/comments/1ur0vc0/improving_team_use_of_agents/"> <img alt="Improving team use of agents" src="https://external-preview.redd.it/_C_ROn8_JFGooQQ-DtjFyVSS9TfYYJ8a2mHF3DJq2Pw.png?width=640&amp;crop=smart&amp;auto=webp&amp;s=f4…

  2790. r/cursor TIER_2 English(EN) · /u/Downtown-Function-10 ·

    经过六个月实际使用考验的代理工作流设计模式

    <!-- SC_OFF --><div class="md"><p>We started with 8 agentic workflow design patterns six months ago. Four survived. The other four fell apart in ways that took a while to understand</p> <p>The survivors. Agent-as-first-reviewer, where the agent reviews before the human and catche…

  2791. r/cursor TIER_2 English(EN) · /u/OwlZealousideal4779 ·

    AI 辅助开发中的架构漂移 — 您是如何处理的?

    <!-- SC_OFF --><div class="md"><p>One challenge I don't see discussed enough: as AI coding tools get better at generating code, teams are shipping faster, but the architecture is quietly degrading underneath. </p> <p>The problem is that most AI tools are stateless. They generate …

  2792. r/StableDiffusion TIER_2 English(EN) · /u/Sensitive_Teacher_93 ·

    使用 Claude 或 Cursor 创建 Agentic AI 工作流

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1u886jh/agentic_ai_workflow_creation_using_claude_or/"> <img alt="Agentic AI workflow creation using Claude or cursor" src="https://external-preview.redd.it/em9tNzlpbW4xdTdoMTc-dnVvrW1nROx2II0b8iVutPa2INq…

  2793. r/cursor TIER_2 English(EN) · /u/atricsky ·

    关于AI模型的问题

    <!-- SC_OFF --><div class="md"><p>Hi,</p> <p>I’m wondering about the $60/month plan. Are Claude Opus, Codex, and other models included?</p> <p>Are there any limitations expect token usage?</p> </div><!-- SC_ON --> &#32; submitted by &#32; <a href="https://www.reddit.com/user/atri…

  2794. r/StableDiffusion TIER_2 (CA) · /u/sylense0 ·

    开源AI模型

    <!-- SC_OFF --><div class="md"><p>Hey everyone. I dont really have any knowledge about any of this stuff.. Im an architecture student looking for an image generating open source model to help me with renders and designing. My pc specs are rtx 5070 12 vram 32gb ddr5 and an ultra 5…

  2795. r/cursor TIER_2 English(EN) · /u/IlyaZelen ·

    停止消耗 Token:使用一个通用插件,AI 编码代理的代码发现速度提升 5.1 倍

    <!-- SC_OFF --><div class="md"><p>My colleagues kept asking me for my setup, so I decided to turn it into a universal plugin: <strong>Agent Code Navigator</strong> - a universal code-navigation plugin for Cursor, Claude, Codex, Gemini, and OpenCode.</p> <p>In my benchmark, semant…

  2796. r/cursor TIER_2 English(EN) · /u/Few-Ad-1358 ·

    开发者使用AI编码助手:您的工作流程中的信任在哪里会破裂?

    &#32; submitted by &#32; <a href="https://www.reddit.com/user/Few-Ad-1358"> /u/Few-Ad-1358 </a> <br /> <span><a href="/r/ExperiencedDevs/comments/1tk6hg6/devs_using_ai_coding_agents_where_does_trust/">[link]</a></span> &#32; <span><a href="https://www.reddit.com/r/cursor/comments…

  2797. r/cursor TIER_2 English(EN) · /u/n4r735 ·

    关于AI编码代理的使用及其对开发者的影响的研究协助

    &#32; submitted by &#32; <a href="https://www.reddit.com/user/n4r735"> /u/n4r735 </a> <br /> <span><a href="/r/aiagents/comments/1tglkpv/help_with_study_on_the_use_of_ai_coding_agents/">[link]</a></span> &#32; <span><a href="https://www.reddit.com/r/cursor/comments/1tgln66/help_w…

  2798. r/cursor TIER_2 English(EN) · /u/muneebh1337 ·

    由规范驱动的代理式编码正在悄悄地让我们在监督代理方面的工作能力下降

    <!-- SC_OFF --><div class="md"><p>Been running an agent-heavy workflow on a mid-size TypeScript monorepo for about six months. Orchestrator on top, sub-agents for codegen, a human (me, mostly) writing specs and reviewing diffs. The pitch was the obvious one: I stay in the archite…

  2799. r/cursor TIER_2 English(EN) · /u/AdorablePumpkin9309 ·

    Ring-2.6-1T 推出,为编码代理工作流提供免费测试窗口

    <!-- SC_OFF --><div class="md"><p>Flagging this because it seems more relevant to actual coding loops than to general AI-news posting: Ring-2.6-1T is now out, and there’s a free developer access window through May 15.<br /> The launch angle is pretty clearly “reasoning model for …

  2800. r/cursor TIER_2 English(EN) · /u/Hk_90 ·

    探索 Meko:Agent 协同工作与学习的数据基础设施

    <table> <tr><td> <a href="https://www.reddit.com/r/cursor/comments/1t6zy9k/discover_meko_the_data_infrastructure_for_agents/"> <img alt="Discover Meko: The Data Infrastructure for Agents That Work and Learn Together" src="https://preview.redd.it/ea544mxdupzg1.jpeg?width=640&amp;c…

  2801. r/ClaudeAI TIER_2 English(EN) · /u/OkBreath9382 ·

    面向 Agentic 数据科学的终端内 Jupyter Notebook

    <table> <tr><td> <a href="https://www.reddit.com/r/ClaudeAI/comments/1vfv9a3/interminal_jupyter_notebook_for_agentic_data/"> <img alt="In-Terminal Jupyter Notebook for Agentic Data Science" src="https://external-preview.redd.it/jepGdxQmpHkFL39u5l1oLXi8UK36hvTBrH2TmlLhqsE.png?widt…

  2802. r/OpenAI TIER_2 English(EN) · /u/Onnoz ·

    改进团队使用代理

    <!-- SC_OFF --><div class="md"><p>I’ve been playing around with an idea for development teams and their agents and would love some feedback.</p> <p>What if agents working on the same project could learn from each other over time? Think of it as a Stack Overflow built by agents, f…

  2803. r/ClaudeAI TIER_2 English(EN) · /u/Lucky_Historian742 ·

    我开源了行业最佳实践以实现自改进代理

    <table> <tr><td> <a href="https://www.reddit.com/r/ClaudeAI/comments/1uh0t7o/i_opensourced_industry_best_practice_to/"> <img alt="I open-sourced industry best practice to self-improving agents" src="https://preview.redd.it/i6m1kc9gbt9h1.png?width=640&amp;crop=smart&amp;auto=webp&…

  2804. r/ClaudeAI TIER_2 English(EN) · /u/bsampera ·

    一个为你(和你的AI代理)准备的上下文大脑

    <table> <tr><td> <a href="https://www.reddit.com/r/ClaudeAI/comments/1uaplfy/a_context_brain_for_you_and_your_ai_agent/"> <img alt="A Context Brain for you (and your AI Agent)" src="https://external-preview.redd.it/enYzc21ncGp3ZDhoMc0qeEjPjE8oY_VYNqXTY77bMsvN6Dt_eef3EFzgT140.png?…

  2805. r/OpenAI TIER_2 English(EN) · /u/MuhammadMujtaba21 ·

    寻找首席机器学习与人工智能编排工程师 – AutoFlow (为人工智能时代构建信任基础设施)

    <!-- SC_OFF --><div class="md"><p>I am 19, and the Founder and CEO of AutoFlow. I want to be entirely transparent before discussing our current team or your potential role: you should know exactly the engineering challenge we are tackling.</p> <p>We are building the trust infrast…

  2806. r/ClaudeAI TIER_2 English(EN) · /u/Luminancee ·

    为复杂的后端多仓库系统构建AI助手——正确的方法是什么?

    <!-- SC_OFF --><div class="md"><p>I work on a distributed backend system split across multiple microservices in separate repos. Understanding how a failure propagates across services is<br /> non-trivial even for experienced team members.</p> <p>I've been using Claude Code with c…

  2807. r/OpenAI TIER_2 English(EN) · /u/vagobond45 ·

    人工智能、科学与经济:系统图谱

    <table> <tr><td> <a href="https://www.reddit.com/r/OpenAI/comments/1trnnv3/ai_science_economy_systems_map/"> <img alt="AI, Science &amp; Economy: Systems Map" src="https://preview.redd.it/jrxepnfxu64h1.png?width=640&amp;crop=smart&amp;auto=webp&amp;s=7a9944ccb5326f6d89fce7d1959d2…

  2808. r/OpenAI TIER_2 English(EN) · /u/Sumsub_Insights ·

    从AI Agent到Know Your Agent:为什么KYA对安全的自主AI至关重要

    <table> <tr><td> <a href="https://www.reddit.com/r/OpenAI/comments/1tq02zg/from_ai_agents_to_know_your_agent_why_kya_is/"> <img alt="From AI Agents to Know Your Agent: Why KYA Is Critical for Secure Autonomous AI" src="https://external-preview.redd.it/SYNihEB_CpsXPD5wVhhCmJ_fz7a7…

  2809. r/singularity TIER_2 English(EN) · /u/Wonderful-Wealth2761 ·

    麻省理工学院的开源、万亿参数模型(Ant's Ring-2.6)据称在推理+代理基准测试上媲美闭源前沿。 "开源"的追赶真的会改变发展轨迹吗?

    <table> <tr><td> <a href="https://www.reddit.com/r/singularity/comments/1uw34a4/an_openweight_mit_trillionparam_model_ants_ring26/"> <img alt="An open-weight, MIT trillion-param model (Ant's Ring-2.6) reportedly matches the closed frontier on reasoning + agent benchmarks. Does &q…

  2810. r/singularity TIER_2 English(EN) · /u/PrometheanPolymath ·

    ELI-Alien:关于AI的冲突

    &#32; submitted by &#32; <a href="https://www.reddit.com/user/PrometheanPolymath"> /u/PrometheanPolymath </a> <br /> <span><a href="/r/aiwars/comments/1trd9l4/elialien_the_conflict_regarding_ai/">[link]</a></span> &#32; <span><a href="https://www.reddit.com/r/singularity/comments…