English(EN)Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists
最新研究探讨LLM智能体在技能选择、自动驾驶和合规性方面的进展
作者PulseAugur 编辑部·[492 个来源]·
arXiv上发布的多篇研究论文探讨了大型语言模型(LLM)智能体的进展,重点在于提高其能力和可靠性。其中一篇论文介绍了用于LLM智能体最优技能选择的最佳前缀选择(BPS),该方法在性能和代币成本方面提供了可证明的保证。另一项研究提出了一个混合框架用于自动驾驶,该框架整合了LLM的常识推理与强化学习和PID控制,以增强决策能力。此外,还有研究通过纵向生命轨迹来缓解LLM智能体中的身份本质主义,并开发了LLM智能体的策略合规性和故障归因方法。
AI
arXiv:2609.21267v1 Announce Type: new Abstract: Production LLM agents are evaluated repeatedly as they evolve, but full agent benchmarks are costly to rerun. We study efficient recurring evaluation for a production analytics agent serving tens of thousands of monthly active users…
arXiv:2609.19425v1 Announce Type: new Abstract: Tool-augmented large language model (LLM) agents fail in a way no tool-selection or tool-security method addresses: they call tools that do not exist and pass arguments no schema declares. Existing defenses either pick the right too…
arXiv cs.AI
TIER_1English(EN)·Zehua Zhang, Ati Priya Bajaj, Divij Handa, Siyu Liu, Arvind S Raj, Hongkai Chen, Hulin Wang, Yibo Liu, Zion Leonahenahe Basque, Souradip Nath, Vishal Juneja, Nikhil Chapre, Tiffany Bao, Yan Shoshitaishvili, Adam Doup\'e, Chitta Baral, Ruoyu Wang·
arXiv:2509.25248v2 Announce Type: replace-cross Abstract: Automatically compiling open-source software (OSS) projects is a vital, labor-intensive, and complex task, which makes it a good challenge for LLM Agents. Existing methods rely on manually curated rules and workflows, whic…
arXiv:2510.25595v2 Announce Type: replace-cross Abstract: While Large Language Model (LLM) agents are often approached from the angle of action planning/generation to accomplish a goal (e.g., given by language descriptions), their abilities to collaborate with each other to achie…
arXiv cs.CL
TIER_1English(EN)·Mikhail Menschikov, Matvey Iskornev, Alexander Kharitonov, Alina Bogdanova, Mikhail Belkin, Ekaterina Lisitsyna, Artyom Sosedka, Victoria Dochkina, Ruslan Kostoev, Ilia Perepechkin, Evgeny Burnaev·
arXiv:2605.13481v2 Announce Type: replace Abstract: We introduce PersonalAI 2.0 (PAI-2), a novel framework designed to enhance LLM-based systems through integration of external knowledge graphs (KGs). The proposed approach addresses key limitations of existing Graph Retrieval-Aug…
arXiv cs.AI
TIER_1English(EN)·Abdelghny Orogat, Ana Rostam, Essam Mansour·
arXiv:2602.03128v2 Announce Type: replace Abstract: Multi-agent LLM frameworks are data-intensive systems that govern how agents orchestrate tasks, manage state, and coordinate decisions. These architectural choices control execution overhead, memory behavior, planning effectiven…
arXiv:2609.20474v1 Announce Type: new Abstract: Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airline pilot in $\tau^2$-bench. The p…
arXiv cs.AI
TIER_1English(EN)·Ruiyang Wang, Hao-Lun Hsu, Swarajh Mehta, Jiwoo Kim, Zhihao Dou, Miroslav Pajic·
arXiv:2609.19315v1 Announce Type: cross Abstract: Large language models (LLMs) provide a flexible interface for long-horizon robot planning, but generated plans often fail to respect embodiment constraints, recover from planning errors, or reason effectively under partial observa…
arXiv cs.AI
TIER_1English(EN)·Haya Halimeh, Sascha Kaltenpoth, Kevin B\"osch, Oliver M\"uller·
arXiv:2609.19843v1 Announce Type: new Abstract: LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind. These graphical user interfaces were designed to support, but also deliberately steer, the behaviour and d…
arXiv:2609.20812v1 Announce Type: cross Abstract: Frontier coding agents are increasingly trusted to work autonomously for long periods, yet an agent's final response is often the only account of that work a user sees. We quantify the propensity of frontier agents to \emph{overcl…
arXiv:2609.20625v1 Announce Type: cross Abstract: Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing state, and on a multi-step traject…
Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing state, and on a multi-step trajectory that a re-run rarely repeats. Record-and-repla…
LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind. These graphical user interfaces were designed to support, but also deliberately steer, the behaviour and decisions of users. While behavioural biases in t…
arXiv:2608.28363v2 Announce Type: replace Abstract: LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a successful mutation may leave persistent effects that cannot b…
arXiv cs.AI
TIER_1English(EN)·Sehee Kim, Yumin Choi, Minki Kang, Sung Ju Hwang·
arXiv:2609.17632v1 Announce Type: new Abstract: Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment. This limits their ab…
arXiv:2609.18123v1 Announce Type: new Abstract: Large language model agents tune GPU kernels and serving engines through a closed loop of propose, measure, and keep, but the measurements behind this loop are not trustworthy. We characterize four failure modes from a four-day pilo…
arXiv:2609.18128v1 Announce Type: new Abstract: Large language model (LLM) agents augmented by tools can automate complex, multi-step tasks, such as web navigation, code generation, and workflow orchestration, by acting on external systems through tool calls. However, hallucinati…
arXiv:2609.18460v1 Announce Type: new Abstract: How does a multi-agent system evolve from a local deviation into collective loss of control? We propose an epidemic explanation organized around accidental mutation, contagion, and recovery. A spontaneous deviation creates a seed; c…
arXiv:2609.18304v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly tackle long-horizon tasks through multi-step environment interaction, yet a single erroneous action can alter subsequent states and observations, causing errors to compound over time. E…
arXiv:2609.18779v1 Announce Type: new Abstract: Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, limiting their ability to respond to evolving agent capabilities. This disconnect prevents routing strategies from …
Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, limiting their ability to respond to evolving agent capabilities. This disconnect prevents routing strategies from adapting to evolving agent capabilities during p…
Large language model (LLM) agents augmented by tools can automate complex, multi-step tasks, such as web navigation, code generation, and workflow orchestration, by acting on external systems through tool calls. However, hallucinations, distributional instability, and adversarial…
arXiv:2609.16091v1 Announce Type: new Abstract: Tabular foundation models deliver strong zero-training predictive performance via in-context learning, but their high inference latency makes them impractical as hot-path decision backends in interactive agentic loops. We distill a …
arXiv cs.AI
TIER_1English(EN)·Zhen Li, Jun Cai, Haoran Gao, An Li, Tan Li·
arXiv:2609.17193v1 Announce Type: new Abstract: Large language model (LLM)-powered agentic AI services increasingly demand low-latency inference, motivating the deployment of LLMs across distributed edge servers. However, heterogeneous communication and computing capabilities, to…
arXiv:2609.17419v1 Announce Type: new Abstract: Long-horizon LLM agents must maintain task state across extended sequences of observations, actions, tool calls, and intermediate beliefs. We study these trajectories through three dynamical views: self-organized criticality, weak c…
arXiv:2609.16098v1 Announce Type: cross Abstract: Large Language Model (LLM) agents have demonstrated impressive capabilities across a variety of domains, particularly when integrated with external tools for multi-step task completion. However, they are increasingly vulnerable to…
arXiv:2609.16436v1 Announce Type: cross Abstract: Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit. However, LLMs are ultimately black boxes based on deep…
Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, limiting their ability to respond to evolving agent capabilities. This disconnect prevents routing strategies from adapting to evolving agent capabilities during p…
LLM-based multi-agent systems can fail even when communication succeeds because agents do not correctly track their peers' roles, knowledge, or intentions. We investigate whether such inter-agent misalignment cases, labelled FC2 in MAST-Data, can be converted into functional part…
arXiv:2609.14767v1 Announce Type: cross Abstract: Hierarchical orchestration, in which a Manager agent reviews worker output and can send it back for revision, is the default coordination pattern in production multi-agent LLM frameworks. Classical organizational theory predicts t…
arXiv:2609.14987v1 Announce Type: cross Abstract: Large language model (LLM) agents interact with external environments through tool invocation, but tool outputs can also expose them to indirect prompt injection (IPI) attacks. Existing defenses mainly rely on prompt hardening, co…
arXiv cs.AI
TIER_1English(EN)·Mirza Samad Ahmed Baig, Syeda Anshrah Gillani, Asher Ali, Muhammad Hamzah Siddiqui·
arXiv:2609.14780v1 Announce Type: cross Abstract: Multi-tenant tools commonly accept a tenant identifier and validate it against the caller's entitlement. For a large language model (LLM) agent, that pattern delegates resource selection to a process whose context may contain atta…
arXiv:2609.14399v1 Announce Type: new Abstract: Natural language prompts and skills serve as the strategic backbone of LLM-based agents. Recent advances in prompt and skill optimization have achieved notable gains, yet all existing methods optimize a \emph{single} text template--…
arXiv:2609.13335v1 Announce Type: cross Abstract: Large Language Models (LLMs) have enabled more natural human-robot interaction, but open-source models often exhibit unstable long-horizon reasoning and inefficient action execution when deployed in agentic robotic frameworks. Thi…
arXiv cs.AI
TIER_1English(EN)·Ala N. Tak, Teruhisa Misu, Kumar Akash, Zhaobo K. Zheng, Kevin H. Joo, Jonathan Gratch·
arXiv:2609.13261v1 Announce Type: cross Abstract: LLM agents are increasingly used for collaborative problem solving and human-group simulation. This makes outcome-only evaluation insufficient: if LLM groups are used as models of human groups, we need to know whether they succeed…
arXiv:2606.22741v2 Announce Type: replace Abstract: A trace records what an LLM agent did at each step. What is gained by also recording what each step relied on? GRADE represents a run as one typed graph: execution edges come free from the trace, and dependency edges are supplie…
arXiv cs.CL
TIER_1English(EN)·Kaiyuan Liu, Qiuyang Mang, Bo Peng, Wenhao Chai, Hanchen Li, Shreyas Pimpalgaonkar, Luke Zettlemoyer, Alex Dimakis, Alvin Cheung·
arXiv:2609.15309v1 Announce Type: new Abstract: Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to stop. This test-time strategy makes it difficult to measure how agent performance …
arXiv:2609.12808v2 Announce Type: replace Abstract: Unlearning benchmarks such as TOFU and MUSE certify forgetting by reading the model's final answer, where a model that refuses to answer already counts as having forgotten. We show that this model-level certificate does not tran…
arXiv:2603.19896v2 Announce Type: replace Abstract: Tool-using large language model (LLM) agents often face a fundamental tension between answer quality and execution cost. Fixed workflows are stable but inflexible, while free-form multi-step reasoning methods such as ReAct may i…
arXiv cs.AI
TIER_1English(EN)·Zixiang Liu, Wenrui Liu, Elsie Dai, Wenhan Yu, Lei Yu, Tong Yang, Jinjun Han, Hong Gao·
arXiv:2512.24565v4 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of external tools via the Model Context Protocol (MCP) is considered a future trend. Current MCP evaluation sets suffer from issue…
Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment. This limits their ability to adapt how they gather evidence, invoke …
Large language models are increasingly deployed as interacting agents, making the persistence of their action policies across repeated interaction critical for reliable multi-agent operation. We investigate whether and how agent-generated, non-binding pre-play communication ("che…
arXiv:2609.11957v1 Announce Type: new Abstract: An LLM agent acts on the world by emitting actions: shell commands to run, edits to apply. A wrong action does not always fail loudly; it can fail silently, producing a plausible but incorrect effect that raises no error. We argue t…
arXiv cs.CL
TIER_1English(EN)·Yuli Qiu, Yutong Li, Wei Su, Zeming Liu, Wanxiang Che, Heyan Huang, Haifeng Wang, Yuang Guo·
arXiv:2609.12655v1 Announce Type: new Abstract: Large language model agents are expected to continuously adapt to new tasks and environments over their lifetime by reusing past experience. However, existing memory-based agents struggle to transfer reusable experience across envir…
Elo-per-token analysis reveals that LLM agents initially scale faster than independent sampling but eventually slow, while parallel short sessions improve performance over single long runs.
Hierarchical orchestration, in which a Manager agent reviews worker output and can send it back for revision, is the default coordination pattern in production multi-agent LLM frameworks. Classical organizational theory predicts that the authority link speeds convergence on decis…
Hierarchical orchestration, in which a Manager agent reviews worker output and can send it back for revision, is the default coordination pattern in production multi-agent LLM frameworks. Classical organizational theory predicts that the authority link speeds convergence on decis…
arXiv:2609.11677v1 Announce Type: cross Abstract: Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing harness evolution methods typically rely on …
arXiv:2609.10964v1 Announce Type: new Abstract: Agentic LLM workflows consist of sequences of model turns interleaved with tool interactions, so their end-to-end completion time depends not only on inference speed but also on when ready turns are released. Most runtimes release e…
arXiv cs.LG
TIER_1English(EN)·Asif Pinjari, Mithun Paul Saint-Germain·
arXiv:2609.10892v1 Announce Type: cross Abstract: When an indirect prompt injection succeeds against an LLM agent, the compromise is visible in the agent's own behavior: a benign prefix of tool calls, a poisoned observation, and a suffix of actions that serve the attacker. An ope…
arXiv cs.AI
TIER_1English(EN)·Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary·
arXiv:2609.09853v1 Announce Type: new Abstract: LLM agents for enterprise systems of record cannot be evaluated on customer production data, and no existing substitute provides ground truth. We present the Era by Eon Benchmark for evaluating LLM agents that use enterprise tools. …
Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing harness evolution methods typically rely on iterative search, repeatedly evaluating and revisi…
arXiv cs.AI
TIER_1English(EN)·Asif Pinjari, Mithun Paul Saint-Germain·
arXiv:2609.06972v1 Announce Type: cross Abstract: LLM agents complete tasks by issuing sequences of tool calls, and every observation they read is a channel through which an indirect prompt injection can enter. A successful injection has a characteristic shape when the trajectory…
arXiv:2510.15047v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) as agents often fail to improve in new environments. We identify and characterize a failure mode we call exploration collapse: under reinforcement learning (RL) in environments whose states are…
arXiv cs.AI
TIER_1English(EN)·Jingyu Zhang, Tianjian Li, William Jurayj, Hongyuan Zhan, Benjamin Van Durme, Daniel Khashabi·
arXiv:2604.09443v4 Announce Type: replace-cross Abstract: Large language model agents receive instructions from many sources-system messages, user prompts, tool outputs, other agents, and more-each carrying different levels of trust and authority. When these instructions conflict…
arXiv:2603.18074v2 Announce Type: replace-cross Abstract: Technical-service LLM agents are entering production workflows, where value depends on whether engineers adopt generated replies. Service tickets hide decision logic, contain noisy single-reference responses, and make rewa…
arXiv:2609.08589v1 Announce Type: cross Abstract: Recent large language models can emit task-progress signals that agent frameworks use to decide whether a task should continue or stop, yet whether a model can reliably report its task progress at every stage of a task, and where …
arXiv:2609.06783v1 Announce Type: cross Abstract: LLM agents operate in workflows where unsafe actions can have real consequences. Existing safety evaluations often reduce behavior to a single score, obscuring risk recognition, pre-action detection, and safe task completion when …
arXiv cs.AI
TIER_1English(EN)·Md Jafrin Hossain, Nur Al Hasan Haldar·
arXiv:2609.05911v1 Announce Type: cross Abstract: Long-horizon LLM agents interact with untrusted content, persistent memory, external state, and sensitive tools. Existing analyses often characterize attacks by the number of execution steps between malicious input and a downstrea…
arXiv:2609.09153v1 Announce Type: new Abstract: Large language models are increasingly deployed as agents that plan over long horizons and act through external tools. Most agents select actions through unconstrained generation over an accumulating history, leaving implicit the pr…
arXiv cs.AI
TIER_1English(EN)·Yuan Gao, Sebastian M\"uller, Mattia Piccinini, Marc Kaufeld, Yuchen Zhang, Finn Rasmus Sch\"afer, Qunying Song, Johannes Betz·
arXiv:2609.08965v1 Announce Type: new Abstract: Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used to validate Autonomous Driving Systems (ADSs), but it remains a fragmented modular pipeline in which scenario gen…
arXiv:2609.08452v1 Announce Type: new Abstract: Multi-agent large language models solve complex tasks by coordinating several policies in a shared environment. However, existing reinforcement learning methods usually optimize each response or trajectory separately, even when seve…
arXiv:2609.07255v1 Announce Type: new Abstract: Language-model agents increasingly rely on skills: reusable procedural knowledge for reasoning, tool use, and interaction. Existing work studies how skills are acquired, retrieved, compressed, or composed, but often assumes that onc…
arXiv cs.AI
TIER_1English(EN)·Katherine Tieu, Dongqi Fu, Yinglong Xia, Hong Li, Hong Yan, Jingrui He·
arXiv:2609.05774v1 Announce Type: new Abstract: Recent multi-agent LLM systems increasingly rely on graph-structured communication to coordinate specialized agents. We revisit multi-agent orchestration from a graph-engineering perspective: rather than optimizing a static topology…
arXiv cs.AI
TIER_1English(EN)·Cen Mia Zhao, Haibo Ruan, Wenjie Chen, Pei-fen Tu, Usman Abbasi, Joel Hesch·
arXiv:2609.05736v2 Announce Type: new Abstract: LLM tool agents can be improved without retraining by modifying the runtime harness around a fixed model: prompts, tool interfaces, middleware, state handling, and recovery logic. We study this setting as resource-bounded harness se…
Large language models are increasingly deployed as agents that plan over long horizons and act through external tools. Most agents select actions through unconstrained generation over an accumulating history, leaving implicit the procedural knowledge of what to do, in what order,…
arXiv cs.MA (Multiagent)
TIER_1English(EN)·Sercan Ö. Arık·
Large language models are increasingly deployed as agents that plan over long horizons and act through external tools. Most agents select actions through unconstrained generation over an accumulating history, leaving implicit the procedural knowledge of what to do, in what order,…
PlannerForge is an LLM-agent framework that unifies all stages of scenario-based autonomous driving testing and improves generation, selection, modification, and planning performance across commercial and open-source models.
Large language models (LLMs) are no longer deployed only for single-turn conversation but increasingly act as agents that plan, call tools, retrieve evidence, maintain memory, and interact over long horizons, often together with other agents through multi-turn conversations. Ther…
arXiv:2609.05261v1 Announce Type: new Abstract: Large language model agents increasingly rely on execution traces to master complex interactive tasks. However, current paradigms are bottlenecked by shallow trajectory retrieval and flat skill summarization, fundamentally ignoring …
arXiv cs.AI
TIER_1English(EN)·Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou·
arXiv:2609.04629v1 Announce Type: new Abstract: A runtime gate for an LLM tool agent is usually cast as a filter. In a ReAct loop a rejected proposal is followed by another at the same state, so the gate is a search operator over the proposal stream whose admission criterion shap…
Traffic analysts must translate diagnosed bottlenecks into executable interventions without allowing local improvements to degrade network-wide performance. This study presents SimTIO, a simulation-grounded multi-agent large language model framework for composing and selecting tr…
arXiv:2609.03383v1 Announce Type: new Abstract: Graph-based policy optimization improves credit assignment for long-horizon LLM agents by organizing rollout trajectories into state-transition graphs. However, existing methods construct graphs independently within each policy upda…
arXiv cs.AI
TIER_1English(EN)·Paul Brookes, Vardan Voskanyan, Rafail Giavrimis, Matthew Truscott, Mina Ilieva, Chrystalla Pavlou, Alexandru Staicu, Manal Adham, Will Evers- Hood, Jingzhi Gong, Kejia Zhang, Matvey Fedoseev, Vishal Sharma, Roman Bauer, Zheng Wang, Hema Nair, Wei Jie, T…·
arXiv:2512.09108v2 Announce Type: replace-cross Abstract: Agentic AI systems built on large language models (LLMs) offer significant potential for automating complex workflows, from software development to customer support. However, LLM agents often underperform due to suboptimal…
arXiv cs.AI
TIER_1English(EN)·Uday Vallabhaneni, Cassie L. Cagwin, David J. Wild·
arXiv:2609.04159v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly proposed as autonomous SOC analysts, but two limitations make them unreliable at enterprise scale: a finite context window cannot hold a multi-thousand-host authentication graph, …
arXiv cs.CL
TIER_1English(EN)·Michael Nguyen, Wei Chen Tan, Nurul Aisyah Hassan, Arvind Raman, Li Hua Lim, Ahmad Faiz Razak·
arXiv:2609.02889v1 Announce Type: new Abstract: A growing body of work improves frozen large language models (LLMs) as agents by evolving their harness: the textual scaffolding around the model, including persona, strategy, format rules, and control heuristics. Existing reflectiv…
Cumulative culture enables humans to preserve, reuse, and extend knowledge and skills across experiences and generations. Inspired by this principle, we introduce \textit{SimSkill}, a self-evolving agent built around the Simulation of Urban MObility (SUMO) traffic simulator. SimS…
arXiv:2609.01736v1 Announce Type: cross Abstract: Large language models (LLMs) augmented with external tools have demonstrated remarkable capability in solving complex real-world tasks. However, existing approaches suffer from two key challenges: brittle multi-step and multi-turn…
arXiv:2609.02094v1 Announce Type: new Abstract: LLM-based multi-agent systems have shown strong performance on complex tasks, yet continual improvement from interaction experience remains challenging. Existing self-reflection methods build experience memories, but memories are mo…
arXiv:2609.02264v1 Announce Type: new Abstract: Adapting the communication topology of an LLM multi-agent system to each query improves both accuracy and efficiency, yet current designers treat this as conditional graph generation: a variational, autoregressive, or diffusion deco…
arXiv:2609.02750v1 Announce Type: new Abstract: Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through textual reflection. Despite strong empirical results, these systems lack a unified account of coordination, memo…
arXiv cs.LG
TIER_1English(EN)·Yanting Yang, Can Jin, Jinman Zhao, Jiahao Wu, Yang Zhou, Zhepeng Wang, Zhendong Wang, Mu Zhou, Dimitris N. Metaxas·
arXiv:2609.02042v1 Announce Type: new Abstract: Large language model (LLM) agents for long-horizon interactive tasks typically follow a ReAct-style protocol, issuing one primitive action per LLM round. While this enables frequent replanning, it is inefficient for long-horizon tas…
Adapting the communication topology of an LLM multi-agent system to each query improves both accuracy and efficiency, yet current designers treat this as conditional graph generation: a variational, autoregressive, or diffusion decoder searches the $N \times N$ adjacency space, a…
arXiv cs.AI
TIER_1English(EN)·Elias Stengel-Eskin, Newton Sander, Carlos Bonetti, Sasha Boguraev, James Bowler, Hale Sirin, Simon Kirby·
arXiv:2609.01491v1 Announce Type: cross Abstract: The growing rate at which LLM agents interact with one another raises key questions about language evolution in multi-LLM-agent settings, with implications for safety and monitorability as well as for linguistic accounts of LLMs. …
arXiv:2609.01437v1 Announce Type: cross Abstract: As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly termed the agent harness. Changing this harness while holding model weights fixe…
arXiv:2511.15755v3 Announce Type: replace Abstract: Large language models (LLMs) promise to accelerate incident response in production systems, yet single-agent approaches generate vague, unusable recommendations. We present MyAntFarm.ai, a reproducible containerized framework de…
arXiv cs.AI
TIER_1English(EN)·Rui Yang, Junjie Xu, Zhengyu Liu, Neil Fendley, Yang Hong, Ziyang Li, Yinzhi Cao·
arXiv:2609.00595v1 Announce Type: cross Abstract: Safe agents can fail together. Multi-agent LLM systems (MAS) move information, state, decisions, and authority across principal boundaries, creating failures that local checks may miss. Without an execution-level view, a multi-age…
arXiv cs.AI
TIER_1English(EN)·Jun Hou, Priya Pitre, Yi Fang, Xuan Wang·
arXiv:2609.01360v1 Announce Type: new Abstract: Large language model (LLM) agent failures often contain multiple related errors rather than a single mistake. Existing attribution methods usually identify a responsible agent, step, or root cause, but do not explicitly model depend…
arXiv cs.AI
TIER_1English(EN)·Wentao Zhang, Syed Shariyar Murtaza, Junaid Ahmad Bhatti, Utkarsh Soni, Yifan Nie, Eugene Wen, Yuntian Deng·
arXiv:2609.00621v1 Announce Type: new Abstract: Prompt optimization can improve multi-agent LLM systems, but the prompts being optimized often serve two entangled roles: generating task-relevant content and specifying execution-critical protocols, such as message routing, output …
arXiv:2609.00455v1 Announce Type: new Abstract: Large language models (LLMs) are being used as policies for autonomous decision-making and planning in many domains. Despite their strong reasoning capabilities, LLMs struggle with long-horizon tasks, especially under partial observ…
arXiv cs.AI
TIER_1English(EN)·Rakibul Hasan Rajib, Mengxing Zheng, Qian Lou·
arXiv:2609.00237v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent systems tackle complex reasoning by orchestrating how multiple agents are configured and how they collaborate. A central challenge is to adapt orchestration to the evolving collaboration …
Large language model (LLM) agents for long-horizon interactive tasks typically follow a ReAct-style protocol, issuing one primitive action per LLM round. While this enables frequent replanning, it is inefficient for long-horizon tasks where many rounds are spent on routine action…
The study formalizes multi-agent LLM coordination via bilevel games and stochastic memory reflection, introducing a grounded evaluation gate and SRMA algorithm with convergence guarantees, validated on SWE-bench.
Large language models (LLMs) augmented with external tools have demonstrated remarkable capability in solving complex real-world tasks. However, existing approaches suffer from two key challenges: brittle multi-step and multi-turn reasoning caused by incompatible tool output type…
The growing rate at which LLM agents interact with one another raises key questions about language evolution in multi-LLM-agent settings, with implications for safety and monitorability as well as for linguistic accounts of LLMs. To address these questions, we introduce GlossoGen…
arXiv cs.AI
TIER_1English(EN)·Wujie Xiong, Rabimba Karanjai, Yang Lu, Weidong Shi, Lei Xu·
arXiv:2608.30041v1 Announce Type: cross Abstract: Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later privileged actions. Existing defenses mainly classify untrusted content or authorize…
arXiv cs.AI
TIER_1English(EN)·Hyomin Lee, Sangwoo Park, Yumin Choi, Sohyun An, Seanie Lee, Sung Ju Hwang·
arXiv:2603.22341v2 Announce Type: replace-cross Abstract: While prior red-teaming efforts have focused on eliciting harmful text outputs from large language models (LLMs), such approaches fail to capture agent-specific vulnerabilities that emerge through multi-step tool execution…
arXiv:2602.13671v2 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, existing works often rely on manual designs or "one-size-fits-all" automation and lack adaptabilit…
arXiv:2608.29309v1 Announce Type: cross Abstract: This work provides an overview of the different strategies that can be used to evaluate the performance of AI models and agents based on large language models (LLMs) for materials synthesis. After providing a brief overview of the…
Prompt optimization can improve multi-agent LLM systems, but the prompts being optimized often serve two entangled roles: generating task-relevant content and specifying execution-critical protocols, such as message routing, output formatting, and termination signals, on which th…
The framework separates structured execution protocols from optimizable language content to prevent prompt optimization from corrupting multi-agent pipelines.
HarnessDev evaluates agents by measuring their ability to build and iteratively improve execution infrastructure rather than final task outputs, revealing that self-built harnesses vary widely in capability and efficiency and transfer poorly across models.
Large language models (LLMs) are increasingly used in software development, but their use in complex hardware design remains limited. This gap stems from both the scarcity of public hardware training data and the fundamentally different methodologies used in hardware design. In p…
Language models have recently been applied to nanophotonic design, but it remains unclear whether they can reliably translate optical objectives into simulation-ready designs, execute electromagnetic analysis, and revise decisions from numerical feedback. We introduce HALO, a phy…
arXiv:2606.07867v2 Announce Type: replace Abstract: Are tool-calling LLM agents equally safe throughout a conversation? We discover they are not: agents are most vulnerable at the very start of a session and become substantially safer after a few regular agentic tasks -- a phenom…
arXiv cs.CL
TIER_1English(EN)·Tatiana Petrova, Andrei Mazniak, Radu State·
arXiv:2608.26130v1 Announce Type: new Abstract: Coding agents built on large language models (LLMs), such as Claude Code, Cursor, OpenAI Codex, GitHub Copilot, and Aider, receive tool responses that routinely exceed the agent's per-turn token budget. The standard remedy, paginati…
arXiv:2608.27427v1 Announce Type: cross Abstract: Large language model (LLM) agents in governed organizations must let the persona (instructions, tone, self-presentation) evolve freely, while keeping execution (stateful, audited work) traceable. A single trust domain does not sat…
arXiv:2608.27260v1 Announce Type: new Abstract: LLM agents increasingly rely on generated interaction data to learn how to interact with external environments. Agentic data generation must maintain consistency among environments, tasks, interactions, and success signals while pro…
arXiv:2608.26867v1 Announce Type: new Abstract: Large language model agents reason, call tools, and act autonomously over many steps, but their agentic skills-correctly sequencing tools, planning under dependencies, judging untrusted inputs, and grounding generated arguments-are …
arXiv:2608.27021v1 Announce Type: cross Abstract: LLM-based agents are increasingly proposed for network fault diagnosis, but existing benchmarks evaluate them only on accurate tickets and always assume a fault is present, conditions rarely met in practice. We present FaulT-Bench…
arXiv:2608.26882v1 Announce Type: cross Abstract: Industrial control systems (ICSs) rely on programmable logic controllers (PLCs) to connect networked computation with physical control. Tool-using large language model (LLM) agents represent an emerging attack threat: can an auton…
arXiv cs.AI
TIER_1English(EN)·Kimberly Milner, Minghao Shao, Nanda Rani, Haoran Xi, Venkata Sai Charan Putrevu, Meet Udeshi, Sandeep K. Shukla, Prashanth Krishnamurthy, Farshad Khorrami, Muhammad Shafique, Ramesh Karri·
arXiv:2608.26237v1 Announce Type: cross Abstract: Capture-the-Flag (CTF) benchmarks are widely used to assess the offensive security capabilities of autonomous language-model agents. Evaluations rely on shallow binary judgments or aggregate scores, overlooking the agent's traject…
arXiv:2608.26117v1 Announce Type: cross Abstract: RDBMS (Relational Database Management System) databases face several limitations, including slow execution with multi-hop queries and a lack of explainability by graphical interpretations. In contrast, Graph database offers a more…
arXiv:2608.26226v1 Announce Type: new Abstract: LLM-based agents are increasingly being developed for time-series problems, but their design choices vary substantially across task settings. This survey adopts a problem-driven taxonomy that organizes these systems by the time-seri…
arXiv cs.AI
TIER_1English(EN)·Abhigya Verma, Amit Kumar Saha, Seganrasan Subramanian, Sai Harshitha Aluru·
arXiv:2608.26623v1 Announce Type: new Abstract: LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJudgeBench, the first benchmark to systematically stud…
arXiv cs.AI
TIER_1English(EN)·Haiteng Wang, Weihao Li, Jing Zhang, Lei Ren·
arXiv:2608.26780v1 Announce Type: new Abstract: Control system design is critical for modern industry, such as chemical process temperature regulation and aero-engine control. However,traditional control design workflows rely heavily on expert knowledge and extensive manual param…
arXiv cs.AI
TIER_1English(EN)·Jiayi Kuang, Yinghui Li, Yunze Song, Keyu Chen, Zhifeng Shen, Yangning Li, Yidong Wang, Di Yin, Ruizhi Qiao, Xing Sun, Kai Jin, Ying Shen, Liang Lin, Philip S. Yu·
arXiv:2608.26950v1 Announce Type: new Abstract: Large Language Models (LLMs) are evolving from performing end-to-end mathematical reasoning to integrating agentic intelligence. However, most existing math benchmarks evaluate only final answers. This outcome-oriented evaluation pr…
arXiv:2608.26990v1 Announce Type: new Abstract: Large language models can summarize financial information, but an operational stock-research system must first assemble heterogeneous evidence, expose unavailable data and model capabilities, and control how generated opinions affec…
arXiv:2608.27146v1 Announce Type: new Abstract: Tool-augmented LLM agents must rely on untrusted runtime Observations to complete open-ended tasks; however, when tool outputs no longer merely provide data but begin to specify concrete actions, they effectively become ``commands''…
arXiv:2608.27141v1 Announce Type: cross Abstract: Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unatten…
EvoUndo evaluates recoverability of self-modifying LLM agents and shows that reliable recovery requires co-designing verification, state grounding, and recovery-language expressivity.
Large language models can summarize financial information, but an operational stock-research system must first assemble heterogeneous evidence, expose unavailable data and model capabilities, and control how generated opinions affect a final report. We present DSA, an evidence-aw…
Industrial control systems (ICSs) rely on programmable logic controllers (PLCs) to connect networked computation with physical control. Tool-using large language model (LLM) agents represent an emerging attack threat: can an autonomous agent convert a network-reachable PLC into s…
arXiv:2608.25570v1 Announce Type: new Abstract: Hardware kernel optimization requires repeated compilation, correctness testing, profiling, and revision. LLM agents can automate parts of this process, and stronger foundation models, longer context windows, and longer execution ho…
arXiv cs.CL
TIER_1English(EN)·Hongqiu Ni, Han Tian, Chi Zhang, Guopeng Li, Haisheng Tan·
arXiv:2608.25523v1 Announce Type: new Abstract: Prefix caching introduces a fundamental tradeoff in multi-agent large language model (LLM) serving: retaining a long system-prompt key-value (KV) cache for an agent accelerates future calls, yet it reduces the GPU memory available f…
arXiv:2608.25277v1 Announce Type: new Abstract: Multi-agent LLM systems coordinate through natural-language messages that consume 40--60\% of their token budget. Replacing these with structured graphs reduces cost but fails on tasks requiring adaptive reasoning. We propose \textb…
Agentic data generation is framed as constrained distribution design over factorized experience tuples, emphasizing execution-grounded accuracy, learner-relative complexity, and diversity rather than scale alone.
Multi-agent large language model (LLM) workflows have emerged as a powerful paradigm for solving complex, open-ended tasks through collaborative reasoning among specialized LLM agents, but they incur substantial operating costs due to repeated LLM invocations and long-horizon con…
Multi-agent systems (MAS) sometimes already have the potential to answer correctly, but still report a wrong answer. Explaining this outcome is difficult because generation, communication and final answer-selection rules usually change simultaneously. We conceptualize multi-agent…
Hardware kernel optimization requires repeated compilation, correctness testing, profiling, and revision. LLM agents can automate parts of this process, and stronger foundation models, longer context windows, and longer execution horizons have improved optimization within individ…
arXiv:2604.01532v3 Announce Type: replace Abstract: LLM agents are beginning to invoke industrial asset-management tools through the Model Context Protocol (MCP), yet whether they can act reliably on this substrate for safety-critical \emph{Prognostics and Health Management (PHM)…
arXiv cs.AI
TIER_1English(EN)·Yuchen Xia, Michael Weyrich, Nasser Jazdi, Johannes St\"umpfle, Johannes Sigel, Akshay Narla, Gavin K. Reynolds, Anna Jawor-Baczynska, Pol Llopart·
arXiv:2608.23622v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require more than plausible text and code generation. They require understanding how a syste…
arXiv cs.AI
TIER_1English(EN)·Roy Ganz, Mor Shpigel Nacson, Adi Kalyanpur, Ron Litman·
arXiv:2608.24358v1 Announce Type: new Abstract: Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As these runs unfold, users face a practical cost-quality trade-off: escalating to a stronger model when a cheaper one struggles, or…
arXiv:2608.24569v1 Announce Type: new Abstract: Large language model (LLM) agents coordinate complex tasks through multi-role and multi-stage workflows. Upstream state is repeatedly transformed into intermediate language artifacts, such as summaries, plans, tickets, memories, and…
arXiv:2608.23953v1 Announce Type: cross Abstract: An agent harness is what turns a language model into an autonomous agent: the surrounding code that builds the model's context, mediates its tools, runs the loop, and persists state across a long-horizon run. This layer, not the m…
arXiv:2608.24017v1 Announce Type: cross Abstract: The emerging W3C WebMCP proposal enables LLM agents to invoke tools exposed by web pages. In multi-party web environments, however, integrating agent execution into a browser security model centered on the Same-Origin Policy (SOP)…
arXiv cs.AI
TIER_1English(EN)·Wonung Kim, Hyunmin Choi, Minsu Kim, Jaehong Cho, Yeongwook Kim, Jongse Park·
arXiv:2608.24650v1 Announce Type: cross Abstract: System-level simulation is an essential tool for exploring the rapidly expanding design space of LLM serving systems, where real deployments remain costly and often infeasible. However, modern LLM serving now evolves faster than h…
arXiv:2605.10059v3 Announce Type: replace Abstract: Agent-based modeling (ABM) has long been used in economics to study human behavior, and large language model (LLM) agents now enable new forms of social and economic simulation. While prior work has discovered strategic deceptio…
Large language model (LLM) agents coordinate complex tasks through multi-role and multi-stage workflows. Upstream state is repeatedly transformed into intermediate language artifacts, such as summaries, plans, tickets, memories, and handoff notes, from which downstream components…
arXiv:2511.12484v2 Announce Type: replace-cross Abstract: With the integration of massive distributed energy resources and the widespread participation of novel market entities, the operation of active distribution networks (ADNs) is progressively evolving into a complex, multi-s…
arXiv:2603.02229v2 Announce Type: replace-cross Abstract: Safety post-training has been studied extensively in single-step "chat" settings where safety typically refers to refusing harmful requests. We study an "agentic" (i.e., multi-step, tool-use) setting where safety refers to…
arXiv:2608.22130v1 Announce Type: cross Abstract: LLM-based multi-agent systems (MAS) solve complex tasks through communication among role-specialized agents. However, inter-agent dependencies introduce reliability risks beyond isolated agent failures. For instance, errors in int…
arXiv:2608.22566v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems are designed to improve reasoning by decomposing tasks across multiple agents with specialized functions, but the presence of multiple agents does not inherently guarantee coherent reas…
arXiv:2608.22152v1 Announce Type: new Abstract: Multi-agent systems built from large language models are deployed widely, yet how much performance is lost when two LLMs must coordinate rather than act alone remains unclear. We formulate the collaboration tax as the team-decentral…
arXiv:2603.29902v2 Announce Type: replace Abstract: Interleaved text-and-image generation represents a significant frontier for Multimodal Large Language Models (MLLMs), offering a more intuitive way to convey complex information. Current paradigms rely on either image generation…
arXiv cs.AI
TIER_1English(EN)·Weiwei Sun, Xuhui Zhou, Weihua Du, Xingyao Wang, Sean Welleck, Graham Neubig, Maarten Sap, Yiming Yang·
arXiv:2511.02208v2 Announce Type: replace Abstract: Despite rapid progress, current AI agents are primarily optimized for isolated task completion. We argue for a paradigm shift toward training agents as collaborators that communicate and adapt to people. To facilitate this shift…
arXiv:2608.23179v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly attractive for automating network configuration, yet their reliability and failure patterns are poorly understood. An essential prerequisite is to assess such agents in a realisti…
arXiv:2608.23104v1 Announce Type: cross Abstract: Molecular science represents an important frontier for LLM-based agents. Unlike general agents that mainly operate over natural language, code, or web environments, molecular LLM agents must perceive, reason about, and act upon ch…
arXiv:2608.22833v1 Announce Type: cross Abstract: Large language models (LLMs) and vision-language models (VLMs) are expanding the range of behaviors that can be represented in agent-based simulations, but many contemporary platforms are difficult to study, modify, or run on ordi…
arXiv:2608.22793v1 Announce Type: cross Abstract: Reliable deployment of LLM agents in user-facing products depends not on raw task-solving ability but on consistency and limit-awareness: behaving the same way across repeated trials, and recognizing when a request cannot, or cann…
arXiv cs.AI
TIER_1English(EN)·Xinyuan Liu, Eren Sadikoglu, Riana Chatterjee, Ransalu Senanayake·
arXiv:2608.22657v1 Announce Type: cross Abstract: Agentic AI frameworks interpret open-ended task goals and decompose them into multi-step plans. Richer information about embodiment-specific capabilities, physical preconditions, and cross-robot coordination improves grounding, bu…
arXiv:2608.21544v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as tool-augmented agents, where responses can depend on tool calls and external observations rather than model parameters alone. This creates an evaluation mismatch for LLM un…
arXiv cs.AI
TIER_1English(EN)·Israt Moyeen Noumi, Tarannum Ahmed Nowshin, Md. Mehedi Hasan Nipu, Mohammad Sakib Mahmood, Md. Jakir Hossain, M. F. Mridha·
arXiv:2608.21423v1 Announce Type: cross Abstract: Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools. As these systems move from demonstrations to deployed products, practitioners repeatedly encounter the same operational failu…
arXiv cs.AI
TIER_1English(EN)·Seth Karten, Alex L. Zhang, Kevin Thomas, Sebastian M\"uller, Elie Bakouch, Daniel Auras, Mika Senghaas, Fares Obeid, Konstantin Dunas, Johannes Hagemann, Sami Jaghouar·
arXiv:2608.23552v1 Announce Type: new Abstract: Language models are sequential processors, but long-horizon agency requires external information and computation beyond model weights and active context. Prime Agent is an open-source harness for long-horizon evaluation and coding-a…
arXiv:2608.23058v1 Announce Type: new Abstract: Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external tools, and iterative prediction. We investigate LLM-based forecasting agents, meanin…
arXiv:2608.22974v1 Announce Type: new Abstract: Large language model (LLM) agents rely heavily on knowledge encoded in model parameters or presented as unstructured context. In domain-specific tasks, this leaves important semantic connections implicit. This often results in incom…
arXiv:2608.21836v1 Announce Type: new Abstract: Large language models have become increasingly capable agents for low-level code and kernel optimization, but isolated kernel benchmarks provide only a proxy for the deployment behavior that matters in language-model inference. We i…
An agent harness is what turns a language model into an autonomous agent: the surrounding code that builds the model's context, mediates its tools, runs the loop, and persists state across a long-horizon run. This layer, not the model it wraps, is increasingly the binding constra…
Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As these runs unfold, users face a practical cost-quality trade-off: escalating to a stronger model when a cheaper one struggles, or downshifting once the hard reasoning is complet…
As LLM agents proliferate, built by different parties and with different capabilities and costs, orchestrating them is more like assembling labor across the economy than a computer calling a subroutine. Existing orchestration is typically centralized, with a single planner assign…
Molecular science represents an important frontier for LLM-based agents. Unlike general agents that mainly operate over natural language, code, or web environments, molecular LLM agents must perceive, reason about, and act upon chemical objects across symbolic strings, molecular …
Large language models (LLMs) and vision-language models (VLMs) are expanding the range of behaviors that can be represented in agent-based simulations, but many contemporary platforms are difficult to study, modify, or run on ordinary computers. We present two intentionally minim…
arXiv:2608.21156v1 Announce Type: cross Abstract: LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolution has produced paradigms including Prompt Engineering to elicit model capabilities, Context Engineering to manage…
arXiv:2608.20631v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated the ability to solve multi-step tasks requiring planning, tool use, and external information access, yet growing execution histories increase inference cost and expose reasoning to…
arXiv:2608.20729v1 Announce Type: new Abstract: Language-model agents can improve after failure or carry text across episodes without revising what counts as success. We study the narrower attribution problem of criterion revision: when criterion K0 accepts an outcome violating a…
arXiv:2608.21027v1 Announce Type: new Abstract: LLM agents are emerging as an important paradigm for real-world tasks that require reasoning, tool use, and sequential decision-making. As these agents operate over longer horizons, runtime intervention offers a way to improve relia…
arXiv:2608.20563v1 Announce Type: cross Abstract: Long-horizon security LLM agents must carry information and decisions across many dependent interactions, where later actions often depend on services, state, or access discovered much earlier. This makes final task success diffic…
arXiv cs.AI
TIER_1English(EN)·Jiajun Wu, Zirui Wang, Jiayu Zhou, Qiang Ye, Steve Drew·
arXiv:2608.20518v1 Announce Type: new Abstract: In Federated Learning (FL), the communication topology is a runtime variable rather than a fixed design choice, since links and edge devices drop in and out during training. Each round, the server must commit three coupled decisions…
arXiv:2608.21101v1 Announce Type: cross Abstract: As large language model (LLM) agents move from conversation to executing code, reading local files, and orchestrating external tools, a single agent hijacked by a malicious third-party skill can cause data exfiltration, privilege …
arXiv cs.LG
TIER_1English(EN)·Quan Wei, Siliang Zeng, Chenliang Li, Zhongruo Wang, William Brown, Oana Frunza, Wei Deng, Anderson Schneider, Yuriy Nevmyvaka, Yang Katie Zhao, Alfredo Garcia, Mingyi Hong·
arXiv:2505.11821v3 Announce Type: replace Abstract: Reinforcement Learning (RL) approaches have been wildly used to enhance the reasoning capabilities of Large Language Model (LLM) agents in long-horizon, multi-turn scenarios. Such interactions can be formalized as turn-level Mar…
Prime Agent is an open-source harness that uses recursive subagents, persistent computation, and agent-to-agent coordination to extend language models' long-horizon capabilities across coding and reasoning tasks.
Agentic AI frameworks interpret open-ended task goals and decompose them into multi-step plans. Richer information about embodiment-specific capabilities, physical preconditions, and cross-robot coordination improves grounding, but does not eliminate infeasible, mistimed, or unsa…
LLM-based multi-agent systems (MAS) solve complex tasks through communication among role-specialized agents. However, inter-agent dependencies introduce reliability risks beyond isolated agent failures. For instance, errors in intermediate messages could be inherited and amplifie…
Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require more than plausible text and code generation. They require understanding how a system responds to intervention, which in practice de…
The core challenge in group recommendation lies in modeling the dynamic evolution of user preferences and explain?ing the consensus formation process. Existing Large Language Model (LLM)-based methods, despite improved interpretability, treat interaction history as fixed text, ig…
LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolution has produced paradigms including Prompt Engineering to elicit model capabilities, Context Engineering to manage information access, Harness Engineering to organi…
LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolution has produced paradigms including Prompt Engineering to elicit model capabilities, Context Engineering to manage information access, Harness Engineering to organi…
arXiv cs.CL
TIER_1English(EN)·Mehdi Azarafza, Faezeh Pasandideh, Ali Ehteshami Bejnordi, Stefan Henkler, Achim Rettberg·
arXiv:2608.20129v1 Announce Type: cross Abstract: Autonomous vehicles require robust perception and decision-making capabilities to operate in diverse and unseen scenarios. While reinforcement learning and rule-based methods can provide effective control and safety mechanisms, th…
arXiv cs.CL
TIER_1English(EN)·Hexi Wang, Yujia Zhou, Bangde Du, Weihang Su, Xinyuan Cao, Qingyi Pan, Qingyao Ai, Yueyue Wu, Min Zhang, Yiqun Liu·
arXiv:2608.19621v1 Announce Type: new Abstract: Large language models (LLMs) offer a scalable approach to social simulation, but their credibility depends on how agents are constructed. Existing methods can partially reproduce population-level patterns, yet often fail to capture …
arXiv:2608.19803v1 Announce Type: cross Abstract: Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision often comes only from final rewards. Existing methods refine trajectory-level signals into step-level credits through step grouping…
arXiv:2608.20318v1 Announce Type: new Abstract: Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inherits the improvement. That process is the training algorithm: a better objective or update rule…
arXiv:2608.20274v1 Announce Type: new Abstract: Large language model (LLM) agents can induce skills from completed tasks and reuse them later to grow more capable with experience. In practice, induced skills may transfer unreliably and can even harm the agent that retrieves them.…
arXiv:2608.19993v1 Announce Type: new Abstract: Loading reusable skill documents into a bounded context window is now the primary way large language model (LLM) agents acquire task-specific capabilities, which makes skill selection a first-order determinant of task performance an…
arXiv cs.AI
TIER_1English(EN)·Yiyang Luo, Yihang Jiang, Qijun Xie, Liang Lan, Lin Willian Cong, Anyi Rao, Yunya Song·
arXiv:2608.19974v1 Announce Type: new Abstract: LLM agents in financial markets may cite rules yet still submit orders that violate executable constraints or misread surveillance evidence. We introduce ReguSim, a controlled financial-compliance environment, and ReguBench, a targe…
arXiv cs.AI
TIER_1English(EN)·Seongjae Kang, Taehyung Yu, Sung Ju Hwang·
arXiv:2608.19861v1 Announce Type: new Abstract: Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures arise from either forbidden actions, such as granting an ineligible change, or omitted procedural requirements, such a…
Graph Engineering organizes multi-agent LLM systems through dynamic graph structures to coordinate specialized agents and manage complex, evolving tasks.
Autonomous vehicles require robust perception and decision-making capabilities to operate in diverse and unseen scenarios. While reinforcement learning and rule-based methods can provide effective control and safety mechanisms, their performance may degrade in situations requirin…
Autonomous vehicles require robust perception and decision-making capabilities to operate in diverse and unseen scenarios. While reinforcement learning and rule-based methods can provide effective control and safety mechanisms, their performance may degrade in situations requirin…
arXiv:2608.18575v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent systems (MAS) often exhibit complex failure modes, which frequently cause agents to produce incorrect outcomes. This motivates the task of Agent Failure Attribution: given a failed multi-…
Autonomous vehicles offload latency-sensitive perception tasks to nearby mobile edge computing (MEC) servers, where a missed safety-critical task is unsafe rather than merely degraded. Large language models (LLMs) are increasingly proposed as adaptive, explainable schedulers, yet…
Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures arise from either forbidden actions, such as granting an ineligible change, or omitted procedural requirements, such as identification or confirmation. Runtime safegu…
arXiv:2608.17433v1 Announce Type: new Abstract: LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take…
arXiv:2608.17310v1 Announce Type: new Abstract: Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyw…
arXiv cs.AI
TIER_1English(EN)·Dayang Liang, Liyuan He, Xuan Feng, Shuxin Li, Bo An, Yunlong Liu·
arXiv:2608.17289v1 Announce Type: new Abstract: Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactive tasks. However, most existing variants fail to distinguish advantages among successful traj…
arXiv cs.MA (Multiagent)
TIER_1English(EN)·Sandeep P. Chinchali·
Multi-agent Large Language Model (LLM) systems often struggle to collaborate with new teammates whose strategies shift mid-task. Because agents execute multi-step or temporally extended skills, they frequently continue executing outdated plans long after public evidence shows tha…
LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take. Existing systems often expose the same compreh…
arXiv:2608.16002v1 Announce Type: cross Abstract: Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments. Existing UQ methods largely rely on local signals, such as token probabilities, predictive …
arXiv:2608.14559v1 Announce Type: new Abstract: Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when? Existing approaches either communicate at every timestep or learn a binary gate through REINFOR…
arXiv:2608.14588v1 Announce Type: new Abstract: Sequential multi-agent LLM pipelines chain specialized agents without verification at handoffs, creating a structural flaw with measurable and severe consequences. We show that hallucinations injected at Stage 1 do not merely persis…
arXiv cs.AI
TIER_1English(EN)·Pierre Dantas, Lucas Cordeiro, Ehsan Nowroozi, Tihanyi Norbert·
arXiv:2608.14590v1 Announce Type: new Abstract: LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use of tools. However, no existing system provides formally grounded, task-level safety guarante…
arXiv:2608.14613v1 Announce Type: new Abstract: Modern LLM-agent frameworks increasingly interoperate through standards such as Anthropic's Model Context Protocol (MCP) for agent-to-tool access and Google's Agent2Agent (A2A) protocol for agent delegation and negotiation. However,…
arXiv:2608.15391v1 Announce Type: new Abstract: Large language model (LLM)-assisted energy-management tools can translate natural-language context into structured grid commands, but syntactic validity does not imply physical admissibility. This paper presents TwinGridShield, a mo…
arXiv:2608.15591v1 Announce Type: new Abstract: Large Language Model (LLM) agents deployed in production environments face a fundamental tension: the agent's behavior is frozen at deployment time, while the business rules and edge cases it must handle continue to evolve. Existing…
arXiv:2608.16637v1 Announce Type: new Abstract: LLMs remain unreliable for long-horizon planning, often generating logically inconsistent or non-applicable plans. Recent hybrid methods instead translate natural language into the Planning Domain Definition Language (PDDL), allowin…
arXiv:2608.14573v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly capable of tool use, code execution, artifact inspection, and iterative revision, creating new opportunities for automating scientific and engineering research. To the best of our…
arXiv cs.AI
TIER_1English(EN)·Kaixiang Wang, Yidan Lin, Jiong Lou, Jie Li·
arXiv:2608.14668v1 Announce Type: cross Abstract: LLM-based multi-agent systems (LLM-MAS) solve complex tasks through specialized collaboration, but inter-agent dependencies can propagate hallucinated or malicious outputs into system-level failures. Auditor agents mitigate these …
arXiv cs.AI
TIER_1English(EN)·Zeyuan Li (Massachusetts Institute of Technology), Lukas Petersson (Andon Labs), Alessandro Acquisti (Massachusetts Institute of Technology), Michiel A. Bakker (Massachusetts Institute of Technology)·
arXiv:2608.14825v1 Announce Type: cross Abstract: Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elicitation ev…
arXiv:2608.15108v1 Announce Type: cross Abstract: Large language model agents are increasingly connected to high-value resources such as computing infrastructure, credentials, usage budgets, identities, private knowledge, communication channels, and organizational workflows. Exis…
arXiv:2608.15549v1 Announce Type: cross Abstract: Programming small social robots from natural-language instructions requires more than invoking isolated APIs. Interactive tasks combine reactive physical behaviors with stateful social behaviors, while existing interfaces often re…
arXiv cs.AI
TIER_1English(EN)·Victor Ye Dong, Reid Pryzant, Yi Liu, Jian Jiao·
arXiv:2608.16068v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as agents that rely on system prompts to use tools and complete tasks. Such deployments impose distinct operational requirements, including appropriate tool use, concise promp…
arXiv:2608.16114v1 Announce Type: new Abstract: As agentic tasks grow in complexity, LLM agents increasingly rely on experiential memory to reuse procedural knowledge across tasks. Effective memory design must jointly address what to store, how memory is structured and retrieved,…
arXiv:2608.14635v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly trained with reinforcement learning in long-horizon, sandboxed environments. Unlike conventional RL, agentic RL couples GPU-intensive rollout engines with stateful environment con…
Agentic ESOpt uses evolution strategies for scalable full-parameter fine-tuning of long-horizon LLM agents via trajectory-level reward-weighted updates and parameter-context co-evolution.
Recent advances in large language models have enabled a new class of agentic data science systems that allow users to complete complex data science workflows through natural language. Although these systems can significantly reduce manual effort, it remains difficult to diagnose …
arXiv:2608.13571v1 Announce Type: cross Abstract: When a language model fails to answer a query on the first attempt, an agentic system retries, consuming additional tokens each time. This retry overhead creates a gap between what a model's per-token price implies and what a full…
arXiv cs.AI
TIER_1English(EN)·Ismail El Hamraoui, Sagar Jose, Nicolas Bureau, Robert Plana·
arXiv:2608.14109v1 Announce Type: new Abstract: Autonomous LLM agents are increasingly deployed in complex real-world workflows, yet they remain vulnerable to runtime behavioral drift, a silent deviation from the original task that can lead to irreversible side effects on externa…
arXiv cs.AI
TIER_1English(EN)·Zhizhao Guan, Chen Huang, Ziming Liu, Hongru Liang, Wenqiang Lei, See-Kiong Ng, Tat-Seng Chua, Anthony G Cohn·
arXiv:2608.14339v1 Announce Type: new Abstract: We study proactive exploration in LLM agents, i.e., the ability to explore an environment to acquire information that improves future decision-making. In this regard, we first identify two fundamental bottlenecks that hinder this ca…
arXiv:2608.14380v1 Announce Type: new Abstract: Many real-world tasks require LLM agents to interact with their environments over long execution horizons. Errors that occur early in execution may propagate through both the agent context and environment state, and their effects ma…
arXiv cs.AI
TIER_1English(EN)·Xiaofan Zhou, Huy Nguyen, Bo Yu, Chenxi Liu, Lu Cheng·
arXiv:2604.01413v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) increasingly rely on multi-turn reasoning and interaction, such as adaptive retrieval-augmented generation (RAG) and ReAct-style agents, to answer difficult questions. These methods improve acc…
arXiv cs.LG
TIER_1English(EN)·Ignacio D. Lopez-Miguel, Andreas Happe, J\"urgen Cito, Ezio Bartocci, Bettina K\"onighofer, Martin Tappler·
arXiv:2608.14352v1 Announce Type: cross Abstract: Large Language Model (LLM)-based agents are increasingly used for complex tasks such as software testing and cybersecurity assessment. While these agents demonstrate impressive capabilities, their behavior is difficult to understa…
RUPA models agent execution as a dependency graph to propagate uncertainty across long trajectories, improving failure detection and confidence estimation for LLM agents.
arXiv cs.MA (Multiagent)
TIER_1English(EN)·M. F. Mridha·
Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools. As these systems move from demonstrations to deployed products, practitioners repeatedly encounter the same operational failures. We systematize these failures through a hands…
What happens when you teach an LLM-based agent the scientific method? Motivation: Scientific discovery emerges from cycles of hypothesis, implementation, empirical testing, and feedback. Can this process be automated? We approach automated algorithm design through the lens of the…
Large language model (LLM)-assisted energy-management tools can translate natural-language context into structured grid commands, but syntactic validity does not imply physical admissibility. This paper presents TwinGridShield, a model-independent runtime authorization layer that…
arXiv cs.MA (Multiagent)
TIER_1English(EN)·Michiel A. Bakker·
Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elicitation evaluations on single agents or stylized tasks. Its …
arXiv cs.MA (Multiagent)
TIER_1English(EN)·Michiel A. Bakker·
Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elicitation evaluations on single agents or stylized tasks. Its …
arXiv cs.MA (Multiagent)
TIER_1English(EN)·Michiel A. Bakker·
Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elicitation evaluations on single agents or stylized tasks. Its …
Autonomous LLM agents are increasingly deployed in complex real-world workflows, yet they remain vulnerable to runtime behavioral drift, a silent deviation from the original task that can lead to irreversible side effects on external systems. Existing approaches address drift at …
arXiv:2608.13173v1 Announce Type: new Abstract: Agent skills are crucial external instructions that enable language agents to execute long procedural tasks such as coding or document processing. Existing agent skills are primarily created through human manual crafting or agent ex…
arXiv:2608.12977v1 Announce Type: cross Abstract: The expanding operational capabilities of large language model (LLM) agents introduce sophisticated security threats. Runtime defenses have emerged as an effective approach to mitigating these risks by integrating security mechani…
arXiv:2605.18852v2 Announce Type: replace-cross Abstract: Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and downstream evaluation signals are noisy. Small observed differences can be compara…
arXiv:2608.12921v1 Announce Type: cross Abstract: The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn communication topologies through b…
arXiv cs.AI
TIER_1English(EN)·Yanwen Peng, Delvin Ce Zhang, Xi Wang, Nikolaos Aletras·
arXiv:2608.13317v1 Announce Type: new Abstract: Large language model based multi-agent systems usually communicate in text, i.e., using discrete tokens. However, text introduces a discrete bottleneck. Converting the sender's continuous hidden states into discrete tokens discards …
arXiv:2608.13179v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignment conflates heterogeneous per-turn outcomes into a…
arXiv:2608.12851v1 Announce Type: new Abstract: Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input disappears. Skill evolution makes this failure measurable by…
The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn communication topologies through black-box optimization driven solely by task-level …
The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn communication topologies through black-box optimization driven solely by task-level …
arXiv:2608.11215v1 Announce Type: new Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the co…
arXiv:2608.11552v1 Announce Type: cross Abstract: Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive …
arXiv:2608.12123v1 Announce Type: cross Abstract: LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enough concurrent work for GPU exe…
arXiv:2608.12273v1 Announce Type: cross Abstract: LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning. This progressive-disclosure design exposes two sequential control points to untrusted publi…
arXiv:2608.11878v1 Announce Type: cross Abstract: Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments,…
arXiv:2608.11232v1 Announce Type: cross Abstract: Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs require ground truth from actual code execution. We present Backtrader-Bench, a fram…
arXiv cs.AI
TIER_1English(EN)·Ziqiang Wan, Li Gu, Zhixiang Chi, Zhi Liu, Seyed Mehdi Ayyoubzadeh, Yuanhao Yu, Yang Wang·
arXiv:2608.11434v1 Announce Type: new Abstract: Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark for s…
arXiv:2608.11381v1 Announce Type: new Abstract: We study whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM specialization. Larix maps a 16-lens European listed-real-estate analysis framework to eight lens-…
arXiv:2608.11241v1 Announce Type: new Abstract: Deploying LLM agents into industrial recommender operations exposes a three-way tension we frame as the autonomy-determinism-efficiency trilemma: general autonomy (interpreting operator intent, generating glue code zero-shot), indus…
arXiv:2608.11888v1 Announce Type: new Abstract: Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, including planning, tool use, problem-solving, and validation. Prior work reported mixed results …
arXiv:2608.11679v1 Announce Type: new Abstract: Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems. Even with skilled operators, interpreting anomalies detected within digital twin pipelines is challenging, as the sheer complexity a…
arXiv cs.AI
TIER_1English(EN)·Alexander Liss, Nicholas Desmond, Santiago Gil Gallego·
arXiv:2608.11207v1 Announce Type: new Abstract: When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach…
LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enough concurrent work for GPU execution, and what changes when a GPU-computed route…
Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, including planning, tool use, problem-solving, and validation. Prior work reported mixed results of agent skills: some skills improve task succes…
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefi…
Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems. Even with skilled operators, interpreting anomalies detected within digital twin pipelines is challenging, as the sheer complexity and volume of raw sensor data make thorough analy…
arXiv cs.AI
TIER_1English(EN)·You Lu, Kun Zhang, Bihuan Chen, Xin Peng·
arXiv:2608.10037v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical grounding resource for LLM agents. Existing studies mainly focus on improving the tool-u…
arXiv cs.AI
TIER_1English(EN)·Vassilis Papadopoulos, McNair Shah, Sam Zimmerman, Jack Lindsey·
arXiv:2608.10218v1 Announce Type: new Abstract: AI agents are becoming more autonomous and increasingly interconnected, exposing them to new emergent risks arising from agent-to-agent interaction. One such risk is the spread of mind viruses: ideas or goals that propagate through …
arXiv:2608.09934v1 Announce Type: cross Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deployment is often limited by the computational cost and instability associated with the…
arXiv:2608.09946v1 Announce Type: cross Abstract: Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints. Although LLM agents offer a promising interface for conversational resource navigation, existin…
arXiv:2605.20173v2 Announce Type: replace Abstract: Production LLM agents combine stochastic model outputs with deterministic software systems, yet the boundary between the two is rarely treated as a first-class architectural object. This paper names that boundary the stochastic-…
arXiv:2608.09988v1 Announce Type: cross Abstract: Large language models are increasingly used to read markets, assess risk, and allocate capital. However, reported results for LLM trading agents can be inflated by look-ahead leakage, optimistic execution, and risk mandates that a…
arXiv:2608.10441v1 Announce Type: cross Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth using…
arXiv:2508.08636v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly expected to act as generalist agents capable of solving complex real-world problems. Training such agents, however, requires stable and diverse environments that support repeated int…
arXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these assets through heuristic reflection or raw success counts. Such updates are brit…
Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying que…
ToolHazard is a scalable framework that synthesizes adversarial environments to test LLM agents against indirect prompt injections, revealing vulnerabilities and improving defensive alignment.
Two studies define measurable GPU control gates for LLM-agent services by analyzing concurrent cohort scheduling and on-device routing versus host redispatch.
arXiv:2608.09857v1 Announce Type: cross Abstract: Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on execution rather than verifying the feasibility actions planning models propose…
arXiv cs.LG
TIER_1English(EN)·Ibne Farabi Shihab, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan·
arXiv:2608.08282v1 Announce Type: new Abstract: Tool-using language-model agents face constraints whose meaning changes with observations and prior actions. We study exact sampling from the model distribution conditioned on a hard stateful validator while reusing invalidity certi…
arXiv:2601.18077v3 Announce Type: replace Abstract: Cooperative reasoning under incomplete information remains challenging for both humans and multi-agent systems. The card game Hanabi embodies this challenge, requiring theory-of-mind reasoning and strategic communication. We ben…
arXiv:2608.09292v1 Announce Type: cross Abstract: Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. However, these methods struggle to learn beyond the inherent capability boundary …
arXiv:2608.08239v1 Announce Type: cross Abstract: LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents. Yet agentic routers are evaluated like single-turn routers: by replaying logge…
arXiv:2505.15062v5 Announce Type: replace-cross Abstract: Knowledge extrapolation is the process of inferring novel information by combining and extending existing knowledge that is explicitly available. It is essential for solving complex questions in specialized domains where r…
arXiv cs.AI
TIER_1English(EN)·Bohan Chen, Shivam N. Patel, Richard Hoffmann, Sam Looi, Tony Yue Yu·
arXiv:2608.00326v2 Announce Type: replace Abstract: Tool calling allows large language models (LLMs) to invoke external computation during problem solving, a useful capability in various fields including AI for mathematics. We study this setting through weighted sum-of-squares (S…
arXiv:2604.02478v2 Announce Type: replace Abstract: Deep learning models excel at detecting anomaly patterns in normal data. However, they do not provide a direct solution for anomaly classification and scalability across diverse control systems, frequently failing to distinguish…
arXiv cs.AI
TIER_1English(EN)·Thassilo M. Schiepanski, Nicholas Pi\"el·
arXiv:2508.04412v3 Announce Type: replace Abstract: The advent of large language models (LLMs) has sparked an evolution of autonomous web browsing agents: given a web browsing task and serialised user interface (UI) state, an LLM is expected to suggest input actions that incremen…
arXiv:2406.14373v3 Announce Type: replace Abstract: The emergence of Large Language Models (LLMs) and advancements in Artificial Intelligence (AI) offer an opportunity for computational social science research at scale. Building upon prior explorations of LLM agent design, our wo…
arXiv:2608.08164v1 Announce Type: cross Abstract: Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of structured information and functional behavior from a large teacher model to a smaller stude…
arXiv:2608.07637v1 Announce Type: new Abstract: Long-running molecular simulation campaigns require repeated continuation from saved states, provenance-aware progression, adaptive assessment, and occasional interpretation of workflow conditions that cannot be resolved safely by f…
arXiv:2608.07925v1 Announce Type: new Abstract: EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing LLMs fail to address. This paper presents ZhuLong, an execution-grounded LLM coding agent for PyAether and SKILL that combines AP…
arXiv cs.AI
TIER_1English(EN)·Florentina Voboril, Stefan Szeider·
arXiv:2608.08127v1 Announce Type: new Abstract: The runtime of Constraint Programming (CP) solvers is highly sensitive to modeling choices, such as symmetry breaking, implied constraints, global constraints, constraint reformulation, and variable representation. Improving these c…
arXiv:2608.08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligati…
arXiv:2608.09168v1 Announce Type: new Abstract: Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has substantially improved skill retrieval due to the increasing skill libraries, retrieving a pl…
arXiv:2608.09253v1 Announce Type: new Abstract: LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and tool use. Although skills provide reusable procedural knowledge, agents may still execute them unreliably. Even when an agent…
arXiv cs.AI
TIER_1English(EN)·Fatemeh Seyedin, Adrian Weller, Jinhyuk Yun, Mahmoudreza Babaei·
arXiv:2608.09574v1 Announce Type: new Abstract: LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move from individual tools to participants in multi-agent organizations, an important qu…
arXiv cs.AI
TIER_1English(EN)·Hao Sui, Simeng Qin, Jie Liao, Xiaojun Jia, Bing Chen, Yang Liu·
arXiv:2608.09577v1 Announce Type: new Abstract: Agent skills, bundles of instructions and resources that an LLM agent loads on demand, form an emerging supply chain where a single poisoned skill can persistently compromise every agent that installs it. However, existing skill att…
arXiv:2608.09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety mechanisms often treat the harnes…
arXiv cs.AI
TIER_1English(EN)·Yingpeng Ma, Jianhao Yan, Bei Shi, Ka Hou Kam, Runnan Wang, Xuebo Liu, Yulong Chen, Yue Zhang, Derek F. Wong·
arXiv:2608.08160v1 Announce Type: cross Abstract: The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the critical challenge of maintaini…
Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth using. Our thesis is a distinction that is easy to miss…
Does the cooperative bias documented for Western frontier LLM agents extend to a different alignment lineage, and should the Chinese models that embody it be treated as a single bloc or as distinct laboratories? We study four frontier-tier Chinese models - DeepSeek V4 Pro, Qwen3-…
Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. However, these methods struggle to learn beyond the inherent capability boundary of the agents, since the agents cannot sample corr…
Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has substantially improved skill retrieval due to the increasing skill libraries, retrieving a plausible skill bundle does not guarantee that exe…
arXiv cs.LG
TIER_1English(EN)·Elizaveta D. Moskovskaya, Anton D. Moscowsky·
arXiv:2509.10317v2 Announce Type: replace-cross Abstract: The article describes the development of a hybrid social robot control architecture to overcome the limitations of traditional approaches, where behavior scripts manually synchronize the robot's actions and text, and exist…
arXiv cs.AI
TIER_1English(EN)·Jiacheng Miao, Jin Mu, Guanhua Chen, James Zou·
arXiv:2608.07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can inspect datasets, generate code, and produce analyses e…
arXiv cs.CL
TIER_1English(EN)·Mingguang Chen, Licheng Wang, Bo Qu·
arXiv:2608.06663v1 Announce Type: new Abstract: Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour tasks: losing track of earlier decisions, declaring half-finished work done, or…
arXiv cs.AI
TIER_1English(EN)·Karolina Rudnicka, Thomas Stephan Juzek·
arXiv:2608.06589v1 Announce Type: cross Abstract: While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that this perspective coexists with distinct, model-specific linguistic signatures akin to human id…
Business Arena evaluates LLM agents running a realistic cross-border shop, revealing large performance gaps versus human strategies and enabling detailed attribution of business decisions.
The study introduces a benchmark and formalizes narrative commitment preservation to evaluate long-horizon logical consistency in interactive storytelling with large language models.
arXiv cs.AI
TIER_1English(EN)·Wuya Chen, Yihao yang, Yang Cao, Yue Lin·
arXiv:2608.05886v1 Announce Type: cross Abstract: Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the file to patch, rather than patching it. On SWE-Bench Verified, a 30B OpenHands agent average…
arXiv:2608.05204v1 Announce Type: new Abstract: LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code, tools, references, and operational workflows. As skills become marketplace artifacts, auditin…
arXiv cs.AI
TIER_1English(EN)·Jie Wu, Ming Gong, Feixiang Cheng, Qinqin Zhao·
arXiv:2608.05519v1 Announce Type: new Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local lookup, broad search, composite research tool, stronger model, or human escalation i…
arXiv cs.AI
TIER_1English(EN)·Linfang Shang, Ming Xu, Yiding Sun, Tianle Xia, Lingxiang Hu, Lan Xu, Ning Zheng·
arXiv:2608.05810v1 Announce Type: new Abstract: Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not monotonic: past a critical pool size, newly added skills degrade performance instead of impr…
arXiv cs.AI
TIER_1English(EN)·Zihan Xu, Haolin Tian, Hai Jiang·
arXiv:2608.05791v1 Announce Type: cross Abstract: Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inference, and their execution strategies directly affect system accuracy, latency, and computation…
arXiv:2608.05446v1 Announce Type: cross Abstract: Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, invoke tools, verify outcomes, and reuse experience across interactions. However, effective harness use raises two coupled …
Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the file to patch, rather than patching it. On SWE-Bench Verified, a 30B OpenHands agent averages 23 rounds and 631K tokens per resolved issue, wi…
Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inference, and their execution strategies directly affect system accuracy, latency, and computational cost. Parallel execution provides a means to im…
arXiv:2608.04588v1 Announce Type: cross Abstract: Agentic systems have emerged as a promising paradigm for solving complex tasks by coordinating specialized LLM-based agents. However, most existing systems primarily optimize task success while giving limited consideration to exec…
arXiv cs.AI
TIER_1English(EN)·J. de Curt\`o, I. de Zarz\`a·
arXiv:2608.04265v1 Announce Type: cross Abstract: Evaluations of LLM planning agents largely ask whether a task succeeds or a declared plan is followed. In strategic cyber-physical systems, a stronger question is whether the planning architecture remains appropriate after autonom…
arXiv:2608.04192v1 Announce Type: cross Abstract: Closed source agent skills may encode proprietary instructions, scripts, constants, and data. Providers may offer their capabilities as services while keeping the underlying packages hidden. Prior work focuses on prompt injection …
arXiv:2608.04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) tool set, each engineered to probe one specific tool-…
arXiv:2608.04872v1 Announce Type: cross Abstract: Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses heterogeneous search failures into a scalar score and a single prompt. We …
arXiv:2608.04828v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed. Existing evaluations mostly judge the quality of a skill or its co…
Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses heterogeneous search failures into a scalar score and a single prompt. We propose A-SR, a self-evolving agentic framework th…
Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) tool set, each engineered to probe one specific tool-selection weakness. A six-type taxonomy (semanti…
Can cooperation among large language model (LLM) agents be evolutionarily stable against free-rider invasion? We study an indirect reciprocity donation game where LLM agents observe behavioral traces and donate on a continuous scale. Strategies, represented as natural language pr…
arXiv:2608.03206v1 Announce Type: cross Abstract: Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a l…
arXiv:2608.03239v1 Announce Type: new Abstract: Large language model-based multi-agent systems (LLM-MAS) are designed through roles, debate protocols, and aggregation rules. These choices create implicit social expectations: agents may be expected to trust, challenge, defer to, o…
arXiv:2608.02356v2 Announce Type: replace Abstract: Large language model agents increasingly solve complex tasks by composing reusable skills from a library. To address this, the key challenge is not merely to retrieve individually relevant skills, but to identify a complete and …
arXiv:2607.28048v2 Announce Type: replace Abstract: Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often treat skills as experience summaries, memory entries, or direct summaries of su…
arXiv:2608.03874v1 Announce Type: new Abstract: Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve…
arXiv:2608.02650v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external tools to complete complex real-world tasks. However, reliable tool-use planning remains challenging due to the limitations of implicit reasoning and the evolving nature…
We present a framework that gives LLM agents two mechanisms for searching large solution spaces autonomously. First, a leaderboard scored on held-out data acts as a reward signal that drives each agent to refine its solutions over repeated submissions, a loop that operates even w…
Agentic code generation has shown promise in automating and accelerating software development by utilizing Large Language Models (LLMs) to generate, test, and deploy code. For engineers and scientists, such systems have the potential to accelerate the development of applied and s…
arXiv cs.MA (Multiagent)
TIER_1English(EN)·I. de Zarzà·
Evaluations of LLM planning agents largely ask whether a task succeeds or a declared plan is followed. In strategic cyber-physical systems, a stronger question is whether the planning architecture remains appropriate after autonomous participants respond and physics constrains th…
Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge this gap, …
Large language model-based multi-agent systems (LLM-MAS) are designed through roles, debate protocols, and aggregation rules. These choices create implicit social expectations: agents may be expected to trust, challenge, defer to, or collaborate with peers. We study the effects o…
Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often treat skills as experience summaries, memory entries, or direct summaries of successful demonstrations. This creates a mismatch for…
Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge this gap, …
arXiv:2607.29678v1 Announce Type: new Abstract: LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, which resubmit a long transcript after each small tool result, and reuse is hard …
arXiv:2607.29626v1 Announce Type: new Abstract: As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmarks typically focus on static code generation, paper r…
arXiv:2607.28956v1 Announce Type: new Abstract: Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to p…
arXiv:2607.28942v1 Announce Type: new Abstract: Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retrieval-augmented generation, and scientific discovery. In these settings, agents must act based o…
Geographic Information System (GIS) professionals rely on multi-step spatial analysis workflows to support decision-making in urban planning, disaster response, and environmental monitoring. The process is tedious, time-consuming, and error-prone. While recent large language mode…
arXiv:2607.26724v1 Announce Type: new Abstract: Large language model (LLM) agents have been widely applied in automating data science tasks. However, existing methods typically rely on a limited set of provided datasets, and they face challenges in data-intensive scenarios that r…
arXiv cs.CL
TIER_1English(EN)·I. Kennedy, T. Kennedy·
arXiv:2607.28196v1 Announce Type: new Abstract: Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original, downstream accuracy (for example MMLU) inside a confidence interval, and data-fr…
arXiv cs.CL
TIER_1English(EN)·Sebastian Pohl, Harsh Mehta, Pranav Mambayil, Abdul Ghafoor, Franziska Lesigang, Yufang Hou, Christian Hilbe·
arXiv:2607.28347v1 Announce Type: new Abstract: LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been tested directly. We test whether six LLMs can simulate individual human belief u…
arXiv:2607.27275v1 Announce Type: new Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless. We test this claim for multi-turn, tool-calling agents, where it now matters most. On $\tau^2$-bench, across two open-weight model families in den…
arXiv:2607.27973v1 Announce Type: new Abstract: Recently, Reinforcement Learning (RL) has emerged as a crucial paradigm for the post-training of Large Language Model (LLM) agents. However, existing methods predominantly rely on sparse task rewards for policy optimization, failing…
arXiv:2607.26120v1 Announce Type: new Abstract: Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric information and strategic deception due to conflicting or hidden objectives. In t…
arXiv:2607.26773v1 Announce Type: new Abstract: Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-r…
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended hori…
Agentic commerce is moving from concept to deployed infrastructure: payment networks, retailers, and AI platforms are setting the stage for agents to transact on behalf of merchants and consumers. Yet whether the LLMs behind these agents can price competently in real markets, whe…
A single human must audit $N$ LLM agents under a budget of $B \ll N$ audits per round, guided by self-reported confidence that may be adversarially miscalibrated and by correlated errors. We model this as budgeted noisy inspection over a two-level Gaussian copula and locate the m…
Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original, downstream accuracy (for example MMLU) inside a confidence interval, and data-free output-fidelity signals that compare the comp…
LLM-based multi-agent systems have the potential to enable collective intelligence and scale toward solving highly complex tasks through coordinated ensembles of specialized agents. However, despite their theoretical potential, the architectural design space remains largely non-s…
arXiv cs.LG
TIER_1English(EN)·Yicheng Feng, Yan Zhang, Yan Cheng, Wei Qi·
arXiv:2607.27083v1 Announce Type: new Abstract: As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental tool-selection challenge: acquiring too few tools leaves the task under-informed, w…
arXiv:2605.16986v2 Announce Type: replace Abstract: Additional test-time compute can give LLM agents access to more past experience, yet expanding the context or adding rollouts does not necessarily yield greater agent capability. We call this challenge test-time compute-to-capab…
LLM-agent systems can solve complex tasks through dynamic self-organization and emergent cooperation. Auditing this process is essential because plausible intermediate or final outputs can conceal incomplete or unsupported work and poorly allocated responsibility, ultimately comp…
Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-relevant information. End-task performance alone …
arXiv:2607.25408v1 Announce Type: new Abstract: A growing body of 2026 work applies control theory to LLM agents: Lyapunov-certified stability for tool-mediated controllers (Prinos et al., "Stable Agentic Control", 2026), sample-complexity bounds for sparse policies over massive …
arXiv:2607.25415v1 Announce Type: new Abstract: Production LLM agents are increasingly assembled from a frozen model wrapped in a harness: a prompt template, a tool set, a memory/retrieval layer, a planning strategy, and a verification policy. Two 2026 systems, Meta-Harness (Lee …
arXiv cs.AI
TIER_1English(EN)·Mingwei Zheng, David OBrien, Siwei Cui, Pardis Pashakhanloo, Rajdeep Mukherjee, Myeongsoo Kim, Sachit Kuhar·
arXiv:2607.22711v1 Announce Type: cross Abstract: LLM coding agents operate by constructing trajectories that accumulate reasoning, tool calls, and results to enable multi-step decision-making. However, the conventional append-only trajectory architecture found in practice tightl…
arXiv:2607.22962v1 Announce Type: new Abstract: LLM agents that operate over many turns accumulate facts in an external memory store and reuse them as premises for downstream reasoning. A hallucinated fact written at one step therefore persists as a false premise for every subseq…
arXiv:2607.23394v1 Announce Type: new Abstract: Recent work shows that fine-tuning language models on even a small amount of poisoned data can install targeted misbehavior, and ostensibly benign data can transmit hidden preferences that generalize broadly. Standard defenses, such…
arXiv cs.AI
TIER_1English(EN)·Yihui Zhang (Beihang University), Tianyu Wo (Beihang University), Jinghao Wang (Beihang University), Xiaoyang Sun (University of Leeds), Menghao Zhang (Beihang University), Cangzhou Yuan (Beihang University), Li Li (Beihang University), Chunming Hu (Beih…·
arXiv:2607.23933v1 Announce Type: cross Abstract: As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces a fundamental tension between resource utilization and interactive tail latency…
arXiv:2607.22165v1 Announce Type: cross Abstract: LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation and production operations: live-environment fidelity (multi-turn read-write inter…
arXiv:2607.20827v1 Announce Type: new Abstract: LLM agents choose tools and arguments from context that mixes user requests, tool outputs, retrieved records, memory, and untrusted text. Evidence can be relevant without being authorized to determine a decision, so a correct action…
arXiv:2607.20531v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score the final answer or a fixed "ground-truth" list of tools, both of which are fragil…
arXiv:2607.19356v1 Announce Type: new Abstract: Tool-using LLM agents increasingly execute high-impact actions, making runtime safety monitoring essential. We present NEXUS (Neural EXecution Utility and Safety), a structured-plan safety monitor that applies a formal intervention …
Multi-agent LLM systems increasingly let one agent choose which other agents to work with, and agents are increasingly given personalities through personas. We test whether Big Five personality alone influences partner selection when capability is explicitly held constant. Host a…
arXiv:2508.14751v2 Announce Type: replace Abstract: We study goal-conditioned reinforcement learning in partially observable environments with sparse rewards and large, structured goal spaces. In such settings, complex goals often require composing simpler skills, but learning th…
arXiv cs.AI
TIER_1English(EN)·Philipp J. Schneider, Lin Tian, Marian-Andrei Rizoiu·
arXiv:2510.19299v2 Announce Type: replace Abstract: Can large language model (LLM) agents reproduce the complex social dynamics that characterize human online behavior -- shaped by homophily, reciprocity, and social validation -- and what memory and learning mechanisms enable suc…
arXiv:2607.18806v1 Announce Type: new Abstract: This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Language Model (LLM)-based agents. The agents are instantiated with distinct personas and collaboratively seek an itinerary that satisf…
arXiv cs.AI
TIER_1English(EN)·Artem Maryanskyy, Dmitry Budnikov, Alibek T. Kaliyev·
This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Language Model (LLM)-based agents. The agents are instantiated with distinct personas and collaboratively seek an itinerary that satisfies their constraints and preferences through na…
arXiv:2607.16215v1 Announce Type: new Abstract: Existing guardrail systems for large language model agents operate as binary classifiers that block unsafe content, leaving organizations to discard failing outputs and retry from scratch. We introduce RAIL Guard, a closed-loop resp…
arXiv:2603.22455v5 Announce Type: replace Abstract: Reusable skills let LLM agents package task-specific procedures, tool affordances, and execution guidance into modular building blocks. As skill ecosystems grow to tens of thousands of entries, exposing every skill at inference …
arXiv:2505.11765v5 Announce Type: replace-cross Abstract: Agents powered by advanced large language models (LLMs) have demonstrated impressive capabilities across diverse complex applications. Recently, Multi-Agent Systems (MAS), wherein multiple agents collaborate and communicat…
arXiv cs.AI
TIER_1English(EN)·Jing-Jing Li, Jianfeng He, Chao Shang, Devang Kulshreshtha, Xun Xian, Yi Zhang, Hang Su, Sandesh Swamy, Yanjun Qi·
arXiv:2509.25624v3 Announce Type: replace-cross Abstract: As LLMs advance into autonomous agents with tool-use capabilities, they introduce security challenges that extend beyond traditional content-based LLM safety concerns. This paper introduces Sequential Tool Attack Chaining …
arXiv cs.AI
TIER_1English(EN)·Roshan Klein-Seetharaman, Daniel Wang, Andrew Xu·
arXiv:2607.16961v1 Announce Type: new Abstract: Existing tool-use benchmarks report a single success rate for complex, multistep tasks. Inspired by ideas from cognitive science, we distinguish tool use from tool discovery and decompose the latter into curiosity (the model's abili…
arXiv:2607.16617v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. We call this…
Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual substrate engineering, each new domain demands a bespoke pipelin…
Verify-repair loops are a standard means for large language model (LLM) agents to correct faulty plans in code generation, mathematical reasoning, and tool use. When both the verifier and the repairer are noisy, repair can damage already-correct plans, and reported acceptance kee…
arXiv:2607.15593v1 Announce Type: cross Abstract: LLM agents increasingly rely on tool calling to act on external systems, and the Model Context Protocol (MCP) has quickly become its de facto interface. Operating MCP at cloud scale, however, becomes difficult. On the tool provide…
arXiv cs.CL
TIER_1English(EN)·Yanze Wang, Pengfei Yao, Tianyi Sun, Chuanrui Hu, Yan Xiao, Yunyun Han, Jun Sun, Yafeng Deng·
arXiv:2607.15557v1 Announce Type: new Abstract: Agent skills, SKILL.md files that package reusable procedural knowledge for an LLM agent, are a popular mechanism for extending agent capabilities. Public repositories now host them in large and growing numbers, yet these artifacts …
Replacing individual LLM agents with low-parameter surrogates fitted from cheap queries enables scalable society simulations, with validity predicted by an interaction-order and memory taxonomy.
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. We call this disconnect the NL2Pipeline gap. To bridge it, we …
arXiv:2607.14197v1 Announce Type: new Abstract: Artificial Intelligence (AI) answer engines now field a growing share of the questions that analysts, scholars, and the public ask about issues of peace and conflict. Large Language Models (LLMs) are known to hallucinate under certa…
arXiv:2607.11079v1 Announce Type: new Abstract: Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, s…
Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, statistical inference, mechanistic explanation, e…
arXiv:2609.05224v1 Announce Type: new Abstract: Recent progress in multimodal large language models (MLLMs) has fueled significant enthusiasm in their potential to act as autonomous agents for real-world tasks. However, scenarios requiring agents to fulfill users' complex, struct…
arXiv:2608.24087v1 Announce Type: cross Abstract: Current LLM agent systems decide delegation before reasoning begins (a router picks a model) or after a response is complete (a verifier scores it and may retry). We study a third regime: an agent that recognises, during its own r…
arXiv:2504.18587v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrated by systems such as OpenAI's O1~\cite{o1} and DeepSeek-R1~\cite{r1}. However, wide…
arXiv:2607.26865v1 Announce Type: new Abstract: LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. Yet, when deployed at the edge, they must tight…
Hacker News — AI stories ≥50 points
TIER_1English(EN)·rellem·
<p>Retrieval, agents, and workflows now ship as visual and plain-English tools. This roundup covers 10 open-source no-code and low-code platforms for building LLM apps, RAG systems, and AI agents, each with its verified license, repository, and best-fit use case.</p> <p>The post …
<h1> Tool-Call Injection in LLM Agents: Why Your MCP Server Is the New Attack Surface </h1> <p>An LLM agent that can read email, browse the web and run shell commands is useful precisely because it acts on untrusted input. That combination is also why the Model Context Protocol (…
<p>Your AI governance problem probably isn't a policy problem; it’s a routing one. </p> <p>Let’s paint a common scenario: Your team uses OpenAI, another adds Anthropic, then the product team connects another model. Then someone builds an agent with MCP tools. A few months later, …
<h1> Solving LLM Parameter Hallucinations in AI Sales Agents with Native MCP Tools </h1> <p>The most efficient way to eliminate LLM parameter hallucinations when retrieving B2B firmographics is by leveraging a native Model Context Protocol (MCP) server with strict Zod-enforced sc…
<h1> Agent Gateway in 60 Seconds: Governed LLM Traffic with TrustGate </h1> <p>Most teams start with a direct OpenAI (or Anthropic) SDK call. That works until you have three apps, two providers, and a security review asking who can call which model, at what rate, with what audit …
<p><em>Written by </em><a href="https://www.linkedin.com/in/matteo-rossi-280391/"><em>Matteo Rossi.</em></a></p><p>The monthly LLM bill jumped, and nobody on the team can say which agent, which user, or which workflow caused it. The provider dashboard breaks usage down by organiz…
Medium — fine-tuning tag
TIER_1English(EN)·Mikhail Borodastov·
<p>The bleeding edge of AI automation isn't just about making Large Language Models (LLMs) smarter; it's about giving them hands and eyes. When building vision-driven agentic architectures, we cross a massive chasm: bridging the high-level semantic reasoning of an LLM with the lo…
<h1> Eliminating LLM Hallucinations in B2B Sales Agents with the Lead Enrichment MCP API </h1> <p>To stop LLMs from hallucinating company data or fabricating contact details, developers must shift from loose prompt-based retrieval to a Model Context Protocol (MCP) architecture th…
Medium — Claude tag
TIER_1English(EN)·Ashishmohanka·
<h1> Eliminating LLM Parameter Hallucinations in Sales Agents with Native MCP B2B Enrichment </h1> <p>The most efficient way to stop LLMs from hallucinating firmographic data or misinterpreting complex API schemas is to deploy a Model Context Protocol (MCP) native B2B lead enrich…
Medium — MLOps tag
TIER_1English(EN)·Tedi Ikonomi·
<h1> PromptOT MCP: Manage and version LLM prompts from your AI tools </h1> <p>Prompts often start as simple strings in code.</p> <p>Then the product grows.</p> <p>You add a better system prompt. Then a guardrail. Then a different version for production. Then a customer-specific v…
Medium — MLOps tag
TIER_1English(EN)·Neelopphersyed·
<div class="medium-feed-item"><p class="medium-feed-snippet">The Model Context Protocol hit 97 million monthly SDK downloads by December 2025. That number alone tells you something important is…</p><p class="medium-feed-link"><a href="https://medium.com/@euroamericanmalta/…
<h4>There is no shortage of articles about building AI agents. What remains much rarer is a practical discussion of how to run them safely in production.</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*tnltjwGYfOIZX6KgpLNUEg.png" /></figure><p>This article…
<h2> How AI Fallback Strategies Keep Uncertain Systems Safe and Useful </h2> <p>The real problem with AI systems is not that they fail. It is that they fail with confidence. When uncertainty rises, the system should catch it early, switch to a controlled fallback, and keep the us…
<p>[ EXECUTIVE TEARDOWN // TL;DR ]</p> <ul> <li> Emit one span per agent step so a request becomes a readable waterfall — total at the root, a labelled child per stage.</li> <li> Use the OpenTelemetry GenAI semantic conventions (gen_ai.*) so any backend renders traces the same wa…
dev.to — LLM tag
TIER_1English(EN)·Manoranjan Rajguru·
<h1> Qwen 3.8 27B: The Frontier LLM That Fits on Your Laptop — Architecture, Reasoning Control & Agentic Integration </h1> <p><em>Published August 18, 2026 · 18 min read</em></p> <h2> Table of Contents </h2> <ol> <li>The Benchmark Moment Nobody Saw Coming</li> <li>How Good Is…
<p>A practical comparison of the four categories of human-in-the-loop tooling for LLM agents — framework-native interrupts, workflow-engine nodes, general approval software, and dedicated gates — so you pick by category, not by name.</p> <h2> The four categories </h2> <p>"HITL to…
<p>I’ve been working on a side project called F*ckCTF — an autonomous agent designed to solve black-box Capture The Flag challenges. It runs inside an isolated Kali Linux container and uses a Multi-Agent architecture to interact with terminals, run scripts, and hunt for flags.</p…
<!-- SC_OFF --><div class="md"><p>LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a successful mutation may leave persistent effects that cannot be safely reversed …
<p>Most production LLM control sits between two extremes: prompt engineering (brittle, context-dependent) and fine-tuning (expensive, slow iteration). Mentat, a YC F24 launch, introduces a third path: runtime intervention that modifies token probabilities mid-generation without r…
<p>Multi-stage LLM agent workflows have a silent failure mode. A hard constraint enters the pipeline at stage one. By stage three, it has become a suggestion. The executor reads it, acknowledges it, and proceeds anyway.</p> <p>The problem is not hallucination or context loss. The…
LLM agents run controlled experiments on pharma simulations A multi-agent framework from the University of Stuttgart lets LLM agents design, run, and interpret experiments on pharmaceutical simulation models. https://www. notatechguy.com/llm-agents-run -controlled-experiments-on-…
<blockquote> <p><strong>The Pain</strong>: Your agent forgets to query the knowledge base. The same knowledge base, the same model — different orchestration produces wildly different outputs. Change one prompt line and everything breaks. Agents still "work by feel", and every fix…
<p>Last month, someone pasted a benchmark table into our team chat. Model B beat model A by 12 points on "general agent tasks," so we swapped models the same afternoon.</p> <p>Two weeks later, the triage agent was mislabeling about a third of the issues it touched.</p> <p>The tab…
<p>When you give a language model the ability to call tools — run code, query databases, browse the web — you've created an autonomous execution surface. Most tutorials skip the part where that surface gets exploited.</p> <p>This post covers practical steps for sandboxing LLM too…
<blockquote> <p><em>This article was originally published on <a href="https://www.buildzn.com/blog/my-2x2-llm-cost-performance-strategy-for-ai-agents" rel="noopener noreferrer">BuildZn</a>.</em></p> </blockquote> <p>Everyone's chasing the biggest LLMs, throwing cash at Claude or …
<p>Pydantic AI is the official agent framework from the Pydantic team, built around typed, validated LLM output. After 45 days of using it for saas.pet's content QA agent and data extraction scripts, here is the real story on structured output, tool calling, and why it beats Lang…
TrueForge: открытая обвязка, которая превращает LLM в полноценного агента Собрать демо-агента сегодня несложно. Подключаешь модель, добавляешь пару инструментов — и она уже читает файлы, вызывает API и бодро обещает выполнить любую задачу. Сложности начинаются, когда такого агент…
Beyond # LLMs : Creating Real-World # AI Agents with Lang Chain Deep Agents Discover how LangChain DeepAgents transform LLMs into production-ready AI systems with memory, skills, sub-agents, context management, and human oversight. https:// hackernoon.com/beyond-llms-cre ating-re…
SkillOpt (Microsoft) : un optimizer de skills en langage naturel pour agents LLM. Le skill s'améliore via des rollouts scorés, sans toucher aux poids du modèle. Le fichier best_skill.md est portable d'un modèle à l'autre. Open source, MIT. ⬇️ https:// github.com/microsoft/SkillOp…
<h2> What Happened </h2> <p>Zero‑Mem lets LLM agents read and write external memory without generating or consuming any tokens. Traditional agents fetch context through token‑based prompts, adding latency and cost. Zero‑Mem replaces that with a lightweight, token‑free interface t…
<h1> Running LLM Agents at the Edge: A Practical Guide with NeoMind + Ollama </h1> <p>Everyone's building AI agents right now. Most of them live in the cloud — you send a request to OpenAI or Anthropic, get a response back, and hope the latency and cost stay reasonable. But what …
<p>Large language models have changed how developers build applications.</p> <p>However, even the most advanced LLMs have one fundamental limitation:</p> <p>They do not have access to real-time information.</p> <p>A model may understand programming, reasoning, and language extrem…
<h2> Introduction </h2> <p>Large Language Models (LLMs) are powerful, but most enterprises rely on pre‑packaged APIs. I wanted to go deeper: train my own LLM model and build an agent layer on top of it that could interact with real systems securely.</p> <p>This post walks through…
LLM agent: code-only verification flips goal abandonment 100% to 0% New arXiv preprint: a deterministic executive owns all agent belief, the LLM only files proposals, and zero ARC-AGI-3 completions are honestly disclosed. https://www. notatechguy.com/llm-agent-code -only-verifica…
<p>An industry-level LLM agent is not simply an API call that returns a response. It needs to be resilient to transient failures, malformed outputs, and schema violations.<br /> To improve the stability of my agent system, I introduced two decorators around my LLM calls. They han…
dev.to — LLM tag
TIER_1English(EN)·Lorena Dávila Ermus·
<p>If you want to run AI models on your own machine and learn the basic concepts with me to do it effectively, then this is the right article :).</p> <p>This is part one of the series. In the next one we build local AI workflows with n8n and Ollama. This article is the vocabulary…
📄 ‘MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations’ hit 80 upvotes on Hugging Face. Tests LLMs on sustained task chains in shopping. https:// huggingface.co/papers/2607.289 56 # AI # MachineLearning # Research
<p>In this post, I'll explain how prompt chaining, tools/skills, and iteration actually make agnets to produce results which are not usually possible when we use simple LLMs.</p> <h2> LLM chaining </h2> <p>You know how sometimes you write one massive prompt like:</p> <blockquote>…
<p>Over the last 18 months our ML team has been doing some very interesting things: building AI agents on top of PostgreSQL, while the infrastructure evolves, the industry matures, and quality expectations keep rising. We started with a single A100 in a managed cloud and fairly m…
<p><strong>Short answer</strong></p> <p><strong>You can't unit-test an LLM to correctness, because the same input can take a different path on the next run.</strong> Evals are the test suite for probabilistic systems: a scored, repeatable check of whether the system reached an ac…
dev.to — LLM tag
TIER_1English(EN)·Hiroshi Toyama·
<p>llms.txt ist ein maschinenlesbarer Standard, der KI-Systemen strukturierte Informationen über Ihre Website bereitstellt. Aufbau, Best Practices und praktische Implementierung für bessere KI-Sichtbarkeit bei ChatGPT, Claude, Gemini und Perplexity.</p> <h3> Zusammenfassung </h3>…
dev.to — LLM tag
TIER_1English(EN)·Jules Robineau·
<blockquote> <p><strong>TL;DR</strong>: you only truly understand a system once you rebuild it. I recoded TCP at school, then the DNS protocol, then Modbus, each time to understand it from the inside. A colleague just went through this with LLMs. He wrote a small agent in Go, and…
dev.to — LLM tag
TIER_1English(EN)·Yusuf Al-Rashidi·
Just added local LLM support to the AI agent boilerplate, you can now run it on top of Ollama instead of relying on cloud APIs 👇 https:// github.com/christopheduc-me/ai -agent-boilerplate # buildinpublic # ai # dev # tech
<p>AI features ship fast and then the bill arrives. The good news: most LLM spend is avoidable waste - the same prompt paid for a thousand times, a frontier model doing work a cheap one could handle, tokens generated that nobody reads. Here are six levers that cut real money, ord…
dev.to — LLM tag
TIER_1English(EN)·Apache SeaTunnel·
<p>OpenAI API'ına token başına para ödemek yerine, açık kaynak dil modellerini (Llama 3, DeepSeek, Mistral, Qwen) kendi sunucunda çalıştırabilirsin. Verin dışarı çıkmaz, sabit maliyet, sınırsız istek. Bu yazıda Ollama ile pratik kurulumu ve gereken donanımı anlatıyorum.</p> <h2> …
dev.to — LLM tag
TIER_1English(EN)·Learn AI Resource·
<p>So you want an AI coding assistant but you're tired of getting dinged for API calls every time you ask for help debugging a regex? Yeah, I get it.</p> <p>Here's the thing: you don't actually need to pay OpenAI or Anthropic to get decent AI pair programming. You can run a solid…
<blockquote> <p>TL;DR: Moderation catches harm and many injection attempts. It does not enforce domain or operational policy. A policy reasoning layer (LLM-as-a-judge) closes that gap, especially in multi-turn conversations.</p> </blockquote> <p><strong>Abstract</strong><br /> Mo…
<h2> Local LLMs, Open Agents & Self-Hosted Deployment Platforms Trending </h2> <h3> Today's Highlights </h3> <p>Today's top stories highlight the growing trend of local and self-hosted AI deployments, featuring an architectural guide for secure "Local Sovereign LLMs" in enter…
<p><em>Organizations deploying AI applications face challenges in accurately tracking LLM usage and spend across different teams and projects. <a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer">Bifrost</a> offers a comprehensive AI gateway solution with virtual k…
dev.to — LLM tag
TIER_1English(EN)·Stéphane Derosiaux·
<p>Ever handed an LLM a full web page and watched the amount of tokens being used?</p> <p>A single product listing is 20-30K tokens of </p> soup before the model finds what it needs: wrapper divs, css class, SVG, JSON blobs etc. The agent needs maybe 300 tokens of that (the items…
<p>I built a system where an LLM talks to a customer about a silicone casting mold, and a<br /> deterministic geometry kernel — OpenCASCADE, three decades of production C++ — does the<br /> actual mass-solving, boolean surgery, and part-splitting. The LLM never touches the kernel…
<h2> LLM Inference & RAG Optimization, Open-Source Voice AI for Local Deployments </h2> <h3> Today's Highlights </h3> <p>This week's highlights feature a new framework for LLM inference and fine-tune optimizations, including KV-cache improvements, alongside an open-source voi…
dev.to — LLM tag
TIER_1English(EN)·Kuldeep Singh Arya·
<p>Executive Summary<br /> As generative AI transitions from experimental prototypes to high-scale production systems, the primary bottleneck for engineering teams has shifted from model capability to unit economics. The pricing structure of modern Large Language Model (LLM) APIs…
<p><strong>TL;DR:</strong> Sir Shortoken is a system prompt that constrains frontier models to operate within information budgets (Quick/Balanced/Deep), never silently escalate capabilities, and prove execution. Tested across Claude, GPT, Gemini. 40-60% token reduction on technic…
dev.to — LLM tag
TIER_1English(EN)·Praveen Maurya·
<blockquote> <p>I didn't build SafeDevTools by asking AI to "build me a website." I built it by treating a local LLM like a junior engineer who never gets tired of writing boilerplate.</p> </blockquote> <p>A few weeks ago, I challenged myself with a simple experiment:<br /> <stro…
dev.to — LLM tag
TIER_1Deutsch(DE)·Uhltak Therestismysecret·
<h1> Lokale LLMs mit Ollama – Modelle selbst hosten und per API anbinden </h1> <p><strong>Hook:</strong> Stell dir vor, du könntest ChatGPT für deine Firma betreiben, ohne einen teuren Cloud‑Vertrag oder ein Datenleck‑Szenario. Du hast die volle Kontrolle, die Kosten liegen bei d…
<p>Если ты открыл эту статью с вопросом «где посмотреть, какая модель реально влезет в мой Mac и не будет тормозить», то короткий ответ такой: 2 июля 2026 года Exo Labs на конференции AI Engineer World's Fair анонсировала сервис local.ai, который отслеживает, какая модель лучше в…
<p>If you've heard about LangChain and assumed it was a Python thing, that's fair. It mostly was.</p> <p>LangChain became popular because building with an LLM turns out to involve a lot of repetitive plumbing. You need to manage conversation history, split documents into chunks, …
<p>From data preparation and tokenizer selection to pretraining, LoRA, RLHF, evaluation, and production monitoring, this guide covers the major stages involved in training an AI model.</p> <p>DEHA Research · July 14, 2026 · 18 min read</p> <p>Training an artificial intelligence m…
<h2> Browser LLM Agents, Rust Engine for Apple Silicon, & Local AI Code Interpreter </h2> <h3> Today's Highlights </h3> <p>This week, we spotlight tools bringing LLM inference directly to your devices. Dive into browser-based agents, a Rust-native engine for Apple Silicon, an…
dev.to — LLM tag
TIER_1English(EN)·Innocent Oyebode·
<h2> The Problem </h2> <p>Most Nigerian small businesses have no web presence at all. When they do get a website, it is usually a stale brochure-ware page that took a freelancer three weeks to deliver and costs ₦150,000 they could not really afford. The freelancer is long gone; t…
<p>A slow AI feature rarely fails all at once. It starts with a longer prompt, then a bigger retrieval result, then one more tool call, then a retry path nobody measured. The demo still works, but users feel the delay before your dashboard explains it.</p> <p>That is why small AI…
<h2> Self-Hosted AI Companion & Open-Source Model API Insights </h2> <h3> Today's Highlights </h3> <p>This week's highlights feature a trending self-hosted AI companion, empowering users with personal, locally-run AI experiences. We also explore a bootcamp grad's practical in…
dev.to — LLM tag
TIER_1English(EN)·Delafosse Olivier·
<blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/from-demos-to-durable-systems-ai-engineering-techniques-that-make-llms-truly-product-ready?utm_source=devto&utm_medium=syndication&utm_campaign=kb-incidents" rel="noopener noreferrer">…
<h2> Self-Hosted LLM Apps, Offline AI Systems, and Local Automation Foundations </h2> <h3> Today's Highlights </h3> <p>This week, we spotlight practical approaches to self-hosting AI, from extensive curated lists of runnable LLM applications to ambitious projects building fully o…
<h1> Published: true </h1> <h1> Description: A deep dive into the final version of EdgeSync-LLM—bringing fast, secure, synchronized Large Language Models straight to edge hardware. </h1> <h1> Tags: ai, open source, architecture, edgecomputing, webdev </h1> <p>The cloud dependency…
Exploring open-source AI models like Llama, Mistral, and Phi is a must for anyone in the tech world! These models are changing the landscape of AI by promoting collaboration and innovation. Dive into the world of deep learning! 🤖 # AI # AITürkiye # DeepLearning # Teknoloji # Mach…
dev.to — LLM tag
TIER_1English(EN)·Delafosse Olivier·
<blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/gpt-5-6-in-the-wild-how-openai-s-new-model-and-custom-silicon-will-reshape-production-llm-systems?utm_source=devto&utm_medium=syndication&utm_campaign=kb-incidents" rel="noopener noref…
dev.to — LLM tag
TIER_1English(EN)·Delafosse Olivier·
<blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/gpt-5-6-jalapeno-and-the-next-generation-of-openai-optimized-llm-infrastructure?utm_source=devto&utm_medium=syndication&utm_campaign=kb-incidents" rel="noopener noreferrer">CoreProse K…
<p>I learned the hard way that a working LLM pipeline and a production LLM pipeline are two different things.</p> <p>When I first built the scoring system for a job board platform, I thought: throw GPT-4 at each listing, ask it to rate relevance, done. It worked for 100 listings.…
dev.to — LLM tag
TIER_1English(EN)·Delafosse Olivier·
<blockquote> <p>Originally published on <a href="https://www.coreprose.com/kb-incidents/inside-gpt-5-6-how-openai-s-new-flagship-model-and-custom-silicon-will-reshape-llm-operations?utm_source=devto&utm_medium=syndication&utm_campaign=kb-incidents" rel="noopener noreferre…
<p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frp4yqktt436b7wu1fox2.png"><img alt="9 Ways an AI Gat…
<p>I once built an AI pipeline that was shut down after a single month. The LLM costs were unsustainable, and worse, the outputs were unreliable enough that we couldn't trust them in production. That failure taught me something I still use today: evaluation isn't a phase you add …
🤖 【Hugging Face Papers】PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used t... # AI # TechNews #... 🔗 https:// huggingf…
Apple brought ceramic back for the Series 12 The new ceramic colors: pearl white and night blue. | Image: The Verge via Apple The Apple Watch Series 12 is launching with a ceramic option, making it the company's first wearable since 2019's Series 5 that can be con… https://www. t…
Source: pageindex.ai/blog/ocr Vector RAG: Why It’s Winning in Production In a... # ai # llm # machinelearning # rag # software # coding # development # engineering # inclusive # community Vector RAG: Why It’s Winning in Production
Deterministic control over LLM behavior through runtime steering. Architecture, latency trade-offs, and compliance implications for financial agents. # agents # ai # api # llm # software # coding # development # engineering # inclusive # community Runtime Intervention for LLMs: H…