English(EN)Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise
AI 研究探索自主系统、递归安全和人机协作 · 跟踪 8 个来源
作者PulseAugur 编辑部·[50 个来源]·
2026 年 9 月提交给 arXiv 的多篇研究论文探讨了 AI 的进展,重点关注自主系统、递归自我改进和人机协作。一篇论文详细介绍了一个结合连接主义和符号 AI 的自主系统框架,强调了代理架构和可信度。另一篇论文引入了“进化安全”来解决递归自我改进 AI 中的风险,提出了风险发现和评估的分类法。通过悖论和模式研究了人机协作,确定了四种主要类型:指令、委托、协助和共创。此外,还介绍了关于协调 AI 编码代理的“ParallelPilot”、管理企业 AI 代理成本的“Harness Tokenomics”、构建可执行 AI 实验室的“LabFactory”的研究,以及一项关于 AI 编写的代码库可靠性的研究。另一篇论文介绍了通过递归 AI 开发的前沿工业编码模型“iCoder-27B”,该模型在多项基准测试中表现优于现有模型。
AI
arXiv:2609.30291v1 Announce Type: new Abstract: The purpose of this article is to highlight the central role of autonomous systems as the ultimate stage in the development of AI, to explain the underlying technical challenges that require a combination of connectionist AI and sym…
arXiv:2609.31186v1 Announce Type: new Abstract: Artificial intelligence is advancing rapidly, with increasingly capable systems taking larger roles in reasoning, decision-making, scientific discovery, and autonomous development. As AI begins to participate in its own improvement,…
Evidence shows that humans and AI systems perform better together, by collaborating, than alone. This paper examines two key design dimensions of human-AI collaboration (autonomy and initiative) and explores the collaboration patterns that they generate. Documenting these pattern…
As coding assistants become increasingly autonomous, developers run multiple sessions in parallel, shifting the challenge from code generation alone to coordinating and monitoring concurrent agent work. Through a formative study (N=14), we identified PILOT: five supervisory pract…
arXiv:2609.29626v1 Announce Type: new Abstract: Recursive AI, the prospect of AI taking an increasingly complete role in building and improving AI, is a crown jewel of AI for AI. Although recursive self-development has become practical for small models, bounded tasks, and fixed t…
arXiv cs.AI
TIER_1English(EN)·Arian Abbasi, Alan Aqrawi, Ted Kwartler·
arXiv:2609.28919v1 Announce Type: new Abstract: Harnesses, the products that run AI coding agents, are multiplying, and enterprises are rolling them out to their employees: what started as pilots with a few hundred seats is scaling to tens of thousands. Most enterprises do not bu…
arXiv:2609.29744v1 Announce Type: cross Abstract: We present: (i) a new dataset consisting of the full development history of a 21,000-line Python tool built entirely by Claude AI, with no human-authored code or tests, (ii) two code-provenance tracing tools, (iii) three taxonomie…
arXiv cs.LG
TIER_1English(EN)·Jinge Wu, Hongjian Zhou, Mingde Zeng, Jiayuan Zhu, Junde Wu, Jiazhen Pan, Lei Clifton, Andrew Liu, David A. Clifton·
arXiv:2609.28697v1 Announce Type: new Abstract: Scientific tasks specify a desired capability, but realizing it often requires building a computational system tailored to the task---acquiring data, designing representations, training models, implementing tools, and deciding how t…
Manually extracting data from hundreds of vendor contracts doesn't scale, and RAG chat tools fall short on portfolio-wide questions. This post shares a contract intelligence platform on AWS that uses AI agents to extract and verify contract fields, then answers aggregate and sing…
AI Supremacy (Michael Spencer)
TIER_1English(EN)·Michael Spencer·
Weco let an AI coding agent rewrite the harness around another agent for eight days: its code, prompts and tools, while the underlying language model stayed fixed. Tim Scarfe asks Weco co-founder Zhengyao Jiang what the reported gains over two years of human engineering actually …
AI Supremacy (Michael Spencer)
TIER_1English(EN)·Michael Spencer·
As AI pushes rack densities into unprecedented territory, utilities, hyperscalers, and developers are rethinking how to deliver reliable, affordable power at massive scale.
<p>Google Cloud AI Research has open-sourced RRSI, a framework that lets LLM agents rewrite their own prompts, tools and memory while model weights stay frozen. It adds a leakage critic, a noise floor, a cost rule and pruning so gains carry over to new tasks. With Claude Opus 4.8…
dev.to — Claude Code tag
TIER_1English(EN)·Thomas Tartrau·
<p>An AI coding agent reads files written by other people, runs commands with your privileges, and sees your secrets. AI coding agent security is therefore not a theoretical topic: between June 2025 and July 2026, the <a href="https://github.com/advisories?query=%40anthropic-ai%2…
<p>TypeSafe AI's Jev skips text generation and returns typed decisions with calibrated probabilities. Input costs $0.042 per million tokens and output is free. We verified 20 agentic use cases, from model routing and tool-call gating to reranking and injection screening, and comp…
<p>We read the contracts behind GitHub Copilot, AWS Kiro, Cursor, Devin and Windsurf. Copilot and Kiro offer uncapped indemnity on generated code. Cognition's standard terms exclude outputs entirely. Here is how 500 seats compare on legal exposure, prompt storage, audit logs and …
dev.to — Claude Code tag
TIER_1English(EN)·yureki_lab·
<h2> TL;DR </h2> <p>My fully autonomous implementation system runs for days at a time, and every few hours the agent's context window fills up and gets wiped. Early on, each reset meant the agent forgot what it was working on, redid finished tasks, or quietly reversed decisions i…
<p>Imagine having 230 AI specialists at your fingertips – an editorial expert for marketing texts, a security auditor for code reviews, a UX designer for interface feedback, a Reddit community ninja, and a lease contract lawyer. All waiting to be activated by you in Claude Code. …
dev.to — Claude Code tag
TIER_1English(EN)·TerminalBlog·
<blockquote> <p><em>Originally published at <a href="https://terminalblog.com/blog/ai-coding-agents-unfiltered-truth-developers/" rel="noopener noreferrer">terminalblog.com</a>.</em></p> </blockquote> <p>The marketing says AI coding agents make you 10x. The developers who actuall…
Medium — Claude tag
TIER_1English(EN)·Symprio Blogs·
<p>When you integrate AI into a SaaS product, the initial metric everyone looks at is usage frequency. Is the button being clicked? Is the LLM call happening? These are vanity metrics that fail to capture whether your AI features are actually driving user retention or deepening p…
Medium — Claude tag
TIER_1English(EN)·Cristian Marcu·
<h4>When a payment gateway drops a TCP packet, an LLM’s autonomous retry loop turns a routine 504 timeout into an automated financial disaster. Here is the distributed systems architecture required to prevent it.</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/10…
Medium — AI coding tag
TIER_1English(EN)·Sheshikiran Gone·
<div class="medium-feed-item"><p class="medium-feed-snippet">Artificial intelligence is rapidly changing how software is built. What started as “vibe coding” — describing an idea in natural language…</p><p class="medium-feed-link"><a href="https://medi…
Medium — AI coding tag
TIER_1English(EN)·Yash Naik·
<p>A monitoring alarm fired within 12 minutes after an OpenAI research agent reached an external chatbot through DNS, yet the run continued for roughly 2.5 hours before a human stopped it, according to <a href="https://www.techi.com/openai-pauses-ai-model-training-agent-dns-sandb…
Towards AI
TIER_1English(EN)·Louis-François Bouchard·
<div class="medium-feed-item"><p class="medium-feed-snippet">From Code Typist to System Supervisor: Navigating the AI Era as a Senior Developer</p><p class="medium-feed-link"><a href="https://medium.com/@joey.zhou81525/surviving-and-thriving-in-the-ai-coding-era-bca96c1d4884?sour…
dev.to — LLM tag
TIER_1English(EN)·AI Frontier Post·
<p><em>Originally published at <a href="https://aifrontierpost.com/articles/openai-agents-api-hands-on-tutorial" rel="noopener noreferrer">AI Frontier Post</a></em></p> <p>For the past year, shipping an AI agent meant writing the same loop as everyone else: call the model, parse …
<p>Each lab's own courses, cookbooks and docs — Claude, ChatGPT, Gemini, Llama, Mistral — and how to run open models yourself.</p> <p>I keep a directory of free ways to learn at <a href="https://brianpfeil.com/learn/?utm_source=devto&utm_medium=crosspost&utm_campaign=lear…
dev.to — LLM tag
TIER_1English(EN)·Tanuj kumar Tanuku·
<p>In modern B2B sales cycles, account executives juggle multiple conversations with a single prospect over weeks or months. From initial discovery calls to deep-dive technical reviews, every interaction builds a complex web of specific requirements, budget constraints, security …
Autonomous AI is changing how software systems are designed. Future intelligent systems will need more than models — they’ll require persistent memory, distributed infrastructure, autonomous coordination, and adaptive execution. This is one of the areas we’re exploring through An…