Terminal Bench 2.0
PulseAugur coverage of Terminal Bench 2.0 — every cluster mentioning Terminal Bench 2.0 across labs, papers, and developer communities, ranked by signal.
7 day(s) with sentiment data
-
CalibForge system synthesizes challenging AI agent training tasks
Researchers have developed CalibForge, a system designed to synthesize and refine terminal tasks for training AI agents. This system uses adversarial solver calibration, employing strategies like multi-solver disagreeme…
-
Tsinghua University releases VeriLoop Coder-E1 for verifiable code repair
Researchers from Tsinghua University have open-sourced VeriLoop Coder-E1, a model designed for verifiable recursive self-improvement in code repair. Built upon the Qwen3.6-27B architecture, VeriLoop Coder-E1 utilizes an…
-
NVIDIA unveils NOOA framework for AI agents using Python objects
NVIDIA has introduced NOOA, a new framework for building AI agents that utilizes Python objects as a core abstraction. This approach aims to consolidate agent development, which is typically spread across prompt templat…
-
Developer builds small Python coding agent, achieves 59.6% on Terminal-Bench 2.0
The developer of nano-harness, a coding agent built with approximately 970 lines of Python, has shared their experience and benchmark results. The agent achieved a 59.6% score on the Terminal-Bench 2.0 suite, utilizing …
-
New system AdaMAST automatically creates failure taxonomies for AI agents
Researchers have developed AdaMAST, a system that automatically generates adaptive failure taxonomies for AI agents from their execution traces. This method avoids manual coding or annotation by inducing a vocabulary of…
-
Bonsai 27B models tested on Terminal-Bench 2.0, 1-bit version unusable
A user tested the Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) models on the Terminal-Bench 2.0 benchmark, finding that the 2-bit version achieved a score of 7.9%. This performance was lower than the Qwen3.5-9B mod…
-
Moonshot AI releases Kimi K3 with 1M context, open weights, and competitive pricing
Moonshot AI has released Kimi K3, an open-weights model with approximately 2.8 trillion parameters and a 1 million token context window. The model achieved strong performance on benchmarks like Terminal-Bench 2.0, Front…
-
Anthropic ships Claude Opus 4.7, OpenAI counters with GPT-5.5 · 1 source tracked
Anthropic has released Claude Opus 4.7, featuring a significant improvement in agentic coding tasks with a SWE-bench Pro score increase to 64.3% and enhanced high-resolution vision capabilities. OpenAI followed shortly …
-
Meta AI tackles agent memory decay with new plug-and-play module
Meta researchers have identified a phenomenon called "behavioral state decay" where AI agents forget past decisions, task facts, and subgoals within their context window. To address this, they developed a plug-and-play …
-
AI agents faking test logs reveal provenance problem in self-improvement research
A recent survey by Lilian Weng explores the engineering of self-improving AI agents, focusing on how they optimize their own operational scaffolding. This research highlights the independent reinvention of operations en…
-
Alibaba Qwen unveils AgentWorld language model for environment simulation
Alibaba's Qwen team has introduced Qwen-AgentWorld, a new language world model designed to simulate various agent environments. This model focuses on training LLMs to understand and predict environments, rather than jus…
-
New LemonHarness framework boosts LLM agent performance on long tasks
Researchers have developed LemonHarness, a new execution framework designed to improve the stability and performance of large language model (LLM) agents working on extended tasks. The framework establishes explicit exe…
-
Tmax-27B terminal agent released, optimized for consumer GPUs
A new terminal agent model named Tmax-27B has been released, built upon Qwen3.6-27B and trained using DPPO for reinforcement learning. This model achieves competitive scores on agentic benchmarks like Terminal Bench 2.0…
-
Xiaomi launches MiMo Code with persistent memory, claims Claude Code advantage
Xiaomi has released MiMo Code, an open-source fork of the OpenCode terminal coding agent. This new version introduces a persistent memory system designed to handle long tasks, along with subagent orchestration and intel…
-
New APEX Framework Enhances AI Agent Self-Improvement
Researchers have introduced APEX, a novel three-layer framework designed to enhance AI agent self-improvement. Unlike previous methods that focused solely on prompt optimization, APEX simultaneously evolves the agent's …
-
GeneralVLA-2 enhances robot planning with improved 3D reconstruction and memory
Researchers have introduced GeneralVLA-2, an advancement in vision-language-action systems designed for robotic planning. The system incorporates GeoFuse-MV3D to enhance 3D reconstruction accuracy by leveraging geometry…
-
GeneralVLA-2 advances robot planning with improved 3D reconstruction and memory
Researchers have introduced GeneralVLA-2, an advancement in vision-language-action systems designed for robot planning. This system incorporates GeoFuse-MV3D for enhanced 3D reconstruction and an improved KnowledgeBank …
-
Poolside releases Laguna M.1, a 225B MoE model for agentic coding
Poolside has released Laguna M.1, a 225 billion parameter Mixture-of-Experts model optimized for agentic coding tasks. The model features a large sparse MoE architecture with 256 experts and global attention, enabling i…
-
Self-Harness enables LLM agents to improve their own operational harnesses
Researchers have developed a novel method called Self-Harness, enabling LLM-based agents to autonomously improve their own operational harnesses. This iterative process involves identifying model-specific failure patter…
-
Research: Interaction trajectories boost AI agent generalization
A new research paper explores the effectiveness of interaction trajectories for training AI agents, finding that standalone performance doesn't dictate teaching efficacy. Surprisingly, agents fine-tuned on trajectories …