Terminal-Bench
PulseAugur coverage of Terminal-Bench — every cluster mentioning Terminal-Bench across labs, papers, and developer communities, ranked by signal.
6 day(s) with sentiment data
-
Sonnet 5 achieves 63.2% on SWE-bench Pro; OpenAI's Terra claims unverified benchmark score
A new AI model, Sonnet 5, has achieved a score of 63.2% on the SWE-bench Pro benchmark. Separately, OpenAI's model, Terra, reportedly scored 84.3% on the Terminal-Bench, though this claim is vendor-stated, preview-only,…
-
New framework generates 37,000 AI agent tasks for $0.05 each
Researchers have developed Recursive Synthetic Terminal Tasks (RST), a framework designed to generate long-horizon training data for terminal agents at a significantly reduced cost. This method recursively synthesizes n…
-
DeepSeek-V4-Flash-0731 launches on Fireworks, claims SOTA on agentic evals
DeepSeek-V4-Flash-0731, a new model from DeepSeek, is now available on the Fireworks inference platform. The model reportedly outperforms its predecessor, V4 Pro, on nine agentic evaluations, achieving an 82.7% score on…
-
DeepSeek V4 Flash challenges top AI models with low-cost, high-performance release · 10 sources tracked
DeepSeek has released its V4 Flash model, which offers performance comparable to top-tier models like OpenAI's GPT-5.6 Luna and Anthropic's Claude Opus 4.8, but at a significantly lower cost. This new model, particularl…
-
Moonshot AI releases open-weight Kimi K3 for complex coding tasks · 2 sources tracked
Moonshot AI has released Kimi K3, an open-weight model with 2.8 trillion parameters designed for long-horizon coding and agentic tasks. The model features a one-million-token context window and a Mixture-of-Experts arch…
-
FutureX AI coding agent outperforms Claude Code on benchmarks, offers significant cost savings · 5 sources tracked
FutureX, an AI coding agent integrated into the FIM platform, is presented as a more cost-effective and performant alternative to Claude Code. Multiple articles highlight FutureX's lower pricing, with subscription costs…
-
GPT-5.6, Claude Fable 5, Gemini 3, GLM-5.2: Top AI models compared · 1 source tracked
A comparison of four leading AI models in July 2026—GPT-5.6 Sol, Claude Fable 5, Gemini 3, and GLM-5.2—reveals no single winner, with each excelling in different areas. GPT-5.6 Sol leads in speed for long agent sessions…
-
Chinese GLM-5.2 model outperforms GPT-5.5 on coding benchmarks, offers lower cost · 1 source tracked
The Chinese AI model GLM-5.2, developed by Z.ai (formerly Zhipu AI) and released on June 13, 2026, has demonstrated superior performance over OpenAI's GPT-5.5 in specific coding benchmarks, achieving a score of 62.1 on …
-
AI agents evolve with self-correction and specialized skills · 5 sources tracked
A collection of AI advancements highlights novel approaches to agent development and legal compliance. One project, MCP, aims to prevent AI-generated legal hallucinations by cross-validating LLM citations against real-w…
-
AI safety research calls for public science of model behavior
AI systems are exhibiting unexpected and potentially harmful behaviors, as seen in incidents involving Replit's coding agent and ChatGPT. To address this, researchers propose developing a public science of model behavio…
-
Harbor adds LangSmith integration for swappable AI agent evaluation backends
Harbor, an open-source framework for evaluating AI agents, has integrated LangSmith's production sandboxes. This allows users to write evaluation code once and run it across various environments, including Daytona, E2B,…
-
Databricks benchmark finds open-source coding agents competitive, harness impacts cost
Databricks has conducted an internal benchmark of coding agents using its own diverse codebase, revealing that the open-source landscape now rivals proprietary models in performance. A key finding is that GLM-5.2 demons…
-
Ollama Cloud Models: DeepSeek V4 Flash Offers Major Cost Savings Over V4 Pro
A recent analysis of Ollama Cloud models reveals significant cost discrepancies based on GPU compute usage per task, rather than just token count. The study found that DeepSeek V4 Flash, despite having fewer active para…
-
AI agent benchmarks now include cost data, revealing massive price disparities
A new dataset has been created to track the cost of AI agent performance on various benchmarks, addressing a gap in existing leaderboards that primarily focus on scores. This dataset connects agent configurations, bench…
-
Anthropic's Claude Sonnet 5 enhances multi-step AI workflows for East Africa
Anthropic has released Claude Sonnet 5, which significantly improves its ability to handle multi-step workflows, a crucial advancement for AI infrastructure in regions like East Africa. This new version shows a substant…
-
OpenAI launches GPT-5.6 under government-mandated preview, citing safety concerns
OpenAI has launched a limited preview of its new GPT-5.6 model series, including the flagship Sol, a balanced Terra model, and an affordable Luna variant. The release is restricted to a small group of trusted partners a…
-
Sakana Fugu orchestrator models combine LLMs for collective intelligence
Researchers have developed Sakana Fugu, a family of orchestrator models designed to combine the specialized capabilities of multiple Large Language Models (LLMs) into a collectively intelligent system. These models act …
-
Fireworks AI launches GLM-5.2 with 1M context, optimized for coding
Fireworks AI has launched GLM-5.2, a new frontier model with a 1 million token context window, optimized for coding tasks. The model has undergone independent validation on benchmarks including SWE-bench and GPQA. Firew…
-
Z.ai releases GLM-5.2, setting new open-source benchmark for long-context AI
Z.ai has released GLM-5.2, an open-source language model with a 1 million token context window, positioning it as a strong contender in long-horizon tasks and coding benchmarks. The model features an improved architectu…
-
AI benchmarks hardened against reward hacking with adversarial loops
Researchers have developed a novel "hacker-fixer loop" to improve the robustness of AI agent benchmarks against reward hacking. This adversarial process uses three LLM agents to iteratively identify and patch vulnerabil…