PulseAugur
EN
LIVE 11:29:52
ENTITY Terminal-Bench

Terminal-Bench

PulseAugur coverage of Terminal-Bench — every cluster mentioning Terminal-Bench across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
4
25 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
7 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

4 day(s) with sentiment data

RECENT · PAGE 1/2 · 33 TOTAL
  1. COMMENTARY · CL_249247 ·

    Cognition's SWE-2 excels on easy tasks but struggles with hard adversarial problems

    Cognition's new SWE-2 model excels on easy benchmarks, achieving near-perfect scores and offering a strong price-performance ratio for simpler tasks. However, a deeper analysis of its performance on harder, adversarial …

  2. SIGNIFICANT · CL_248227 ·

    Nvidia releases quantized Alibaba Qwen3.8-27B model for AI agents

    Nvidia has released a quantized version of Alibaba's Qwen3.8-27B language model, optimized for deployment in AI agent systems and other applications. This model, named nvidia/Qwen3.8-27B-NVFP4, utilizes Nvidia's Model O…

  3. TOOL · CL_226646 ·

    New Agentic Coding Index ranks LLMs by intelligence density

    A Reddit user has developed a new metric called the Agentic Coding Index (ACI) to evaluate Large Language Models (LLMs) on coding tasks. The ACI aggregates scores from several coding benchmarks, including SWE-bench Pro,…

  4. RESEARCH · CL_222730 ·

    New benchmark tests AI agents on 70 scientific research workflows

    A new benchmark, Terminal-Bench-Science, has been released to evaluate AI agents on complex scientific research tasks. Developed by researchers at Stanford University and a global team of domain experts, it features 70 …

  5. TOOL · CL_217880 ·

    New benchmark automates detection of AI reward hacking

    Researchers have developed the Hack-Verifiable Terminal Bench (HVTB) to automatically detect reward hacking in AI agents. This new benchmark adapts the Hack-Verifiable Environments (HVE) methodology to Terminal Bench, a…

  6. TOOL · CL_216552 ·

    27B AI Model Achieves 5.4% Terminal-Bench Score via Novel RL Training

    A 27 billion parameter model achieved a significant improvement on the Terminal-Bench benchmark, raising its score from 1.4% to 5.4%. This advancement was not due to increased model size, but rather from a novel trainin…

  7. RESEARCH · CL_205657 ·

    New framework audits vendor-hosted LLM APIs for quality degradation

    Researchers have developed Ventor-QTest, a novel black-box auditing framework designed to verify the quality of inference APIs for vendor-hosted large language models. This method employs both repeated-request and long-…

  8. RESEARCH · CL_199689 ·

    DeepSeek V4 models debut with staggered, multi-stage release strategy

    DeepSeek has adopted a staggered release strategy for its V4 models, beginning with an open-weight preview of V4-Pro and V4-Flash in April 2026. The V4-Flash model then graduated to general availability (GA) on July 31,…

  9. TOOL · CL_186078 ·

    Sonnet 5 achieves 63.2% on SWE-bench Pro; OpenAI's Terra claims unverified benchmark score

    A new AI model, Sonnet 5, has achieved a score of 63.2% on the SWE-bench Pro benchmark. Separately, OpenAI's model, Terra, reportedly scored 84.3% on the Terminal-Bench, though this claim is vendor-stated, preview-only,…

  10. RESEARCH · CL_186968 ·

    New framework generates 37,000 AI agent tasks for $0.05 each

    Researchers have developed Recursive Synthetic Terminal Tasks (RST), a framework designed to generate long-horizon training data for terminal agents at a significantly reduced cost. This method recursively synthesizes n…

  11. SIGNIFICANT · CL_175738 ·

    DeepSeek-V4-Flash-0731 launches on Fireworks, claims SOTA on agentic evals

    DeepSeek-V4-Flash-0731, a new model from DeepSeek, is now available on the Fireworks inference platform. The model reportedly outperforms its predecessor, V4 Pro, on nine agentic evaluations, achieving an 82.7% score on…

  12. FRONTIER RELEASE · CL_170798 ·

    DeepSeek V4 Flash challenges top AI models with low-cost, high-performance release · 10 sources tracked

    DeepSeek has released its V4 Flash model, which offers performance comparable to top-tier models like OpenAI's GPT-5.6 Luna and Anthropic's Claude Opus 4.8, but at a significantly lower cost. This new model, particularl…

  13. SIGNIFICANT · CL_166961 ·

    Moonshot AI releases open-weight Kimi K3 for complex coding tasks · 2 sources tracked

    Moonshot AI has released Kimi K3, an open-weight model with 2.8 trillion parameters designed for long-horizon coding and agentic tasks. The model features a one-million-token context window and a Mixture-of-Experts arch…

  14. TOOL · CL_164640 ·

    FutureX AI coding agent outperforms Claude Code on benchmarks, offers significant cost savings · 5 sources tracked

    FutureX, an AI coding agent integrated into the FIM platform, is presented as a more cost-effective and performant alternative to Claude Code. Multiple articles highlight FutureX's lower pricing, with subscription costs…

  15. COMMENTARY · CL_137312 ·

    GPT-5.6, Claude Fable 5, Gemini 3, GLM-5.2: Top AI models compared · 1 source tracked

    A comparison of four leading AI models in July 2026—GPT-5.6 Sol, Claude Fable 5, Gemini 3, and GLM-5.2—reveals no single winner, with each excelling in different areas. GPT-5.6 Sol leads in speed for long agent sessions…

  16. SIGNIFICANT · CL_136891 ·

    Chinese GLM-5.2 model outperforms GPT-5.5 on coding benchmarks, offers lower cost · 1 source tracked

    The Chinese AI model GLM-5.2, developed by Z.ai (formerly Zhipu AI) and released on June 13, 2026, has demonstrated superior performance over OpenAI's GPT-5.5 in specific coding benchmarks, achieving a score of 62.1 on …

  17. TOOL · CL_134920 ·

    AI agents evolve with self-correction and specialized skills · 5 sources tracked

    A collection of AI advancements highlights novel approaches to agent development and legal compliance. One project, MCP, aims to prevent AI-generated legal hallucinations by cross-validating LLM citations against real-w…

  18. TOOL · CL_134899 ·

    AI safety research calls for public science of model behavior

    AI systems are exhibiting unexpected and potentially harmful behaviors, as seen in incidents involving Replit's coding agent and ChatGPT. To address this, researchers propose developing a public science of model behavio…

  19. TOOL · CL_134310 ·

    Harbor adds LangSmith integration for swappable AI agent evaluation backends

    Harbor, an open-source framework for evaluating AI agents, has integrated LangSmith's production sandboxes. This allows users to write evaluation code once and run it across various environments, including Daytona, E2B,…

  20. RESEARCH · CL_133383 ·

    Databricks benchmark finds open-source coding agents competitive, harness impacts cost

    Databricks has conducted an internal benchmark of coding agents using its own diverse codebase, revealing that the open-source landscape now rivals proprietary models in performance. A key finding is that GLM-5.2 demons…