PulseAugur
EN
LIVE 10:48:50
ENTITY SWE-bench

SWE-bench

PulseAugur coverage of SWE-bench — every cluster mentioning SWE-bench across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
35
115 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
15
57 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

22 day(s) with sentiment data

RECENT · PAGE 1/6 · 115 TOTAL
  1. FRONTIER RELEASE · CL_192575 ·

    Meta releases open-weight Muse Glimmer model for agentic tasks

    Meta has released Muse Glimmer, a new 30B parameter open-weight model licensed under Apache 2.0. The model is designed for end-to-end agentic task completion, reliable tool use, and multi-step reasoning, showing strong …

  2. TOOL · CL_191613 ·

    Developer creates custom harness to benchmark coding LLMs on personal bugs

    A developer has created a Python-based harness to evaluate coding LLMs against a personal corpus of bugs, rather than relying on public benchmarks like SWE-bench. This approach aims to provide more relevant performance …

  3. FRONTIER RELEASE · CL_191809 ·

    Meta releases Muse Glimmer, a 30B open-weight model for local AI agents

    Meta has released Muse Glimmer, a 30-billion-parameter open-weight model optimized for local agentic workflows. This model is designed to run on consumer hardware, such as a single GPU, making it accessible for personal…

  4. TOOL · CL_189473 ·

    Token-saving tools for AI coding agents over-promise, benchmark reveals

    A recent benchmark of five token-saving tools across coding agents revealed that their advertised savings of 60-90% did not hold up in real-world agent workloads. The study, which used 48 Django questions from SWE-bench…

  5. TOOL · CL_188540 ·

    NVIDIA releases NOOA framework for unified AI agent development

    NVIDIA Labs has released NOOA, an open-source Python framework designed to simplify AI agent development by encapsulating all components within a single Python class. This approach integrates prompt templates, tool sche…

  6. SIGNIFICANT · CL_187112 ·

    OpenAI launches GPT-5 with real-time reasoning, boosting performance

    OpenAI has launched GPT-5, its latest language model, which features enhanced real-time reasoning capabilities allowing it to think through problems step-by-step before responding. This new model demonstrates significan…

  7. COMMENTARY · CL_186257 ·

    Sonnet 5 with Graft tool outperforms Opus 5 for coding tasks

    A user on Reddit's r/ClaudeAI subreddit has found that combining Sonnet 5 with a tool called Graft outperforms Opus 5 for most coding tasks. Graft creates a context graph of a code repository, allowing Claude Code to au…

  8. TOOL · CL_185260 ·

    COMPAS method optimizes code generation by jointly tuning models, prompts, and settings

    Researchers have developed COMPAS, a novel method for optimizing code generation by jointly searching over models, prompts, and decoding settings. This difficulty-aware approach learns group-specific quality-cost fronts…

  9. TOOL · CL_184575 ·

    LLM benchmark scores fail in production due to "saturation paradox"

    Static academic benchmarks are becoming less effective for evaluating enterprise LLMs due to the "Benchmark Saturation Paradox," where models scoring highly on leaderboards like MMLU and SWE-bench perform poorly on real…

  10. TOOL · CL_183242 ·

    New tool PAIChecker identifies PR-issue misalignment in LLM benchmarks

    Researchers have developed PAIChecker, a multi-agent system designed to identify and correct misalignments between pull requests (PRs) and their associated issues in benchmarks used to evaluate large language models (LL…

  11. TOOL · CL_181244 ·

    Bespoke Labs seeks researcher for long-horizon agent benchmarks

    Bespoke Labs is seeking a researcher to develop and evaluate reinforcement learning environments and benchmarks for long-horizon agent tasks. The role requires demonstrated experience with multi-step reasoning agents, s…

  12. TOOL · CL_178266 ·

    New research identifies "Agentic Formalism Trap" in LLM evaluators

    A new research paper introduces the "Agentic Formalism Trap" and the "Evaluative Dissonance Index" ($D_E$) to measure how Large Language Model (LLM) based evaluation systems can be misled by consensus mimicry under adve…

  13. COMMENTARY · CL_176974 ·

    LLaMA community seeks best benchmarks for local model performance validation

    Users on the r/LocalLLaMA subreddit are seeking recommendations for the best benchmarks to evaluate the performance of their locally run large language models. The original poster is looking for ways to measure their mo…

  14. COMMENTARY · CL_172646 ·

    2026 LLM Benchmark: No Single Winner, Specialized Leaders Emerge · 1 source tracked

    A comprehensive benchmark of 20 leading LLMs in 2026 reveals no single dominant model, but rather specialized leaders across different tasks. Claude Opus 5 leads the overall Artificial Analysis Intelligence Index, while…

  15. COMMENTARY · CL_172249 ·

    LLM routing shifts from prediction to verification for cost savings

    The dominant inefficiency in LLM serving is not model speed or quantization, but rather the misallocation of compute resources. Research indicates that ex-ante prediction of request difficulty is unreliable for routing …

  16. RESEARCH · CL_178294 ·

    LLMs struggle to delete code, hindering maintainability, new research finds

    A new research paper identifies "deletion avoidance" as a key issue in large language models' code editing capabilities, where models tend to retain code that should be removed. This behavior, observed across leading mo…

  17. TOOL · CL_167466 ·

    CodexGraph system enhances LLM interaction with code repositories

    Researchers have developed CodexGraph, a novel system designed to improve how large language models (LLMs) interact with entire code repositories. Unlike existing methods that rely on similarity retrieval or task-specif…

  18. TOOL · CL_167390 ·

    WISERouter framework optimizes LLM routing with budget constraints

    Researchers have introduced WISERouter, a novel framework designed to optimize Large Language Model (LLM) routing by balancing performance and cost. This system addresses limitations in current methods, such as heuristi…

  19. TOOL · CL_164640 ·

    FutureX AI coding agent outperforms Claude Code on benchmarks, offers significant cost savings · 5 sources tracked

    FutureX, an AI coding agent integrated into the FIM platform, is presented as a more cost-effective and performant alternative to Claude Code. Multiple articles highlight FutureX's lower pricing, with subscription costs…

  20. TOOL · CL_161824 ·

    SWE-bench highlights need for human-in-the-loop reasoning in AI models

    Medium-difficulty software engineering issues require a structured human-in-the-loop process for AI models to learn effective reasoning. These problems challenge models to understand underlying code intent, expected beh…