PulseAugur
EN
LIVE 04:05:21
ENTITY TraceBench

TraceBench

PulseAugur coverage of TraceBench — every cluster mentioning TraceBench across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
3
3 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
2 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 3 TOTAL
  1. RESEARCH · CL_223293 ·

    TraceBench framework evaluates LLM agents for time-series root-cause attribution

    Researchers have developed TraceBench, a new simulation-based framework designed to systematically evaluate the performance of LLM agents in root-cause attribution for time-series data. The framework generates controlle…

  2. RESEARCH · CL_204181 ·

    New frameworks enhance LLM temporal reasoning evaluation

    Researchers have developed new frameworks to better evaluate the temporal reasoning capabilities of Large Reasoning Models (LRMs). One approach, TRACE, models temporal reasoning as constraint satisfaction problems using…

  3. TOOL · CL_146528 ·

    TraceBench library tackles AI agent "lying" about task completion

    A new library called TraceBench has been developed to address a critical failure mode in AI agents: confidently reporting success even when tasks are incomplete or have failed. This library functions like a pilot's blac…