\tau2-bench
PulseAugur coverage of \tau2-bench — every cluster mentioning \tau2-bench across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Paper questions LLM-agent leaderboard validity
A new paper questions the validity of LLM-agent leaderboards, arguing that direct comparisons of ranked agents can be misleading. The authors highlight that differences in task mixtures, data sources, release details, a…
-
TaichuAI releases ZDTaichu5.0-9B multimodal foundation model
TaichuAI has released ZDTaichu5.0-9B, a multimodal foundation model designed for visual understanding, spatial reasoning, and agentic tasks. This model integrates a Qwen3.5-9B language backbone with a C-RADIOv4-H vision…
-
Apple's PROOF-Gen method enhances AI model distillation from failures
Apple Machine Learning Research has introduced PROOF-Gen, a novel method for improving the distillation of tool-calling capabilities into deployable AI models. This technique addresses the limitations of traditional gen…
-
New research evaluates open-source AI model routers across benchmarks
A new paper published on arXiv evaluates four open-source model routers, which are systems that delegate model selection in agentic frameworks. The study introduces a common measurement protocol to compare these routers…
-
llama.cpp fixes muse-glimmer tool call parsing issue
The llama.cpp project has released an update (b10380) to fix an issue where the muse-glimmer model incorrectly handled tool calls. Previously, the model would sometimes absorb tool call markup into its response content,…
-
New method improves multi-turn AI agents with preference learning · 2 sources tracked
Researchers have developed a novel method called ToolGraph, which enhances multi-turn tool-using agents by integrating schema-derived topology and transition weights from successful rollouts. This approach improves the …