Psi-Bench
PulseAugur coverage of Ψ-Bench — every cluster mentioning Ψ-Bench across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
OpenAI's GPT-6 Astra shows major gains, especially in graphics, per analysis
Sebastian Raschka's analysis suggests OpenAI's new GPT-6 Astra model demonstrates significant improvements over its predecessor, GPT-5.6 "Sol," particularly in graphical tasks and achieving a near-perfect score on the A…
-
LLM agents slash token use by tracking state, not history · 1 source tracked
A recent preprint, SKILL.state, introduces a novel approach to LLM agent memory management, significantly reducing token usage by tracking structured state instead of conversational history. This method, tested on vario…
-
Meta releases open-weight Muse Glimmer model for agentic tasks
Meta has released Muse Glimmer, a new 30B parameter open-weight model licensed under Apache 2.0. The model is designed for end-to-end agentic task completion, reliable tool use, and multi-step reasoning, showing strong …
-
New benchmark \u03a8-Bench tests LLMs' persuasive dialogue skills
Researchers have introduced \u03a8-Bench, a new benchmark designed to evaluate the persuasive capabilities of large language models (LLMs) in conversational settings. The benchmark focuses on persona-sensitive influenci…