LMSys Arena
PulseAugur coverage of LMSys Arena — every cluster mentioning LMSys Arena across labs, papers, and developer communities, ranked by signal.
-
New PRACT-120 benchmark aims to evaluate AI chatbots holistically
A new benchmark called PRACT-120 has been proposed to evaluate AI chatbots more comprehensively than existing tests like MMLU or GPQA. The benchmark aims to assess not just the core model's capabilities but also the int…
-
Prospector Labs proposes 'luckrig' for LLM hardware rig testing
Prospector Labs has introduced "luckrig," a concept for evaluating the specific hardware configurations, or "rigs," used to run large language models, rather than just the models themselves. This system aims to fill a g…
-
AI model performance chart reveals hidden degradation trends
A new chart visualizes the performance history of major AI models, tracking their capabilities over time rather than just their latest release. This tool aims to expose hidden trends like performance degradation or "ner…
-
Brightwave raises $6M seed round for AI research assistant in finance
Brightwave, an AI research assistant for investment professionals, has secured $6 million in seed funding led by Decibel. The company, which serves clients managing over $120 billion in assets, is focusing on practical …