long-context reasoning
PulseAugur coverage of long-context reasoning — every cluster mentioning long-context reasoning across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
Sarvam 30B model performance metrics revealed across benchmarks
Sarvam AI has released its Sarvam 30B model, with performance metrics now available for several benchmarks. The model achieved 63.3% on GPQA, 7.5% on Humanity's Last Exam, and 19.2% on SciCode. Notably, it scored 0% on …
-
LLM performance benchmarks released for Llama, GLM, and Mistral models · 4 sources tracked
Independent benchmarks reveal performance metrics for several large language models, including Llama 3.2 Instruct 90B, GLM-4.7-Flash, Mistral Large 2, and Llama 3.1 Instruct 8B. The data highlights scores across various…
-
LLM benchmarks show mixed results for Kimi, Qwen3, and Exaone models
Independent benchmarks reveal varying performance across several large language models. Kimi K2 0905 achieved strong scores on GPQA and MMLU-Pro, while Qwen3 235B A22B also performed well on these metrics but struggled …
-
LLM benchmark results reveal performance across multiple models · 9 sources tracked
A recent independent benchmark evaluation has revealed performance metrics for several large language models, including Kimi K2, Sarvam Maya, NVIDIA Nemotron 3 Super 120B, DeepSeek V3.2, Falcon H1R-7B, GLM-5.2, GLM-5.1,…
-
DBRX Instruct and Mistral Medium 3 benchmark results revealed
Independent benchmarks reveal performance metrics for two large language models. DBRX Instruct achieved scores of 33.1% on GPQA, 39.7% on MMLU-Pro, 6.6% on Humanity's Last Exam, and 9.3% on LiveCodeBench. Mistral Medium…
-
Open-source LLMs show strong benchmark performance across multiple metrics · 4 sources tracked
Several open-source AI models have demonstrated strong performance on various benchmarks, according to independent measurements. Mi:dm K 2.5 Pro achieved 70.1% on GPQA and 80.9% on MMLU-Pro, while MiMo-V2-Flash showed 8…