PulseAugur
EN
LIVE 08:42:15
ENTITY long-context reasoning

long-context reasoning

PulseAugur coverage of long-context reasoning — every cluster mentioning long-context reasoning across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
4
6 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
1
1 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

3 day(s) with sentiment data

RECENT · PAGE 1/1 · 6 TOTAL
  1. TOOL · CL_203015 ·

    Sarvam 30B model performance metrics revealed across benchmarks

    Sarvam AI has released its Sarvam 30B model, with performance metrics now available for several benchmarks. The model achieved 63.3% on GPQA, 7.5% on Humanity's Last Exam, and 19.2% on SciCode. Notably, it scored 0% on …

  2. RESEARCH · CL_201933 ·

    LLM performance benchmarks released for Llama, GLM, and Mistral models · 4 sources tracked

    Independent benchmarks reveal performance metrics for several large language models, including Llama 3.2 Instruct 90B, GLM-4.7-Flash, Mistral Large 2, and Llama 3.1 Instruct 8B. The data highlights scores across various…

  3. RESEARCH · CL_188546 ·

    LLM benchmarks show mixed results for Kimi, Qwen3, and Exaone models

    Independent benchmarks reveal varying performance across several large language models. Kimi K2 0905 achieved strong scores on GPQA and MMLU-Pro, while Qwen3 235B A22B also performed well on these metrics but struggled …

  4. RESEARCH · CL_161409 ·

    LLM benchmark results reveal performance across multiple models · 9 sources tracked

    A recent independent benchmark evaluation has revealed performance metrics for several large language models, including Kimi K2, Sarvam Maya, NVIDIA Nemotron 3 Super 120B, DeepSeek V3.2, Falcon H1R-7B, GLM-5.2, GLM-5.1,…

  5. RESEARCH · CL_152721 ·

    DBRX Instruct and Mistral Medium 3 benchmark results revealed

    Independent benchmarks reveal performance metrics for two large language models. DBRX Instruct achieved scores of 33.1% on GPQA, 39.7% on MMLU-Pro, 6.6% on Humanity's Last Exam, and 9.3% on LiveCodeBench. Mistral Medium…

  6. RESEARCH · CL_137517 ·

    Open-source LLMs show strong benchmark performance across multiple metrics · 4 sources tracked

    Several open-source AI models have demonstrated strong performance on various benchmarks, according to independent measurements. Mi:dm K 2.5 Pro achieved 70.1% on GPQA and 80.9% on MMLU-Pro, while MiMo-V2-Flash showed 8…