SciCode
PulseAugur coverage of SciCode — every cluster mentioning SciCode across labs, papers, and developer communities, ranked by signal.
5 day(s) with sentiment data
-
Sarvam 30B model performance metrics revealed across benchmarks
Sarvam AI has released its Sarvam 30B model, with performance metrics now available for several benchmarks. The model achieved 63.3% on GPQA, 7.5% on Humanity's Last Exam, and 19.2% on SciCode. Notably, it scored 0% on …
-
Meta's Muse Spark 1.2 shows rapid performance gains, rivals top AI models
Meta's latest foundational model, Muse Spark 1.2, has achieved high scores in third-party performance analyses, demonstrating rapid improvement since the Muse series' debut four months ago. The model notably surpassed G…
-
Gemma 4's top ranking on SciCode benchmark questioned by users
A user on Reddit's r/LocalLLaMA community is questioning the ranking of Gemma 4 above Qwen-3.6 27B on the SciCode benchmark, as reported by artificialanalysis.ai. The user expresses surprise, stating that this ranking c…
-
New method enhances LLM scientific computing by consolidating experience
Researchers have developed a new method called SciConsolidate to improve the scientific computing capabilities of large language models. This technique converts runtime experience from solving problems into transferable…
-
LLM benchmark results reveal performance across multiple models · 9 sources tracked
A recent independent benchmark evaluation has revealed performance metrics for several large language models, including Kimi K2, Sarvam Maya, NVIDIA Nemotron 3 Super 120B, DeepSeek V3.2, Falcon H1R-7B, GLM-5.2, GLM-5.1,…
-
Open-source LLMs show strong benchmark performance across multiple metrics · 4 sources tracked
Several open-source AI models have demonstrated strong performance on various benchmarks, according to independent measurements. Mi:dm K 2.5 Pro achieved 70.1% on GPQA and 80.9% on MMLU-Pro, while MiMo-V2-Flash showed 8…
-
AI research introduces new methods for benchmark evolution and agent self-reconfiguration
Two new research papers introduce novel methods for advancing AI capabilities. BenchEvolver focuses on creating more challenging coding benchmarks by evolving existing problems, aiming to overcome benchmark saturation a…
-
NVIDIA quantizes Alibaba's Qwen3.6-35B model for efficient deployment
NVIDIA has released a quantized version of Alibaba's Qwen3.6-35B-A3B model, named nvidia/Qwen3.6-35B-A3B-NVFP4. This model utilizes the NVFP4 data type, reducing memory requirements by approximately 3.06x while maintain…
-
No Test Cases, No Problem: Distillation-Driven Code Generation for Scientific Workflows
Researchers have developed MOSAIC, a novel framework for generating code for scientific workflows without relying on traditional input/output test cases. This new approach utilizes a knowledge distillation technique, wh…