SciCode
PulseAugur coverage of SciCode — every cluster mentioning SciCode across labs, papers, and developer communities, ranked by signal.
6 day(s) with sentiment data
-
DeepSeek V4 Pro 0813 benchmarks show 10.6% on Humanity's Last Exam
DeepSeek V4 Pro 0813 has demonstrated specific performance metrics on benchmarks including Humanity's Last Exam, Long Context Reasoning, and SciCode. The model achieved 10.6% on Humanity's Last Exam, 50.7% on Long Conte…
-
Open-source project Livenerf tracks Opus 5.5 performance for potential 'nerfs'
An open-source project called Livenerf has been developed to independently track and measure changes in the performance of Anthropic's Opus 5.5 model over time. The project uses established benchmarks like SWE-bench and…
-
DeepSeek V4 Pro and Nova 2.0 Lite benchmarks revealed · 2 sources tracked
Independent benchmarks reveal the performance of two large language models, DeepSeek V4 Pro and Nova 2.0 Lite. DeepSeek V4 Pro achieved high scores in reasoning-focused tasks, with 92.8% on GPQA and 80.3% on Long Contex…
-
New Quantum-Classical Hybrid AI Architecture Boosts Long-Horizon Reasoning
Researchers have introduced QART, a novel quantum-classical hybrid architecture designed to improve long-horizon reasoning in AI models. QART integrates a backbone language model with quantum encoding, optimization, and…
-
Artificial Analysis Intelligence Index v4.2 released, Anthropic leads rankings
Artificial Analysis has released version 4.2 of its Intelligence Index, introducing new evaluations like AA-Briefcase for agentic knowledge work and GDP.pdf for long-context document reasoning. This update increases the…
-
Nvidia releases quantized Alibaba Qwen3.8-27B model for AI agents
Nvidia has released a quantized version of Alibaba's Qwen3.8-27B language model, optimized for deployment in AI agent systems and other applications. This model, named nvidia/Qwen3.8-27B-NVFP4, utilizes Nvidia's Model O…
-
Qwen3.5 2B shows mixed results in independent benchmarks
Independent benchmarks reveal that Qwen3.5 2B, a non-reasoning model, achieves 43.8% on the GPQA benchmark. However, its performance significantly drops on more complex tasks, scoring only 5% on HLE, 15% on Long Context…
-
Kimi K2.5 achieves strong benchmark scores with competitive pricing
Kimi K2.5 has achieved notable performance on several benchmarks, including GPQA, HLE, Long Context, and SciCode. The model offers competitive pricing at 30 integer points per dollar across these evaluations. These resu…
-
AI models nearing saturation on scientific coding benchmarks, Reddit discussion reveals
A discussion on Reddit's r/singularity forum explores the saturation point of AI models on scientific coding benchmarks, specifically SciCode, HLE, and CritPt. The analysis suggests that while HLE and CritPt show monthl…
-
Sarvam 30B model performance metrics revealed across benchmarks
Sarvam AI has released its Sarvam 30B model, with performance metrics now available for several benchmarks. The model achieved 63.3% on GPQA, 7.5% on Humanity's Last Exam, and 19.2% on SciCode. Notably, it scored 0% on …
-
LLM performance benchmarks released for Llama, GLM, and Mistral models · 4 sources tracked
Independent benchmarks reveal performance metrics for several large language models, including Llama 3.2 Instruct 90B, GLM-4.7-Flash, Mistral Large 2, and Llama 3.1 Instruct 8B. The data highlights scores across various…
-
Meta's Muse Spark 1.2 shows rapid performance gains, rivals top AI models
Meta's latest foundational model, Muse Spark 1.2, has achieved high scores in third-party performance analyses, demonstrating rapid improvement since the Muse series' debut four months ago. The model notably surpassed G…
-
Gemma 4's top ranking on SciCode benchmark questioned by users
A user on Reddit's r/LocalLLaMA community is questioning the ranking of Gemma 4 above Qwen-3.6 27B on the SciCode benchmark, as reported by artificialanalysis.ai. The user expresses surprise, stating that this ranking c…
-
New method enhances LLM scientific computing by consolidating experience
Researchers have developed a new method called SciConsolidate to improve the scientific computing capabilities of large language models. This technique converts runtime experience from solving problems into transferable…
-
LLM benchmark results reveal performance across multiple models · 9 sources tracked
A recent independent benchmark evaluation has revealed performance metrics for several large language models, including Kimi K2, Sarvam Maya, NVIDIA Nemotron 3 Super 120B, DeepSeek V3.2, Falcon H1R-7B, GLM-5.2, GLM-5.1,…
-
Open-source LLMs show strong benchmark performance across multiple metrics · 4 sources tracked
Several open-source AI models have demonstrated strong performance on various benchmarks, according to independent measurements. Mi:dm K 2.5 Pro achieved 70.1% on GPQA and 80.9% on MMLU-Pro, while MiMo-V2-Flash showed 8…
-
AI research introduces new methods for benchmark evolution and agent self-reconfiguration
Two new research papers introduce novel methods for advancing AI capabilities. BenchEvolver focuses on creating more challenging coding benchmarks by evolving existing problems, aiming to overcome benchmark saturation a…
-
NVIDIA quantizes Alibaba's Qwen3.6-35B model for efficient deployment
NVIDIA has released a quantized version of Alibaba's Qwen3.6-35B-A3B model, named nvidia/Qwen3.6-35B-A3B-NVFP4. This model utilizes the NVFP4 data type, reducing memory requirements by approximately 3.06x while maintain…
-
No Test Cases, No Problem: Distillation-Driven Code Generation for Scientific Workflows
Researchers have developed MOSAIC, a novel framework for generating code for scientific workflows without relying on traditional input/output test cases. This new approach utilizes a knowledge distillation technique, wh…