PulseAugur
EN
LIVE 13:49:28
ENTITY Math-500

Math-500

PulseAugur coverage of Math-500 — every cluster mentioning Math-500 across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
5
24 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
3
21 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

5 day(s) with sentiment data

RECENT · PAGE 1/3 · 43 TOTAL
  1. TOOL · CL_255981 ·

    7B model ZGCM-1 prioritizes tool use and large context over memorization

    Researchers from Zhongguancun Academy and Zhongguancun Institute of AI have developed ZGCM-1, a 7.39B parameter model that prioritizes tool use and a large context window over memorizing vast datasets. This approach all…

  2. TOOL · CL_248458 ·

    Qwen 3 4B Base model sees 31% boost on MATH-500 after puzzle fine-tuning

    A fine-tuned version of the Qwen 3 4B Base model demonstrated a 31% improvement on the MATH-500 benchmark after being trained on 100 zebra puzzles. The process for reproducing this result, which took approximately 6.5 m…

  3. TOOL · CL_234641 ·

    R1 1776 achieves 95.4% on MATH-500 benchmark

    R1 1776, an AI model, has achieved a score of 95.4% on the MATH-500 benchmark. This performance was independently verified and is available for comparison against other models on the olud.ai leaderboard.

  4. TOOL · CL_228710 ·

    AI reasoning systems trained with multi-solver disagreement reward show improved performance

    Researchers have developed a novel method for training AI reasoning systems by using disagreement among multiple models to generate challenging questions. This approach, called multi-solver disagreement reward, contrast…

  5. RESEARCH · CL_225917 ·

    AI research advances inference, optimization, and mobile benchmarking

    Researchers are exploring advanced techniques for improving AI inference and statistical analysis, particularly in resource-constrained environments. One paper introduces IMABO, a framework for Online Hyperparameter Opt…

  6. TOOL · CL_216092 ·

    New RODE optimizer decouples neural network training dynamics

    Researchers have introduced RODE, a novel optimization engine for neural networks that decouples the radial and directional components of matrix updates. This separation allows for distinct update rules and step sizes, …

  7. TOOL · CL_216033 ·

    New research reveals multilingual bias in LLM math training rewards

    A new research paper identifies a significant bias in multilingual reinforcement learning with verifiable rewards (RLVR), a common technique for training large language models on mathematical reasoning. The study found …

  8. TOOL · CL_200005 ·

    New method TAM reduces language model memory usage for reasoning

    Researchers have developed a new method called Thought-Aware Attention Matching (TAM) to address the memory bottleneck caused by lengthy reasoning sequences in language models. TAM segments reasoning trajectories into b…

  9. TOOL · CL_198086 ·

    Indian AI models strong on old benchmarks, lag in new evaluations

    A new paper assesses the progress of Indian foundation models by analyzing publicly reported benchmark results. While Indian models show strong performance on established benchmarks like MMLU and MATH-500, they lag in p…

  10. TOOL · CL_195936 ·

    New metric measures semantic abstractness of LLM features

    Researchers have introduced a new metric called Feature Nonlocality (FNL) to better understand the semantic abstractness of features within Sparse Autoencoders (SAEs) used in Large Language Models (LLMs). FNL measures t…

  11. TOOL · CL_193780 ·

    New stress test method reveals vulnerabilities in AI reward models

    Researchers have developed a new method for stress-testing process reward models (PRMs) used in AI training and search. This quality-diversity search approach, utilizing MAP-Elites, aims to identify and quantify vulnera…

  12. TOOL · CL_193612 ·

    New DRBENCHER benchmark tests AI agents' combined browsing and math skills

    Researchers have introduced DRBENCHER, a new benchmark designed to evaluate AI agents' ability to combine web browsing with multi-step mathematical computations. Unlike previous benchmarks that assess these skills in is…

  13. TOOL · CL_193611 ·

    New framework detects hidden behavioral entanglement in LLMs

    Researchers have developed a new statistical framework to detect and quantify behavioral entanglement among large language models (LLMs). This framework uses information-theoretic metrics, specifically a Difficulty-Weig…

  14. RESEARCH · CL_191155 ·

    New research explores test-time scaling for LLM reasoning

    Two new research papers introduce novel methods for improving the reasoning capabilities of large language models (LLMs) through test-time scaling. The first paper, 'Consilience,' addresses limitations in existing confi…

  15. TOOL · CL_187337 ·

    New ABC-GRPO algorithm enhances LLM training stability and performance

    Researchers have introduced All-Quadrant Bounded Clipping GRPO (ABC-GRPO), a novel algorithm designed to improve the stability and generalizability of reinforcement learning for large language models. ABC-GRPO addresses…

  16. TOOL · CL_185321 ·

    G-Boost framework enhances edge SLMs via LLM collaboration

    Researchers have developed G-Boost, a novel framework designed to enhance the performance of small language models (SLMs) deployed on edge devices. This system enables collaboration between resource-constrained edge SLM…

  17. TOOL · CL_182645 ·

    DeepSeek-V2.5 achieves 76.3% on MATH-500 benchmark

    DeepSeek-V2.5, slated for release in December 2024, has achieved a score of 76.3% on the MATH-500 benchmark. This performance metric was independently verified, distinguishing it from self-reported figures often found i…

  18. RESEARCH · CL_180530 ·

    New research boosts LLM speculative decoding speed and efficiency · 4 sources tracked

    Four new research papers published on arXiv introduce novel techniques to enhance speculative decoding for large language models. These methods aim to improve generation speed and efficiency without requiring additional…

  19. TOOL · CL_167436 ·

    New LOCKS method drastically cuts LLM long-context decoding latency

    Researchers have developed a new method called LOCKS (Page-Local Compact Key Summaries) to improve the efficiency of long-context decoding in large language models. This technique addresses the bottleneck caused by the …

  20. RESEARCH · CL_154325 ·

    New methods tackle LLM long-context efficiency challenges · 3 sources tracked

    Researchers are developing new methods to improve the efficiency of long-context reasoning in large language models. One approach, LISA, combines linear attention with a sparse attention mechanism to reduce computationa…