AIME2025
PulseAugur coverage of AIME2025 — every cluster mentioning AIME2025 across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
AI agent Intern-S1-MO tackles Olympiad-level math problems
Researchers have developed Intern-S1-MO, a novel long-horizon reasoning agent designed to tackle complex mathematical problems at the Olympiad level. This agent employs a multi-round, hierarchical reasoning approach usi…
-
New method links LLM evaluation failures to targeted data fixes
Researchers have developed a novel method to bridge the gap between model capability evaluation and data curation in large language models. Their approach, termed the "capability slice," allows for precise localization …
-
New VISTA framework enhances LLM prompt optimization
Researchers have developed VISTA, a new framework for automatically optimizing prompts used with large language models. This method aims to overcome limitations in existing reflective prompt optimization techniques, whi…
-
PiCSAR method boosts LLM reasoning chain accuracy with probabilistic confidence scoring
Researchers have introduced PiCSAR, a novel method for improving the accuracy of large language and reasoning models. This training-free approach enhances performance on reasoning tasks by selecting the best candidate s…