test-time scaling
PulseAugur coverage of test-time scaling — every cluster mentioning test-time scaling across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
Layer pruning harms LLM long-chain reasoning, new study finds
A new arXiv paper reveals that layer pruning, a common technique for optimizing large language models (LLMs), can significantly harm their ability to perform long-chain reasoning. While pruning may not affect general kn…
-
New GAIA system trains critic models to improve GUI agent performance
Researchers have developed GAIA, a data flywheel system designed to improve the performance of GUI agents by training an Intuitive Critic Model (ICM). This ICM evaluates the correctness of an agent's actions, selecting …
-
New research questions excessive sampling in AI models
A new research paper explores the concept of test-time scaling in language models and reasoning systems, arguing that excessive sampling can lead to worse performance. The paper introduces the 'modal ceiling' and 'corre…
-
Survey paper maps Test-Time Scaling for multimodal AI models
A new survey paper details the emerging field of Test-Time Scaling (TTS) for Multimodal Foundation Models (MFMs). The paper categorizes existing TTS methods into sampling-based, feedback-based, and search-based approach…
-
Stochastic backtracking boosts language model reasoning efficiency
Researchers have developed a new method called stochastic backtracking to improve the efficiency of test-time scaling in language models. This technique allows models to revisit previously generated states, rather than …
-
New framework enables LLMs to share insights during reasoning
Researchers have introduced Collaborative Parallel Thinking (CPT), a novel training-free framework designed to enhance the efficiency of test-time scaling (TTS) for large language models. CPT addresses the issue of redu…