PulseAugur
EN
LIVE 18:10:58
ENTITY GSM8K

GSM8K

PulseAugur coverage of GSM8K — every cluster mentioning GSM8K across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
45
133 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
37
120 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

20 day(s) with sentiment data

RECENT · PAGE 1/7 · 133 TOTAL
  1. RESEARCH · CL_193469 ·

    New methods tackle LLM and VLM hallucinations with internal analysis · 2 sources tracked

    Researchers have developed new methods to detect hallucinations in large language and vision-language models. UniProbe, a technique for Large VLMs, uses a graph neural network, a Vision Transformer, and a gated recurren…

  2. TOOL · CL_193280 ·

    New EL-DGR framework improves LLM judge performance in reasoning pipelines

    A new research paper introduces Evidence-Locked Derive-Gate-Repair (EL-DGR), a novel decision-making framework for LLM judges within reasoning pipelines. The study demonstrates that EL-DGR significantly improves perform…

  3. RESEARCH · CL_191278 ·

    Diffusion language models research tackles efficiency and confidence gaps · 6 sources tracked

    Recent research explores methods to improve the efficiency and effectiveness of diffusion language models (DLMs). One paper investigates when classifier-free guidance (CFG) is truly necessary during decoding, suggesting…

  4. TOOL · CL_191183 ·

    New bandit algorithm tackles LLM refinement with reward decay modeling

    Researchers have developed a new contextual bandit algorithm designed to improve iterative refinement in Large Language Models (LLMs). This algorithm explicitly models reward decay, addressing the issue of over-exploita…

  5. TOOL · CL_188515 ·

    LLM JSON optimization shows mixed results across models

    An optimization involving a change in JSON field representation for LLMs showed promising results on the Qwen2.5-7B model, improving correctness on the GSM8K benchmark. However, this optimization failed to translate to …

  6. TOOL · CL_185321 ·

    G-Boost framework enhances edge SLMs via LLM collaboration

    Researchers have developed G-Boost, a novel framework designed to enhance the performance of small language models (SLMs) deployed on edge devices. This system enables collaboration between resource-constrained edge SLM…

  7. TOOL · CL_185297 ·

    Research audits latent communication in multi-agent LLMs

    A new research paper investigates the effectiveness of latent communication in multi-agent large language models, specifically examining the role of relayed key-value (KV) caches. The study causally audits these systems…

  8. TOOL · CL_184575 ·

    LLM benchmark scores fail in production due to "saturation paradox"

    Static academic benchmarks are becoming less effective for evaluating enterprise LLMs due to the "Benchmark Saturation Paradox," where models scoring highly on leaderboards like MMLU and SWE-bench perform poorly on real…

  9. TOOL · CL_184013 ·

    Qwen2.5-7B-Instruct accuracy drops with JSON constraints, but can be recovered

    A study on the Qwen2.5-7B-Instruct model revealed that enforcing strict JSON output schemas, while ensuring compliance, can reduce mathematical accuracy by up to 18.4 percentage points. This reduction was attributed to …

  10. TOOL · CL_184016 ·

    Karpathy's nanochat uses simplified GRPO for RL loop

    Andrej Karpathy's nanochat project includes a simplified reinforcement learning loop, labeled GRPO, that deviates from the standard GRPO algorithm. This loop uses a basic policy gradient method, essentially REINFORCE wi…

  11. TOOL · CL_183278 ·

    AI Model Robustness Analysis Reveals Layer Dissociation

    A new research paper analyzes the perturbation robustness of language models, revealing that sensitivity, causality, and repair capacity do not align across model layers. The study found two distinct propagation regimes…

  12. TOOL · CL_183067 ·

    New LLM prompting methods may outperform Chain-of-Thought

    A new paper suggests that standard Chain-of-Thought (CoT) prompting may be becoming less effective for advanced large language models (LLMs). Researchers found that for certain reasoning tasks, particularly in mathemati…

  13. TOOL · CL_180590 ·

    New Structured Recurrent Mixer architecture boosts AI sequence generation efficiency

    Researchers have introduced the Structured Recurrent Mixer (SRM), a novel architecture designed to enhance sequence generation efficiency. SRMs can switch between parallel processing during training and recurrent proces…

  14. RESEARCH · CL_180530 ·

    New research boosts LLM speculative decoding speed and efficiency · 4 sources tracked

    Four new research papers published on arXiv introduce novel techniques to enhance speculative decoding for large language models. These methods aim to improve generation speed and efficiency without requiring additional…

  15. TOOL · CL_180488 ·

    New LLM method uses hidden-state geometry for better reasoning

    Researchers have developed Cloud-ScPO, a novel framework for semi-supervised preference optimization in large language models (LLMs) that leverages the geometric structure of internal model states. This method uses a sm…

  16. TOOL · CL_178938 ·

    TAPR refines LLM prompts to boost benchmark accuracy

    TAPR, a new system, automatically refines user prompts for large language models using reinforcement learning. This process enhances the accuracy of LLM responses, particularly on benchmarks like Natural Questions and GSM8K.

  17. TOOL · CL_171824 ·

    Constitutional Midtraining Enhances AI Alignment Durability

    Researchers have developed a method called constitutional midtraining to improve the durability of AI alignment. By integrating principled, values-based content into the midtraining phase of AI development, models demon…

  18. RESEARCH · CL_171885 ·

    New methods enable language models to perform "any-order inference"

    Researchers have developed new methods to enable language models to perform "any-order inference," a non-causal reasoning process similar to how programmers draft code by moving between high-level concepts and specific …

  19. TOOL · CL_179971 ·

    Constitutional Midtraining boosts AI alignment durability

    Researchers have introduced "Constitutional Midtraining," a novel approach to enhance the durability of AI alignment. By embedding values-based content into the model's training phase, rather than solely after, this met…

  20. TOOL · CL_167185 ·

    Masked distillation trains LLMs to internalize reasoning steps

    Researchers have developed a new method called masked distillation to train language models to internalize the computational steps of reasoning, thereby reducing latency and cost. This technique trains a student model t…