PulseAugur
EN
LIVE 05:49:47
ENTITY mathematics-dataset

mathematics-dataset

PulseAugur coverage of mathematics-dataset — every cluster mentioning mathematics-dataset across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
4
38 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
32 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

4 day(s) with sentiment data

RECENT · PAGE 1/2 · 38 TOTAL
  1. TOOL · CL_212075 ·

    New method optimizes LLM evaluation panels for efficiency and accuracy

    A new research paper proposes a method for optimizing the selection and deployment of Large Language Model (LLM) evaluation panels. The approach formulates judge-panel design as a role-conditioned allocation problem, es…

  2. TOOL · CL_210217 ·

    New framework models LLM delegation contracts and technology choice

    Researchers have developed a new framework to model the Principal-Agent problem when agents select from various technologies, each with different cost-capability profiles. This model is particularly relevant for Large L…

  3. COMMENTARY · CL_198847 ·

    Financial Times flags poor numeracy as AI blind spot

    The Financial Times highlights a critical gap in AI development: poor numeracy. This deficiency in mathematical understanding poses a significant blind spot, potentially hindering the effective and responsible advanceme…

  4. COMMENTARY · CL_195428 ·

    Ethan Mollick: LLMs are revolutionizing science by synthesizing diverse ideas

    Ethan Mollick argues that the revolutionary impact of Large Language Models (LLMs) on science, particularly in mathematics, is already profound. He posits that LLMs excel at synthesizing diverse ideas from different sci…

  5. TOOL · CL_156456 ·

    New distillation method trains LLMs efficiently with soft prompts

    Researchers have developed a new method called Multi-Task On-Policy Distillation via Soft-Prompt Privileged Context ("method") to train large language models. This technique uses a teacher model that differs from the st…

  6. RESEARCH · CL_154170 ·

    LLMs inherit narrative patterns and commit to answers before reasoning, new papers reveal

    Recent research papers explore how narrative structures and experiential abstractions influence Large Language Model (LLM) behavior. One study suggests LLMs inherit narrative patterns from training data, leading to pote…

  7. TOOL · CL_154133 ·

    New SOS-LoRA method boosts LLM performance on reasoning and math tasks

    Researchers have introduced SOS-LoRA, a novel parameter-efficient fine-tuning method designed to enhance the performance of large language models. This technique decomposes the total rank across multiple low-rank expert…

  8. TOOL · CL_145837 ·

    New MASPRM model optimizes multi-agent AI systems without human step-level input

    Researchers have developed a Multi-Agent System Process Reward Model (MASPRM) designed to optimize compute usage in multi-agent systems. This model scores intermediate messages between agents to identify progress, actin…

  9. RESEARCH · CL_143331 ·

    Research: Training duration impacts LLM merging effectiveness

    A new research paper explores the impact of expert training duration on the effectiveness of merging multiple expert models into a single, more capable large language model. The study challenges the standard practice of…

  10. COMMENTARY · CL_138645 ·

    AI writes math papers, humans struggle to understand

    AI models are now capable of writing mathematical papers, a development that has left human researchers struggling to comprehend the generated content. This advancement raises concerns about the effectiveness of current…

  11. RESEARCH · CL_139219 ·

    KV-PRM paper introduces efficient reward modeling for multi-agent LLMs

    Researchers have introduced KV-PRM, a novel method for improving the efficiency of Process Reward Models (PRMs) used in multi-agent systems. Unlike existing text-based PRMs that re-encode entire trajectories, KV-PRM dir…

  12. TOOL · CL_135358 ·

    New 'Representation-as-a-Judge' method uses small models for evaluation

    Researchers have proposed a new evaluation method for language models called Representation-as-a-Judge, which utilizes the internal representations of smaller models rather than their generative output. This approach is…

  13. RESEARCH · CL_131228 ·

    DeepSeek V4 Pro challenges GPT-5 and Claude 4 on benchmarks, offering superior value · 2 sources tracked

    New benchmarks from mid-2026 indicate that Chinese LLM providers, particularly DeepSeek, are now competitive with or surpassing top-tier models from OpenAI and Anthropic in performance and cost-effectiveness. DeepSeek V…

  14. COMMENTARY · CL_113119 ·

    AI's growing math prowess prompts reevaluation of mathematicians' roles

    The increasing capability of artificial intelligence in performing mathematical tasks is prompting a reevaluation of the role of mathematicians. As AI systems become more adept at solving complex problems, human mathema…

  15. COMMENTARY · CL_104324 ·

    AI era prompts debate on the future of mathematics in research

    Researchers are questioning the future necessity of traditional mathematics in the age of advanced AI. As AI models become increasingly capable of performing complex calculations and problem-solving, some academics are …

  16. TOOL · CL_100107 ·

    AI math reasoning benchmarks have a 'sampling blind spot', study finds

    A new research paper published on arXiv explores a critical limitation in evaluating the difficulty of math reasoning problems for AI models. The study reveals that standard benchmarks, which rely on the success rate of…

  17. TOOL · CL_99348 ·

    Nate Soares introduces Gaussian Natural Latents research direction

    Nate Soares has introduced a new research direction called Gaussian Natural Latents, aiming to develop a rigorous theory of concepts and abstraction. This approach leverages Gaussian distributions as a simplified model …

  18. TOOL · CL_96181 ·

    New EngTrace benchmark tests LLMs on verifiable engineering reasoning

    Researchers have introduced EngTrace, a new symbolic benchmark designed to rigorously evaluate the engineering reasoning capabilities of large language models (LLMs). Unlike existing benchmarks that focus on isolated sk…

  19. RESEARCH · CL_89191 ·

    HRM-Text: 1B parameter model with novel architecture challenges LLM paradigms

    A new language model called HRM-Text, developed by Sapient Intelligence, is gaining attention for its innovative architecture that focuses on internal reasoning rather than simply increasing model size or training data.…

  20. COMMENTARY · CL_73169 ·

    AI's Impact on Math and IPOs Explored on Hard Fork

    The latest Hard Fork podcast episode delves into the potential for a "hot IPO summer" driven by AI companies, exploring how the burgeoning field of artificial intelligence is impacting traditional mathematics. The discu…