HumanEval
PulseAugur coverage of HumanEval — every cluster mentioning HumanEval across labs, papers, and developer communities, ranked by signal.
- used by GSM8K 90%
- instance of large-language models 90%
- developed CatalyzeX 90%
- developed alphaXiv 90%
- instance of GSM8K 70%
- used by alphaXiv 70%
- instance of DagsHub 70%
- instance of ScienceCast 70%
- instance of Gotit.pub 70%
- used by large-language models 70%
- used by ScienceCast 70%
- used by DagsHub 70%
5 day(s) with sentiment data
-
Ollama tag for Qwen2.5-Coder 3B model fails to generate code
A user encountered an issue where a specific Ollama tag for the Qwen2.5-Coder 3B model, quantized at q3_K_M, failed to generate functional code. Despite the model loading and processing requests normally, it produced no…
-
Mistral Small 3.2 released with enhanced function calling and 128K context
Mistral AI has released Mistral Small 3.2, an open-weight model featuring improved function calling and a 128K context window. This update enhances the model's ability to handle tool distinctions and provides cleaner JS…
-
OpenAI's GPT OSS 20B leads in speed and coding benchmarks on Mac
A comparison of three open-weight LLMs—OpenAI's GPT OSS 20B, Alibaba's Qwen3 14B, and Mistral AI's Mistral-Small 24B—was conducted on an Apple M2 machine with 24GB of RAM. GPT OSS 20B emerged as the fastest, outperformi…
-
LLM reasoning exhibits irrationality beyond value alignment, study finds
A new research paper from arXiv explores the concept of "rational value risk" in large language models, suggesting that even well-aligned models can exhibit irrationality during reasoning. This risk is quantified as a d…
-
New PRACT-120 benchmark aims to evaluate AI chatbots holistically
A new benchmark called PRACT-120 has been proposed to evaluate AI chatbots more comprehensively than existing tests like MMLU or GPQA. The benchmark aims to assess not just the core model's capabilities but also the int…
-
New DARTS technique improves decoder LLM merging with entropy-weighted loss
Researchers have developed a new technique called DARTS (Decoder-Aware Representation Tuning via Surgery) to improve model merging for decoder-based large language models. Unlike previous methods for encoder models, DAR…
-
Research details serving challenges for faster diffusion language models
A new research paper on arXiv explores the challenges of serving masked diffusion language models (dLLMs), which can generate text faster than traditional autoregressive models by denoising multiple tokens simultaneousl…
-
AI benchmark flawed by token limits, corrected results show models struggle with new tasks
A benchmark designed to evaluate LLM routing capabilities, named OmnisBench, was found to have a flaw where its output token limit inadvertently penalized models for taking too long to reason. Initially, the benchmark r…
-
New LoRA-GA^2 Algorithm Enhances Large Model Fine-Tuning
Researchers have introduced LoRA-GA^2, a novel fine-tuning algorithm designed to improve upon existing Low-Rank Adaptation (LoRA) methods for large models. This new approach utilizes multi-step gradient information, whi…
-
New method automatically evolves AI evaluation metrics
Researchers have developed a novel method called EvalCEGAR to automatically generate evaluation metrics for AI agents, particularly for tasks like report generation where human scoring is difficult. This approach uses a…
-
Self-correction methods fail to improve LLM code generation without verification
A new study on arXiv investigates the effectiveness of self-correction methods for large language models (LLMs) in code generation. Researchers found that while some uncertainty estimation techniques correlate weakly wi…
-
Anthropic's Claude 3.5 Sonnet overtakes Opus in benchmarks, becomes default for builders
Anthropic has released Claude 3.5 Sonnet, a new mid-tier model that surpasses its previous flagship, Claude 3 Opus, in reasoning and coding benchmarks. This release offers a significant improvement in the cost-performan…
-
Qwen3.8-27B model shows strong performance across multiple hardware setups · 4 sources tracked
Users are reporting impressive performance and capabilities with the Qwen3.8-27B model across various hardware configurations. One user achieved a 262K context window on a single RTX 5090 using vLLM, demonstrating funct…
-
ISO-grounded NFRs improve LLM code quality but not always correctness
A new research paper explores how to improve Large Language Model (LLM) code generation by grounding Non-Functional Requirements (NFRs) in the ISO/IEC 25010 Quality Model. The study found that using either rich natural …
-
New system tackles LLM behavioral relapse in dialogues
Researchers have developed a new system called \"sysname\" to address the issue of behavioral relapse in large language models (LLMs) during multi-turn dialogues. This relapse occurs when LLMs continue to adhere to with…
-
LLM-generated code fails to meet developer intent, study finds
A new study published on arXiv introduces DevIntent, a benchmark designed to measure how often Large Language Models (LLMs) generate code that violates implicit developer intentions. The research found that both Claude …
-
New method retrofits linear attention to speed up diffusion language models
Researchers have developed a method to retrofit linear attention into diffusion language models (dLLMs) to accelerate inference. This new approach, called block-hybrid attention, combines exact softmax attention within …
-
LLM performance varies; task-specific capabilities matter more than rankings
A recent experiment revealed that the performance of large language models can vary significantly even when using the same tasks and parameters, challenging the notion of a single "best" model. Across two runs on 164 Hu…
-
New ABC-GRPO algorithm enhances LLM training stability and performance
Researchers have introduced All-Quadrant Bounded Clipping GRPO (ABC-GRPO), a novel algorithm designed to improve the stability and generalizability of reinforcement learning for large language models. ABC-GRPO addresses…
-
Qwen3 235B leads agentic benchmark, highlighting tool-use differences
The Agentic Index, a benchmark for multi-step task completion involving tool use and error recovery, shows a significant divergence from traditional chat leaderboards. Qwen3 235B, a Mixture-of-Experts model, has achieve…