WinoGrande
PulseAugur coverage of WinoGrande — every cluster mentioning WinoGrande across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
New FPO method adapts LLMs without backward pass, boosting throughput
Researchers have developed a new method called Forward-Pass-Only (FPO) training that adapts large language models without requiring a backward pass through the model's layers. This technique achieves significantly highe…
-
AMD releases open Instella-MoE-16B LLM with 2.8B active parameters
AMD has released Instella-MoE-16B-A3B, an open-source Mixture-of-Experts language model. This model features 16 billion total parameters but only activates 2.8 billion per token, utilizing architectural innovations like…
-
New UNIFUSION method adapts autoregressive models to discrete diffusion
Researchers have developed UNIFUSION, a novel method to adapt autoregressive language models into discrete diffusion models. This approach unifies existing diffusion objectives under a single generalized Kullback--Leibl…
-
New Benchmark Suite Evaluates LLMs on Kyrgyz Language Understanding
Researchers have developed KyrgyzLLM-Bench, a new benchmark suite designed to evaluate large language models (LLMs) on the Kyrgyz language. This suite includes natively authored datasets like KyrgyzMMLU and KyrgyzRC, al…
-
Prism Transformer introduces progressive head schedules for hierarchical attention
Researchers have introduced the Prism Transformer, a novel architecture that modifies the standard multi-head attention mechanism. Instead of allocating equal dimensional space to each attention head at every layer, Pri…