LLaDA~2.1
PulseAugur coverage of LLaDA~2.1 — every cluster mentioning LLaDA~2.1 across labs, papers, and developer communities, ranked by signal.
-
New method retrofits linear attention to speed up diffusion language models
Researchers have developed a method to retrofit linear attention into diffusion language models (dLLMs) to accelerate inference. This new approach, called block-hybrid attention, combines exact softmax attention within …
-
FlowBlock framework accelerates diffusion LLM decoding with parallel processing
Researchers have developed FlowBlock, a novel training-free framework designed to accelerate the decoding process for diffusion large language models (dLLMs). This method introduces parallel decoding by allowing blocks …
-
New decoding methods boost diffusion language model speed and accuracy
Researchers have developed new methods to accelerate the decoding process for diffusion language models (dLLMs). FlowBlock, presented in one paper, uses a training-free approach with "Gated Wavefront Decoding" and "Hete…
-
Order-agnostic language models show order-dependent likelihoods
Researchers have identified that order-agnostic language models (OALMs) do not perfectly factorize joint distributions, meaning the order in which tokens are revealed can impact the generated likelihood by up to 0.49 na…