Diffusion LLMs
PulseAugur coverage of Diffusion LLMs — every cluster mentioning Diffusion LLMs across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
LaCache speeds up diffusion LLMs by 1.3x with token caching
A new caching technique called LaCache has been developed to accelerate diffusion large language models. This method caches unchanged tokens during the denoising process, resulting in a 1.3x speedup in standalone infere…
-
New methods accelerate Diffusion LLMs, addressing speed-quality trade-offs · 3 sources tracked
Researchers are developing new methods to accelerate Diffusion Large Language Models (dLLMs), which are computationally intensive due to their sequence length scaling. Two new frameworks, Dynamic-dLLM and Streaming-dLLM…
-
New DSB method optimizes diffusion LLM scheduling for quality and efficiency
Researchers have introduced Dynamic Sliding Block (DSB), a novel scheduling method for diffusion large language models (dLLMs) that aims to improve both generation quality and inference efficiency. Unlike fixed block sc…
-
New DAPD method speeds up Diffusion LLM decoding
Researchers have introduced Dependency-Aware Parallel Decoding (DAPD), a novel method for accelerating the decoding process in Diffusion Large Language Models (dLLMs). DAPD utilizes self-attention to construct a conditi…
-
New fine-tuning method boosts LLM knowledge injection without paraphrasing
Researchers have developed a new fine-tuning method called Diffusion-Inspired Masked Fine-Tuning (DMT) for autoregressive large language models (LLMs). This technique aims to improve the injection of factual knowledge i…