Gated DeltaNet
PulseAugur coverage of Gated DeltaNet — every cluster mentioning Gated DeltaNet across labs, papers, and developer communities, ranked by signal.
5 day(s) with sentiment data
-
Complex KDA enables rotations in linear attention for advanced state tracking
Researchers have detailed a new type of linear attention mechanism called Complex KDA (CKDA), which allows for rotations in memory updates, enabling more sophisticated state tracking than previous linear models. This me…
-
HLA-WM framework boosts video world models with improved long-range memory
Researchers have developed HLA-WM, a novel training-free framework designed to enhance long-horizon video world models by improving memory retention over extended sequences. The system addresses the issue of information…
-
New attention mechanisms boost long-context sequence modeling
Two new papers introduce novel approaches to enhance long-context sequence modeling in recurrent neural networks. The first paper, "SMat-Attention," proposes Structured Matrix Attention, which uses structured causal mas…
-
New AI model AMOR selectively uses attention based on predictive uncertainty
Researchers have introduced AMOR (Adaptive Metacognitive Output Router), a novel hybrid AI architecture that selectively employs attention mechanisms based on predictive uncertainty. This approach augments a recurrent b…
-
Triadic Linear Attention Enhances RNN Long-Context Modeling
Researchers have introduced Triadic Linear Attention, a novel method that enhances the memory state of Recurrent Neural Networks (RNNs) by utilizing a third-order tensor state. This approach allows for an E-fold increas…
-
SpectralShift enhances Gated DeltaNet context windows via spectral reparameterization · 2 sources tracked
Researchers have introduced SpectralShift, a novel method for extending the context window of Gated DeltaNet (GDN) models, which utilize linear attention mechanisms. Unlike previous approaches that focused on continued …
-
New DASC method slashes AI model state compression by 2.63x
Researchers have developed Decay-Aware State Compression (DASC), a novel method to optimize the serving of hybrid linear-attention models. DASC analyzes the retention timescales of different model components, identifyin…
-
Tail-Replay boosts hybrid LLM inference speed by enabling unconstrained prefix reuse
Researchers have introduced Tail-Replay, a novel prefix caching mechanism designed to enhance the efficiency of hybrid large language models. These models combine full-attention and linear-attention layers to manage lon…
-
New DAMP technique slashes LLM memory use and boosts speed
Researchers have developed a novel quantization technique called DAMP (Decay-Aware Mixed-Precision Recurrent-State Quantization) to reduce the memory footprint and improve the speed of large language models that use rec…
-
Qwen3.8-Flash-Next architecture detailed with efficiency and stability gains · 2 sources tracked
Researchers have detailed the architecture of Qwen3.8-Flash-Next, a 125B parameter sparse mixture-of-experts model. This new model demonstrates improved efficiency and stability compared to its predecessor, the 397B-A17…
-
Qwen3.8-Flash-Next-FP8 VLM Features 125B Parameters and Gated DeltaNet
Qwen3.8-Flash-Next-FP8 is a 125 billion parameter VLM that utilizes 6 billion active MoE units and a Gated DeltaNet architecture. This FP8 variant is distributed across 131 safetensors shards and supports advanced funct…
-
Alibaba previews Qwen4 architecture with cost-efficient Qwen3.8-Flash-Next model
Alibaba's Qwen team has released Qwen3.8-Flash-Next, an open-weight multimodal MoE model that previews the architecture for the upcoming Qwen4. This new model boasts significant cost-efficiency, activating only 6B param…
-
Alibaba's Qwen3.8-27B debuts with hybrid attention for efficient long context
Alibaba's Tongyi Lab has released Qwen3.8-27B, a 27.78-billion-parameter multimodal model featuring a novel hybrid attention architecture. This design strategically replaces three out of every four attention layers with…
-
New TreeWY method enhances speculative verification for hybrid AI models
Researchers have developed a new method called TreeWY for speculative verification in gated delta-net hybrid models. This technique eliminates the need for memory-intensive snapshots of recurrent states, instead using a…
-
Alibaba's Qwen3.8-27B integrates vision and language, rivals larger models
Alibaba's Qwen team has released Qwen3.8-27B, a new open-weight model that integrates vision and language capabilities. This model boasts a large context window of 262,144 tokens, extensible to 1 million, and features f…
-
Alibaba releases open-weight Qwen3.8-Max with 2.4T parameters
Alibaba has released Qwen3.8-2.4T-A95B, marking the first open-weight release of a model in its Qwen-Max class. This new model boasts 2.4 trillion total parameters, with 95 billion active parameters per forward pass, ut…
-
New Modular TTT Framework Simplifies Test-Time Training Design
Researchers have introduced Modular TTT, a new framework designed to simplify the creation and analysis of test-time training (TTT) methods. This framework represents the inner learning process as a directed acyclic gra…
-
New TTCD framework enhances long-context language modeling during inference
Researchers have introduced Test-Time Context Distillation (TTCD), a novel framework for long-context language modeling that optimizes parameter updates during inference. Unlike previous methods, TTCD incorporates a sel…
-
Guide to understanding Moonshot AI's Kimi K3 model architecture
A Reddit post outlines a recommended reading order for understanding the Kimi K3 model by Moonshot AI. The suggested sequence begins with foundational papers on linear transformers and gated delta mechanisms, progressin…
-
Kimi Delta Attention Explained: From Quadratic to Linear Variants
This article delves into the Kimi Delta Attention (KDA) mechanism, a sophisticated variant of linear attention. It traces the evolution from quadratic attention to KDA, explaining how KDA addresses the limitations of ea…