PulseAugur
EN
LIVE 00:45:02
ENTITY Gated DeltaNet

Gated DeltaNet

PulseAugur coverage of Gated DeltaNet — every cluster mentioning Gated DeltaNet across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
33
33 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
28
28 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

5 day(s) with sentiment data

RECENT · PAGE 1/2 · 34 TOTAL
  1. TOOL · CL_283927 ·

    Complex KDA enables rotations in linear attention for advanced state tracking

    Researchers have detailed a new type of linear attention mechanism called Complex KDA (CKDA), which allows for rotations in memory updates, enabling more sophisticated state tracking than previous linear models. This me…

  2. TOOL · CL_283403 ·

    HLA-WM framework boosts video world models with improved long-range memory

    Researchers have developed HLA-WM, a novel training-free framework designed to enhance long-horizon video world models by improving memory retention over extended sequences. The system addresses the issue of information…

  3. RESEARCH · CL_270698 ·

    New attention mechanisms boost long-context sequence modeling

    Two new papers introduce novel approaches to enhance long-context sequence modeling in recurrent neural networks. The first paper, "SMat-Attention," proposes Structured Matrix Attention, which uses structured causal mas…

  4. TOOL · CL_268828 ·

    New AI model AMOR selectively uses attention based on predictive uncertainty

    Researchers have introduced AMOR (Adaptive Metacognitive Output Router), a novel hybrid AI architecture that selectively employs attention mechanisms based on predictive uncertainty. This approach augments a recurrent b…

  5. TOOL · CL_280595 ·

    Triadic Linear Attention Enhances RNN Long-Context Modeling

    Researchers have introduced Triadic Linear Attention, a novel method that enhances the memory state of Recurrent Neural Networks (RNNs) by utilizing a third-order tensor state. This approach allows for an E-fold increas…

  6. RESEARCH · CL_254534 ·

    SpectralShift enhances Gated DeltaNet context windows via spectral reparameterization · 2 sources tracked

    Researchers have introduced SpectralShift, a novel method for extending the context window of Gated DeltaNet (GDN) models, which utilize linear attention mechanisms. Unlike previous approaches that focused on continued …

  7. TOOL · CL_228892 ·

    New DASC method slashes AI model state compression by 2.63x

    Researchers have developed Decay-Aware State Compression (DASC), a novel method to optimize the serving of hybrid linear-attention models. DASC analyzes the retention timescales of different model components, identifyin…

  8. TOOL · CL_228886 ·

    Tail-Replay boosts hybrid LLM inference speed by enabling unconstrained prefix reuse

    Researchers have introduced Tail-Replay, a novel prefix caching mechanism designed to enhance the efficiency of hybrid large language models. These models combine full-attention and linear-attention layers to manage lon…

  9. TOOL · CL_227089 ·

    New DAMP technique slashes LLM memory use and boosts speed

    Researchers have developed a novel quantization technique called DAMP (Decay-Aware Mixed-Precision Recurrent-State Quantization) to reduce the memory footprint and improve the speed of large language models that use rec…

  10. RESEARCH · CL_229109 ·

    Qwen3.8-Flash-Next architecture detailed with efficiency and stability gains · 2 sources tracked

    Researchers have detailed the architecture of Qwen3.8-Flash-Next, a 125B parameter sparse mixture-of-experts model. This new model demonstrates improved efficiency and stability compared to its predecessor, the 397B-A17…

  11. SIGNIFICANT · CL_219953 ·

    Qwen3.8-Flash-Next-FP8 VLM Features 125B Parameters and Gated DeltaNet

    Qwen3.8-Flash-Next-FP8 is a 125 billion parameter VLM that utilizes 6 billion active MoE units and a Gated DeltaNet architecture. This FP8 variant is distributed across 131 safetensors shards and supports advanced funct…

  12. FRONTIER RELEASE · CL_219957 ·

    Alibaba previews Qwen4 architecture with cost-efficient Qwen3.8-Flash-Next model

    Alibaba's Qwen team has released Qwen3.8-Flash-Next, an open-weight multimodal MoE model that previews the architecture for the upcoming Qwen4. This new model boasts significant cost-efficiency, activating only 6B param…

  13. SIGNIFICANT · CL_217028 ·

    Alibaba's Qwen3.8-27B debuts with hybrid attention for efficient long context

    Alibaba's Tongyi Lab has released Qwen3.8-27B, a 27.78-billion-parameter multimodal model featuring a novel hybrid attention architecture. This design strategically replaces three out of every four attention layers with…

  14. TOOL · CL_215893 ·

    New TreeWY method enhances speculative verification for hybrid AI models

    Researchers have developed a new method called TreeWY for speculative verification in gated delta-net hybrid models. This technique eliminates the need for memory-intensive snapshots of recurrent states, instead using a…

  15. SIGNIFICANT · CL_209509 ·

    Alibaba's Qwen3.8-27B integrates vision and language, rivals larger models

    Alibaba's Qwen team has released Qwen3.8-27B, a new open-weight model that integrates vision and language capabilities. This model boasts a large context window of 262,144 tokens, extensible to 1 million, and features f…

  16. SIGNIFICANT · CL_200541 ·

    Alibaba releases open-weight Qwen3.8-Max with 2.4T parameters

    Alibaba has released Qwen3.8-2.4T-A95B, marking the first open-weight release of a model in its Qwen-Max class. This new model boasts 2.4 trillion total parameters, with 95 billion active parameters per forward pass, ut…

  17. RESEARCH · CL_191303 ·

    New Modular TTT Framework Simplifies Test-Time Training Design

    Researchers have introduced Modular TTT, a new framework designed to simplify the creation and analysis of test-time training (TTT) methods. This framework represents the inner learning process as a directed acyclic gra…

  18. TOOL · CL_180518 ·

    New TTCD framework enhances long-context language modeling during inference

    Researchers have introduced Test-Time Context Distillation (TTCD), a novel framework for long-context language modeling that optimizes parameter updates during inference. Unlike previous methods, TTCD incorporates a sel…

  19. COMMENTARY · CL_171584 ·

    Guide to understanding Moonshot AI's Kimi K3 model architecture

    A Reddit post outlines a recommended reading order for understanding the Kimi K3 model by Moonshot AI. The suggested sequence begins with foundational papers on linear transformers and gated delta mechanisms, progressin…

  20. TOOL · CL_168772 ·

    Kimi Delta Attention Explained: From Quadratic to Linear Variants

    This article delves into the Kimi Delta Attention (KDA) mechanism, a sophisticated variant of linear attention. It traces the evolution from quadratic attention to KDA, explaining how KDA addresses the limitations of ea…