PulseAugur
中
实时 06:23:34
English(EN) Understanding Sparse Attention Selectivity in Long-Context Foundation Models via Counterfactual Evaluation

新研究致力于稀疏注意力,以实现高效的长上下文大语言模型 · 跟踪 6 个来源

2026 年 8 月发布的多个研究论文探讨了用于大语言模型稀疏注意力机制的新方法,旨在提高效率和长上下文建模能力。这些研究引入了诸如学习到的 Tsallis 指数、具有流式处理能力的在线高效稀疏注意力以及输入自适应稀疏引擎等技术。目标是降低自注意力的二次计算复杂度,使模型能够更有效地处理更长的上下文并加速推理时间,尤其是在视频扩散模型和通用语言任务方面。 AI

影响 稀疏注意力方面的这些进步旨在显著降低计算成本,从而实现大语言模型更高效的训练和推理,尤其是在处理扩展上下文方面。

排序理由 在 arXiv 上发表了多篇研究论文,详细介绍了 LLM 中稀疏注意力的新方法。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 7 个来源。 我们如何撰写摘要 →

新研究致力于稀疏注意力,以实现高效的长上下文大语言模型 · 跟踪 6 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
在 arXiv 上发表了多篇研究论文,详细介绍了 LLM 中稀疏注意力的新方法。
Source corroboration
7 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
59 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [7]

  1. arXiv cs.AI TIER_1 English(EN) · Kleyton da Costa, Bernardo Modenesi ·

    何时应使图注意力稀疏?学习每条边的 Tsallis 指数

    arXiv:2608.02938v1 Announce Type: cross Abstract: Graph attention normalizes neighborhood scores with softmax, the maximum-entropy choice under Shannon statistics. But homophilic and heterophilic graphs want different attention shapes, and one fixed normalization cannot serve bot…

  2. arXiv cs.AI TIER_1 English(EN) · Pike D. Liu, Chang Liu, Yanxuan Yu ·

    $\pi$-Attention:用于长上下文建模的在线高效稀疏Transformer

    arXiv:2511.10696v3 Announce Type: replace-cross Abstract: Sparse attention is crucial in long-context Transformers, which restricts each token to a limited neighborhood and thereby reduces the quadratic cost of full self-attention. Local windows capture nearby context effectively…

  3. arXiv cs.CL TIER_1 English(EN) · Wen Zan, Jiaqi Zhang, Jianchao Tan, Hong Liu, Cunguang Wang, Xiang Li, Duyue Ma, Guanyu Wu, Yifan Lu, Fengcun Li, Yerui Sun, Peng Pei, Yuchen Xie, Xunliang Cai ·

    LongCat 稀疏注意力:通过流感知分层跨层索引驯服闪电

    arXiv:2608.01662v1 Announce Type: cross Abstract: DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remains constrained by the indexer's expensive $O(L^2)$ scoring overhead and the hardware-ineffic…

  4. arXiv cs.CL TIER_1 English(EN) · Xingyu Ren, Youran Sun, Chugang Yi, Haizhao Yang ·

    通过反事实评估理解长上下文基础模型中的稀疏注意力选择性

    arXiv:2608.01676v1 Announce Type: new Abstract: Sparse attention is widely deployed in long-context serving stacks, yet no framework audits how discarding blocks changes the influence of specific content on model output. We first establish that the phenomenon is real and causal: …

  5. arXiv cs.AI TIER_1 English(EN) · Lin Niu, Xin Luo, Linchuan Xie, Yifu Sun, Guanghua Yu, Jianchen Zhu, S Kevin Zhou ·

    Stem:稀疏注意力中的因果信息流重思

    arXiv:2603.06274v2 Announce Type: replace-cross Abstract: The quadratic computational complexity of self-attention remains a fundamental bottleneck for scaling Large Language Models (LLMs) to long contexts, particularly during the pre-filling phase. In this paper, we rethink the …

  6. arXiv cs.CV TIER_1 English(EN) · Shanghao Liu (Eric), Renze Chen (Eric), Size Zheng (Eric), Yuanqiang Liu (Eric), Yun (Eric), Liang, Hailong Yang ·

    SPADE:面向快速视频扩散模型推理的输入自适应稀疏注意力引擎

    arXiv:2608.03335v1 Announce Type: new Abstract: Video diffusion transformers (vDiTs) generate high quality but pay quadratic self-attention cost, making inference prohibitive at video-token scales. The challenge is input-adaptive sparsity: selecting critical Q/K/V tokens with neg…

  7. r/MachineLearning TIER_1 English(EN) · /u/dttdrv ·

    Monodratic:稀疏因果注意力学习的product-hash路由 [R]

    <!-- SC_OFF --><div class="md"><p>Hi everyone,</p> <p>I'm an independent researcher sharing Monodratic, a sparse causal-attention architecture with learned product-hash routing.</p> <p>The idea is that after RoPE, source blocks are assigned to bounded causal posting lists, while …