PulseAugur
中
实时 21:37:09
Svenska(SV) OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

新方法大幅削减OmniLLM Token成本,提高效率和准确性 · 跟踪9个来源

研究人员开发了多种新颖的方法来压缩全模态大语言模型(OmniLLMs)中的Token序列,以降低内存和推理成本。这些方法,包括OmniDelta、OmniScope、Progressive Cramming、PCA和ReMo,专注于在音频、视频和文本等模态之间智能地分配和修剪Token。通过解耦模态相关性、利用查询相似性以及识别冗余或分布外Token,这些技术旨在在显著降低计算开销的同时,维持甚至提高准确性。实验表明,GPU内存大幅减少,推理速度显著提高,为OmniLLMs树立了新的效率标杆。 AI

影响 这些压缩技术对于通过降低全模态LLM的计算需求,使其更加实用和易于访问至关重要。

排序理由 多篇研究论文介绍了全模态大语言模型中Token压缩的新颖方法。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 9 个来源。 我们如何撰写摘要 →

新方法大幅削减OmniLLM Token成本,提高效率和准确性 · 跟踪9个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇研究论文介绍了全模态大语言模型中Token压缩的新颖方法。
Source corroboration
9 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
86 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [9]

  1. arXiv cs.AI TIER_1 Svenska(SV) · Haoyang Huang, Wenjie Huang, Tianqi Xu, Hongyaoxing Gu, Kang Tan, Yikai Fu, Yuhao Shen, Tianyu Liu, Baolin Zhang, Jun Zhang, Xinyi Hu, Jun Dai, Shuang Ge, Lei Chen, Yue Li, Mingchen Wang, Meng Zhang ·

    OmniDelta:用于 OmniLLMs 中 Token 压缩的驱动技能的预算分配

    arXiv:2607.25669v1 Announce Type: new Abstract: Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compression methods m…

  2. arXiv cs.LG TIER_1 English(EN) · Shishen Gu, Jiequan Cui, Wenbo Hu, Zenglin Shi, Zhenzhen Hu, Richang Hong ·

    视觉令牌压缩增强了MLLM的鲁棒性

    arXiv:2607.22716v1 Announce Type: cross Abstract: In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbreak attacks and hallucinations. Given that vision an…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    OmniScope:模态解耦的Token压缩技术,用于全模态大语言模型

    Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often breaks down: for the same query, audio and video relevance often peaks at different moments. This cro…

  4. Hugging Face Daily Papers TIER_1 Svenska(SV) ·

    OmniDelta:用于 OmniLLMs 中 token 压缩的技能驱动预算分配

    Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compression methods mainly focus on selecting important tokens under …

  5. arXiv cs.CL TIER_1 English(EN) · Dmitrii Tarasov, Timofei Lashukov, Elizaveta Goncharova, Andrey Kuznetsov ·

    渐进式挤压:可靠的令牌压缩及其揭示的意义

    arXiv:2607.21231v1 Announce Type: new Abstract: Token cramming compresses sequences into learned embeddings with near-perfect reconstruction, but fixed token budgets and 99\% accuracy thresholds leave it unclear whether residual errors reflect optimization failures or fundamental…

  6. Hugging Face Daily Papers TIER_1 English(EN) ·

    VisCo:利用大型语言模型作为内在编码器进行视觉令牌压缩

    Vision-language models (VLMs) process large numbers of visual tokens, resulting in substantial inference latency and memory overhead. This has motivated extensive research on visual token compression. While training-free strategies rely on heuristic metrics and suffer significant…

  7. arXiv cs.CV TIER_1 English(EN) · Jinsen Su, Yongdong Luo, Yuexiao Ma, Yibo Hu, Meiguang Jin, Xiaowu Zheng ·

    OmniScope:模态解耦的Token压缩用于全模态大语言模型

    arXiv:2607.23193v1 Announce Type: new Abstract: Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often breaks down: for the same query, audio and video rele…

  8. arXiv cs.CV TIER_1 English(EN) · Zihan Song, Shuo Ye, Bo Zhao, Ruixin Zhang, Jiayu Zhang, Shouhong Ding, Zitong Yu ·

    PCA:面向快速视频大语言模型的持久性感知压缩与聚合

    arXiv:2607.22726v1 Announce Type: new Abstract: Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-duration frames remains a hindrance to efficient reasoning. This paper introduces a tr…

  9. arXiv cs.CV TIER_1 English(EN) · Suho Yoo, Youngjoon Jang, Hyebin Cho, Joon Son Chung ·

    视而不见,心念不绝:面向 Omni-LLM 的 Token 压缩

    arXiv:2607.21179v1 Announce Type: new Abstract: The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at inference time. Omni-LLMs reason jointly over audio, video and text, but the cost of the three streams is highly unbalanced:…