Svenska(SV)OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs
新方法大幅削减OmniLLM Token成本,提高效率和准确性 · 跟踪9个来源
作者PulseAugur 编辑部·[9 个来源]·
研究人员开发了多种新颖的方法来压缩全模态大语言模型(OmniLLMs)中的Token序列,以降低内存和推理成本。这些方法,包括OmniDelta、OmniScope、Progressive Cramming、PCA和ReMo,专注于在音频、视频和文本等模态之间智能地分配和修剪Token。通过解耦模态相关性、利用查询相似性以及识别冗余或分布外Token,这些技术旨在在显著降低计算开销的同时,维持甚至提高准确性。实验表明,GPU内存大幅减少,推理速度显著提高,为OmniLLMs树立了新的效率标杆。
AI
arXiv:2607.25669v1 Announce Type: new Abstract: Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compression methods m…
arXiv:2607.22716v1 Announce Type: cross Abstract: In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbreak attacks and hallucinations. Given that vision an…
Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often breaks down: for the same query, audio and video relevance often peaks at different moments. This cro…
Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compression methods mainly focus on selecting important tokens under …
arXiv:2607.21231v1 Announce Type: new Abstract: Token cramming compresses sequences into learned embeddings with near-perfect reconstruction, but fixed token budgets and 99\% accuracy thresholds leave it unclear whether residual errors reflect optimization failures or fundamental…
Vision-language models (VLMs) process large numbers of visual tokens, resulting in substantial inference latency and memory overhead. This has motivated extensive research on visual token compression. While training-free strategies rely on heuristic metrics and suffer significant…
arXiv:2607.23193v1 Announce Type: new Abstract: Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often breaks down: for the same query, audio and video rele…
arXiv cs.CV
TIER_1English(EN)·Zihan Song, Shuo Ye, Bo Zhao, Ruixin Zhang, Jiayu Zhang, Shouhong Ding, Zitong Yu·
arXiv:2607.22726v1 Announce Type: new Abstract: Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-duration frames remains a hindrance to efficient reasoning. This paper introduces a tr…
arXiv:2607.21179v1 Announce Type: new Abstract: The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at inference time. Omni-LLMs reason jointly over audio, video and text, but the cost of the three streams is highly unbalanced:…