Two new research papers propose methods to compress token sequences in omnimodal large language models (OmniLLMs) to reduce inference costs. The first paper, DASH, uses audio cues to dynamically segment sequences and a tri-signal estimator to retain important tokens, achieving higher compression ratios with competitive accuracy on benchmarks like AVUT and VideoMME. The second paper, OmniFocus, employs a query-guided approach that independently estimates importance for video and audio tokens, aiming to mitigate modality bias and maintain alignment. OmniFocus demonstrated strong performance on the Qwen2.5-Omni model family, offering speedups with minimal accuracy loss. AI
IMPACT These methods could significantly reduce the computational cost of running omnimodal LLMs, making them more accessible and efficient for real-world applications.
RANK_REASON Two research papers published on arXiv propose novel methods for compressing token sequences in omnimodal large language models.
- DailyOmni
- OmniFocus
- OmniLLMs
- omni-modal large language models
- Qwen2.5-Omni
- Qwen2.5-Omni-7B
- arXiv
- AVUT
- Tao Huang
- VideoMME
- WorldSense
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →