Researchers have introduced new methods and benchmarks to improve audio-visual joint reasoning in omni-modal large language models. The OmniReasoning project developed OmniReasoningBench, a benchmark and data engine designed to explicitly test the ability of models to reason using both audio and visual information simultaneously. Their proposed Modality-Factored Self-Distillation (MFSD) method, when applied to the OmniReasoning-30B-A3B model, significantly improved performance on audio-visual reasoning tasks. Separately, the OP-CAD framework focuses on enhancing the robustness of these models against environmental noise and competing speech by using on-policy distillation with selective token supervision, demonstrating improved performance under various noise conditions. AI
IMPACT These advancements could lead to more robust and capable omni-modal AI systems, improving their performance in real-world, noisy environments.
RANK_REASON The cluster contains two arXiv papers detailing new methods and benchmarks for audio-visual reasoning in LLMs.
- Hugging Face
- Modality-Factored Self-Distillation
- OmniQA
- OmniReasoning
- OmniReasoning-30B-A3B
- OmniReasoningBench
- OmniReasoning-RL-19K
- OmniReasoning-SFT-112K
- OmniVideoBench
- On-Policy Clean-Audio Distillation
- OP-CAD
- Qwen3-Omni-30B-A3B-Thinking
- Video-MME-v2
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →