Researchers have developed OmniFusion, a novel approach to simultaneous multilingual multimodal translation. This method fuses pretrained multimodal foundation models (MMFMs) with dedicated translation large language models (LLMs) to create an end-to-end system. OmniFusion, built using Omni 2.5-7B as the MMFM and SeedX PPO-7B as the translation LLM, can handle speech-to-text, speech-and-image-to-text, and text-and-image-to-text translation. Experiments show it reduces latency in simultaneous speech translation by one second compared to cascaded pipelines and enhances overall translation quality by effectively utilizing both audio and visual inputs. AI
IMPACT Enables more efficient and context-aware translation by integrating visual and audio data, potentially improving real-time communication across languages.
RANK_REASON The cluster contains a research paper detailing a new model and methodology for multimodal translation. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →