The MOSS-VL model family has been introduced as an open vision-language model designed for real-time interaction. It achieves this by incorporating gated cross-attention, allowing the model to process visual information while generating text. The models are trained using a synthesized interaction corpus and a staged curriculum, focusing on proactive behavior and reducing time-to-first-token latency. MOSS-VL-Realtime demonstrates strong performance on streaming benchmarks, outperforming other open-source models in proactive behavior tests and showing a significant time-to-first-token advantage over comparable models like Qwen3 VL 8B. AI
IMPACT This release provides a new open-source option for real-time vision-language tasks, potentially accelerating research and development in interactive AI systems.
RANK_REASON The cluster describes the release of a new open-source vision-language model family with technical details and performance benchmarks, published on Hugging Face and arXiv.
- arXiv
- Hugging Face
- MOSS-VL
- MOSS-VL-Instruct
- MOSS-VL-Realtime
- OmniMMI Proactive Alerting
- Qwen3 VL 8B
- OpenMOSS
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →