PulseAugur
EN
LIVE 12:47:08

MOSS-VL family of open vision-language models introduced for real-time interaction

Researchers have introduced MOSS-VL, a novel family of open vision-language models designed for real-time interaction. The model's architecture allows it to perceive visual input while speaking, utilizing gated cross-attention and a synthesized interaction corpus for training. MOSS-VL-Realtime demonstrates strong performance on streaming benchmarks, particularly in proactive behavior, outperforming existing open-source models on metrics like OmniMMI Proactive Alerting. Despite its 11.3B parameters, MOSS-VL achieves a significant time-to-first-token advantage over comparable models like Qwen3 VL 8B, especially as visual context increases. AI

IMPACT This research introduces a new architecture for real-time vision-language interaction, potentially improving agent responsiveness and proactive capabilities in multimodal AI systems.

RANK_REASON The cluster describes a new research paper detailing a novel model family. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

MOSS-VL family of open vision-language models introduced for real-time interaction

COVERAGE [2]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    MOSS-VL Technical Report

    MOSS-VL is an open vision-language model family enabling real-time interaction by attending to vision via gated cross-attention during generation, using a synthesized interaction corpus and staged curriculum to achieve strong streaming performance with reduced time-to-first-token…

  2. arXiv cs.CV TIER_1 English(EN) · Pengyu Wang, Chenkun Tan, Shaojun Zhou, Qirui Zhou, Yanxin Chen, Xingyang He, Huazheng Zeng, Jijun Cheng, Chenghao Wang, Xiaomeng Qian, Pengfei Wang, Zhan Huang, Shanqing Gao, Wei Huang, Longjun Cao, Wu Ran, Jie Liu, Changtai Zhu, Hongkai Wang, Yixian Ti… ·

    MOSS-VL Technical Report

    arXiv:2608.15045v1 Announce Type: new Abstract: We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across the stack: the language decoder attends to vision only…