DeepSeek has released the weights and reference code for its V4 multimodal model, allowing researchers to examine its visual processing capabilities. Unlike simple image-to-text additions, V4 integrates visual tokens directly into its existing architecture, enabling them to participate in attention, Mixture-of-Experts (MoE) routing, and agent reasoning. The model employs a Vision Transformer (ViT) for initial encoding, followed by an Aligner to compress and map visual features into the language model's space. Specialized mechanisms within V4 handle visual tokens, including modified attention rules and distinct MoE routing biases, to better preserve spatial relationships and computational needs of images within the long-context framework. AI
IMPACT Enables deeper research into multimodal integration and agent capabilities.
RANK_REASON Frontier-lab model release with weights and reference code. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →