PulseAugur
EN
LIVE 14:02:58
中文(ZH) DeepSeek V4 多模态开源,我们把它的视觉链路拆了一遍

DeepSeek V4 multimodal model weights released for inspection

DeepSeek has released the weights and reference code for its V4 multimodal model, allowing researchers to examine its visual processing capabilities. Unlike simple image-to-text additions, V4 integrates visual tokens directly into its existing architecture, enabling them to participate in attention, Mixture-of-Experts (MoE) routing, and agent reasoning. The model employs a Vision Transformer (ViT) for initial encoding, followed by an Aligner to compress and map visual features into the language model's space. Specialized mechanisms within V4 handle visual tokens, including modified attention rules and distinct MoE routing biases, to better preserve spatial relationships and computational needs of images within the long-context framework. AI

IMPACT Enables deeper research into multimodal integration and agent capabilities.

RANK_REASON Frontier-lab model release with weights and reference code. [lever_c_demoted from frontier_release: ic=1 ai=1.0]

Read on 雷峰网 (Leiphone) →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

DeepSeek V4 multimodal model weights released for inspection

How we ranked this

Signal score
29 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
Frontier-lab model release with weights and reference code. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    DeepSeek V4 Multimodal Open Source, We Broke Down Its Vision Chain

    <section style="text-align: center; margin: 0px 16px; line-height: 1.75em; display: block;"><img class="rich_pages wxw-img" src="https://static.leiphone.com/uploads/new/images/20260901/6a96aaf3030b8.jpg?imageMogr2/quality/90" style="width: 100%; display: inline-block; text-align:…