Researchers have developed SenseNova-Vision, a unified multimodal model that treats all computer vision tasks as generation problems. This approach uses natural language instructions and visual prompts to specify tasks, allowing the model to generate text, images, or a combination of both. Trained on the newly created SenseNova-Vision Corpus, the model demonstrates performance comparable to specialized systems across a wide array of vision tasks, including detection, segmentation, and pose estimation. This work suggests that unified multimodal generation is a scalable method for integrating diverse computer vision capabilities into general-purpose foundation models, with the model and corpus now publicly available.
AI
IMPACT
This unified approach to computer vision could streamline the integration of visual capabilities into general-purpose foundation models.
RANK_REASON
The cluster contains multiple research papers detailing a new approach to computer vision and a new benchmark.
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under this formulation, SenseNova-Vision uses natural-l…
arXiv:2607.04423v1 Announce Type: cross Abstract: Unified Multimodal Models (UMMs) integrate image understanding and generation within a single architecture, yet how the two tasks interact remains understudied. We investigate $\boldsymbol{\mathsf{transferability}}$ in UMMs: wheth…
A unified multimodal model formulates computer vision tasks as generation problems using natural language and visual prompts, achieving performance comparable to specialized systems across diverse vision tasks.
Unified Multimodal Models (UMMs) integrate image understanding and generation within a single architecture, yet how the two tasks interact remains understudied. We investigate $\boldsymbol{\mathsf{transferability}}$ in UMMs: whether training a capability on one task improves the …
arXiv:2607.08434v1 Announce Type: new Abstract: Current Unified Large Multimodal Models (ULMMs) support interleaved multimodal reasoning through textual reasoning and intermediate visual states, but typically generate each visual state as a full image. This full-image generation …
Current Unified Large Multimodal Models (ULMMs) support interleaved multimodal reasoning through textual reasoning and intermediate visual states, but typically generate each visual state as a full image. This full-image generation paradigm introduces substantial visual-token red…
arXiv:2502.09696v3 Announce Type: replace Abstract: Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals. Despite this, they attain high scores on many popular visual benchma…
arXiv:2607.06560v1 Announce Type: new Abstract: We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under t…
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under this formulation, SenseNova-Vision uses natural-l…