PulseAugur
实时 15:20:49
English(EN) Transferability Between Understanding and Generation in Unified Multimodal Models

SenseNova-Vision 将计算机视觉任务统一为多模态生成 · 跟踪 6 个来源

研究人员开发了 SenseNova-Vision,一个统一的多模态模型,它将所有计算机视觉任务视为生成问题。这种方法使用自然语言指令和视觉提示来指定任务,允许模型生成文本、图像或两者的组合。该模型在新创建的 SenseNova-Vision Corpus 上进行了训练,在检测、分割和姿态估计等广泛的视觉任务上的性能与专用系统相当。这项工作表明,统一的多模态生成是一种可扩展的方法,可以将各种计算机视觉能力集成到通用基础模型中,并且该模型和语料库现已公开提供。 AI

影响 这种统一的计算机视觉方法可以简化视觉能力与通用基础模型的集成。

排序理由 该集群包含多篇详细介绍新计算机视觉方法和新基准的研究论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 9 个来源。 我们如何撰写摘要 →

SenseNova-Vision 将计算机视觉任务统一为多模态生成 · 跟踪 6 个来源

报道来源 [9]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    视觉作为统一的多模态生成

    We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under this formulation, SenseNova-Vision uses natural-l…

  2. arXiv cs.AI TIER_1 English(EN) · Jiwon Kang, Heeji Yoon, Jaewoo Jung, Jaewon Min, Minkyeong Jeon, Biyeon Hwang, Sangwon Jung, Seungryong Kim ·

    统一多模态模型中理解与生成之间的可迁移性

    arXiv:2607.04423v1 Announce Type: cross Abstract: Unified Multimodal Models (UMMs) integrate image understanding and generation within a single architecture, yet how the two tasks interact remains understudied. We investigate $\boldsymbol{\mathsf{transferability}}$ in UMMs: wheth…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    视觉作为统一的多模态生成

    A unified multimodal model formulates computer vision tasks as generation problems using natural language and visual prompts, achieving performance comparable to specialized systems across diverse vision tasks.

  4. Hugging Face Daily Papers TIER_1 English(EN) ·

    统一多模态模型中理解与生成之间的可迁移性

    Unified Multimodal Models (UMMs) integrate image understanding and generation within a single architecture, yet how the two tasks interact remains understudied. We investigate $\boldsymbol{\mathsf{transferability}}$ in UMMs: whether training a capability on one task improves the …

  5. arXiv cs.CV TIER_1 English(EN) · Pengjie Wang, Linger Deng, Zujia Zhang, Shaojie Zhang, Zhenbo Luo, Pei Fu, Jian Luan, Xiang Bai, Yuliang Liu ·

    DeltaV:在统一的大型多模态模型中通过视觉状态更新进行思考

    arXiv:2607.08434v1 Announce Type: new Abstract: Current Unified Large Multimodal Models (ULMMs) support interleaved multimodal reasoning through textual reasoning and intermediate visual states, but typically generate each visual state as a full image. This full-image generation …

  6. arXiv cs.CV TIER_1 English(EN) · Yuliang Liu ·

    DeltaV:统一大型多模态模型中的视觉状态更新思维

    Current Unified Large Multimodal Models (ULMMs) support interleaved multimodal reasoning through textual reasoning and intermediate visual states, but typically generate each visual state as a full image. This full-image generation paradigm introduces substantial visual-token red…

  7. arXiv cs.CV TIER_1 English(EN) · Jonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma, Akash Gupta, Samuel Roberts, Ioana Croitoru, Simion-Vlad Bogolin, Jialu Tang, Florian Langer, Vyas Raina, Vatsal Raina, Hanyi Xiong, Vishaal Udandarao, Jingyi Lu, Shiyang Chen, Sam Purkis, Tianshuo Ya… ·

    ZeroBench:当代大型多模态模型的“不可能”视觉基准

    arXiv:2502.09696v3 Announce Type: replace Abstract: Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals. Despite this, they attain high scores on many popular visual benchma…

  8. arXiv cs.CV TIER_1 English(EN) · Xiaoyang Han, Jianhua Li, Kewang Deng, Zukai Chen, Xuanke Shi, Sihan Wang, Boxuan Li, Linyan Wang, Siyi Xie, Xin You, Jinsheng Quan, Zhongang Cai, Haiwen Diao, Ziwei Liu, Lei Yang, Dahua Lin, Quan Wang ·

    视觉作为统一的多模态生成

    arXiv:2607.06560v1 Announce Type: new Abstract: We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under t…

  9. arXiv cs.CV TIER_1 English(EN) · Quan Wang ·

    视觉作为统一的多模态生成

    We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under this formulation, SenseNova-Vision uses natural-l…