PulseAugur
EN
LIVE 05:53:21

FlowInOne unifies multimodal generation into a single visual model

Researchers have introduced FlowInOne, a novel framework that unifies multimodal generation into a single visual flow-matching model. This approach converts all inputs, including text, into visual prompts, creating an image-in, image-out pipeline. FlowInOne aims to eliminate cross-modal alignment issues and task-specific architectures, handling tasks like text-to-image generation and visual instruction following. The framework is supported by VisPrompt-5M, a dataset of 5 million visual prompt pairs, and VP-Bench, a benchmark for evaluating instruction faithfulness and visual realism. Experiments show FlowInOne achieves state-of-the-art performance among open-source models and is competitive with leading commercial systems. AI

IMPACT Establishes a new foundation for vision-centric generative modeling by unifying diverse generation tasks.

RANK_REASON The cluster contains an academic paper detailing a new model and dataset. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

FlowInOne unifies multimodal generation into a single visual model

COVERAGE [1]

  1. arXiv cs.CV TIER_1 English(EN) · Junchao Yi, Rui Zhao, Jiahao Tang, Weixian Lei, Linjie Li, Qisheng Su, Zhengyuan Yang, Lijuan Wang, Xiaofeng Zhu, Alex Jinpeng Wang ·

    FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching

    arXiv:2604.06757v3 Announce Type: replace Abstract: Multimodal generation has long been dominated by text-driven pipelines where language dictates vision but cannot reason or create within it. We challenge this paradigm by asking whether all modalities, including textual descript…