Researchers have introduced FlowInOne, a novel framework that unifies multimodal generation into a single visual flow-matching model. This approach converts all inputs, including text, into visual prompts, creating an image-in, image-out pipeline. FlowInOne aims to eliminate cross-modal alignment issues and task-specific architectures, handling tasks like text-to-image generation and visual instruction following. The framework is supported by VisPrompt-5M, a dataset of 5 million visual prompt pairs, and VP-Bench, a benchmark for evaluating instruction faithfulness and visual realism. Experiments show FlowInOne achieves state-of-the-art performance among open-source models and is competitive with leading commercial systems. AI
IMPACT Establishes a new foundation for vision-centric generative modeling by unifying diverse generation tasks.
RANK_REASON The cluster contains an academic paper detailing a new model and dataset. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →