Two new research papers explore the internal workings of unified multimodal models (UMMs), questioning whether they truly share a unified semantic space across understanding and generation. The first paper, "Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering," introduces a method to probe UMMs by transferring semantic directions between their understanding and generation branches, finding that understanding-derived semantics are more transferable. The second paper, "See2Think: Do Multimodal Models Really Use Intermediate Visual States?" presents a new benchmark and evaluation framework, See2ThinkBench, to assess how UMMs utilize intermediate visual states during reasoning, revealing that their reliance on these states is highly dependent on the model and environment, with rendering being a key bottleneck. AI
IMPACT These studies offer new tools and insights for understanding the internal representations and reasoning processes of multimodal AI, potentially guiding future model development.
RANK_REASON Two academic papers published on arXiv presenting new methods and benchmarks for analyzing multimodal models.
- arXiv
- cross-branch semantic steering
- See2Think
- See2ThinkBench
- Unified Multimodal Models
- Visual Action-of-Thought
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →