A new research paper introduces the concept of "two clocks" in diffusion Large Multimodal Models (LMMs), distinguishing between when an answer stabilizes and when its rationale is fully generated. The study analyzes this phenomenon across three visual question-answering benchmarks, finding that a significant portion of the rationale remains unwritten at the point of answer stabilization. Experiments with varying block lengths and prompting strategies, including direct instructions, show impacts on accuracy and coverage, suggesting that coverage is a larger component of prompting differences than conditional accuracy. AI
IMPACT This research offers insights into the internal temporal dynamics of LMMs, potentially guiding future model development for improved reasoning and answer generation.
RANK_REASON The cluster contains a research paper detailing novel findings about the internal workings of diffusion LMMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →