Researchers have introduced a new framework called Temporal Causal Drive Evaluation to analyze how different information sources influence the generation process of vision-language models (VLMs). This framework uses causal and temporal metrics to track the evolving roles of visual input, question text, and generated prefixes during decoding. Experiments with models like Qwen3-VL-8B-Instruct and InternVL2-8B demonstrated a shift from early reliance on question and visual guidance to increasing dependence on generated prefixes. The proposed causal-drive metrics showed significant improvements in reducing recovery error compared to observational baselines. AI
IMPACT Provides new diagnostic tools for understanding and improving VLM generation processes.
RANK_REASON The cluster contains an academic paper detailing a new evaluation framework for vision-language models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- InternVL2-8B
- LLaVA-Video-178K
- Mavis
- MiraData
- Qwen3-VL-8B-Instruct
- ScienceCast
- Temporal Causal Drive Evaluation Framework
- VLMBias
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →