Researchers have introduced ClearText-Video (CTVid), a new large-scale dataset designed to evaluate how multimodal large language models (MLLMs) handle text-centric video understanding under varying quality conditions. CTVid includes over 4,600 videos with more than 1.6 million human-verified scene-text annotations and 220,000 question-answer pairs in both Chinese and English. The dataset features degraded and restored versions of videos to test text-centric video restoration and multi-quality video question-answering, revealing that visual enhancement does not always improve textual fidelity or reasoning performance for MLLMs. AI
IMPACT This dataset will enable researchers to better understand and improve MLLM performance on real-world, low-quality video data.
RANK_REASON The item describes a new dataset and associated research paper published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- ClearText-Video
- CTVid
- English
- MLLMs
- Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
- Multi-Quality VideoQA
- optical character recognition
- Standard Chinese
- Text-Centric Video Restoration
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →