English(EN)MindWorldBench: Evaluating Mental-State-to-Behavior Reasoning in Image-to-Video Generation
AI研究在图像/视频生成中推进视觉推理和个性化
作者PulseAugur 编辑部·[7 个来源]·
研究人员正在开发新的基准和模型来提高AI的推理能力,特别是在视觉生成任务中。一篇论文介绍了PEARL,一个用于个性化图像生成的系统,该系统利用用户历史来使图像与生活方式和审美偏好保持一致,在个性化指标上提高了15%。另一项研究提出了MindWorldBench来评估图像到视频生成中的心智状态到行为推理,发现当前模型难以将行为与潜在心智状态对齐,并表现出“全知偏见”。此外,一项关于通过魔方进行推理的视频生成规模化的研究表明,较小的模型可以用更少的计算量实现更高的状态准确性,并且符号监督可以显著提高性能。
AI
arXiv:2610.00737v1 Announce Type: cross Abstract: Personalized image generation has remained narrowly focused on conditional synthesis from curated visual exemplars, rather than capturing who a user is. In practice, however, a user's personal context is much richer, comprising re…
Current image-to-video models achieve visual realism and physical plausibility, but reasoning about mental states remains unexplored. Actions are driven by belief, desire, and perception, requiring inference beyond explicit instructions. We introduce MindWorldBench to evaluate me…
arXiv:2609.36599v1 Announce Type: cross Abstract: We study whether scaling video generation enables models to reason about hidden information from the past frames, and at what computational cost. Our controlled benchmark requires predicting nine prescribed moves of an initially s…
Personalized image generation has remained narrowly focused on conditional synthesis from curated visual exemplars, rather than capturing who a user is. In practice, however, a user's personal context is much richer, comprising reviews, posts, images, captions, and metadata accum…
Chain-of-thought reasoning has revolutionized natural language processing by enabling large language models (LLMs) to decompose problems into intermediate steps before answering. Yet confining reasoning to the textual domain presents limitations for tasks requiring direct manipul…
arXiv cs.CV
TIER_1English(EN)·Zehua Ma, Kun Xiang, Yunshuang Nie, Quanlin Chen, Haoyuan Li, Xiuwei Chen, Jiang Ji, Haijun Wu, Zhenyu Xie, Michael Kampffmeyer, Hanhui Li, Xiaodan Liang·
arXiv:2609.40129v1 Announce Type: new Abstract: Reasoning through video generation offers a promising path toward visual intelligence by modeling latent visual states and their dynamics. However, current video generation models often lack explicit guidance on how these states sho…
arXiv:2609.39147v1 Announce Type: new Abstract: Current image-to-video models achieve visual realism and physical plausibility, but reasoning about mental states remains unexplored. Actions are driven by belief, desire, and perception, requiring inference beyond explicit instruct…