研究人员开发了两种新方法来提高视觉语言模型(VLM)的效率和准确性。一种方法SCOPD使用on-policy自蒸馏来帮助VLM适应剪枝的视觉上下文,即使在视觉信息减少的情况下也能显著提高性能。另一种方法VQS通过使用程序验证的QA生成系统来解决训练自演化VLM的挑战,该系统将图像解析为结构化记录,从而获得更准确的答案和更好的模型性能。
AI
arXiv:2609.38426v1 Announce Type: cross Abstract: We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through sh…
arXiv:2609.39579v1 Announce Type: new Abstract: Deploying vision-and-language navigation (VLN) agents in unseen environments remains challenging because unfamiliar layouts and visual conditions can cause execution to go off track. Rather than relying on continuous human supervisi…
arXiv cs.AI
TIER_1English(EN)·Rafi Ibn Sultan, Xiangyu Zhou, Md. Sajid Alam Chowdhury, Chengyin Li, Prashant Khanduri, Marco Brocanelli, Dongxiao Zhu·
arXiv:2609.38716v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spat…
arXiv:2603.26741v2 Announce Type: replace-cross Abstract: Goal-conditioned visual navigation has been a long-standing testbed for embodied AI. We study a natural language-conditioned variant, language-conditioned visual navigation (LCVN), in which an embodied agent must follow a …
arXiv:2609.39920v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap…
arXiv:2609.37515v1 Announce Type: cross Abstract: Thorough evaluation of vision-language models (VLMs) has become prohibitively expensive, as benchmarks span an ever-broader spectrum of capabilities and new models arrive at a relentless pace. Benchmark compression methods that pr…
arXiv:2609.37591v1 Announce Type: cross Abstract: Test-time adaptation for vision-language navigation (TTA-VLN) enables pretrained policies to adapt online to unseen environments using only test-time observations and interaction history. However, distribution shifts can distort l…
arXiv:2609.32352v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging remain insufficiently characterized. Existing ophthalmic datasets are typically de…
Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of ta…
Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24\% of majority-vote l…
arXiv:2609.38777v1 Announce Type: new Abstract: A central goal of vision-language model (VLM) distillation is to transfer both the teacher's language capabilities and its visual understanding. However, existing methods primarily supervise the student's output, leaving visual unde…
arXiv:2609.39492v1 Announce Type: new Abstract: Matching the same vehicle across front and rear cameras is difficult because the cameras do not share a view and the vehicle's appearance changes substantially. We introduce Front2Back-ReID, a benchmark of 500 manually verified vehi…
arXiv:2609.39563v1 Announce Type: new Abstract: Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes between sampled frames. Codec-aware front-ends read the motion vectors and residuals th…
arXiv:2609.40362v1 Announce Type: new Abstract: We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and quantized images as discrete tokens or combine discrete language prediction with c…
arXiv cs.CV
TIER_1English(EN)·Shuming Liu, Zhifang Zhang, Suqin Yuan, Khin Mi Mi Aung, Zhuoyi Lin, Lei Feng·
arXiv:2609.36680v1 Announce Type: new Abstract: Visual reprogramming adapts pretrained models to downstream tasks by modifying their input and output interfaces while keeping the backbone fixed. In vision-language models, existing methods mainly rely on intra-class prompt aggrega…