New methods boost vision-language model efficiency and accuracy
ByPulseAugur Editorial·[18 sources]·
Researchers have developed two novel approaches to improve the efficiency and accuracy of vision-language models (VLMs). One method, SCOPD, uses on-policy self-distillation to help VLMs adapt to pruned visual contexts, significantly improving performance even with reduced visual information. The other, VQS, addresses the challenge of training self-evolving VLMs by using a program-verified QA generation system that parses images into structured records, leading to more accurate answers and better model performance.
AI
IMPACT
These advancements offer more efficient and accurate ways to train and utilize vision-language models, potentially accelerating their application in various domains.
RANK_REASON
Two distinct research papers introducing new methods for vision-language models.
arXiv:2609.39920v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap…
arXiv:2603.26741v2 Announce Type: replace-cross Abstract: Goal-conditioned visual navigation has been a long-standing testbed for embodied AI. We study a natural language-conditioned variant, language-conditioned visual navigation (LCVN), in which an embodied agent must follow a …
arXiv:2609.39579v1 Announce Type: new Abstract: Deploying vision-and-language navigation (VLN) agents in unseen environments remains challenging because unfamiliar layouts and visual conditions can cause execution to go off track. Rather than relying on continuous human supervisi…
arXiv cs.AI
TIER_1English(EN)·Rafi Ibn Sultan, Xiangyu Zhou, Md. Sajid Alam Chowdhury, Chengyin Li, Prashant Khanduri, Marco Brocanelli, Dongxiao Zhu·
arXiv:2609.38716v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spat…
arXiv:2609.38426v1 Announce Type: cross Abstract: We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through sh…
arXiv:2609.37515v1 Announce Type: cross Abstract: Thorough evaluation of vision-language models (VLMs) has become prohibitively expensive, as benchmarks span an ever-broader spectrum of capabilities and new models arrive at a relentless pace. Benchmark compression methods that pr…
arXiv:2609.37591v1 Announce Type: cross Abstract: Test-time adaptation for vision-language navigation (TTA-VLN) enables pretrained policies to adapt online to unseen environments using only test-time observations and interaction history. However, distribution shifts can distort l…
arXiv:2609.32352v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging remain insufficiently characterized. Existing ophthalmic datasets are typically de…
We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through…
Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of ta…
Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24\% of majority-vote l…
arXiv:2609.39492v1 Announce Type: new Abstract: Matching the same vehicle across front and rear cameras is difficult because the cameras do not share a view and the vehicle's appearance changes substantially. We introduce Front2Back-ReID, a benchmark of 500 manually verified vehi…
arXiv:2609.38777v1 Announce Type: new Abstract: A central goal of vision-language model (VLM) distillation is to transfer both the teacher's language capabilities and its visual understanding. However, existing methods primarily supervise the student's output, leaving visual unde…
arXiv:2609.39563v1 Announce Type: new Abstract: Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes between sampled frames. Codec-aware front-ends read the motion vectors and residuals th…
arXiv:2609.40362v1 Announce Type: new Abstract: We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and quantized images as discrete tokens or combine discrete language prediction with c…
arXiv cs.CV
TIER_1English(EN)·Shuming Liu, Zhifang Zhang, Suqin Yuan, Khin Mi Mi Aung, Zhuoyi Lin, Lei Feng·
arXiv:2609.36680v1 Announce Type: new Abstract: Visual reprogramming adapts pretrained models to downstream tasks by modifying their input and output interfaces while keeping the backbone fixed. In vision-language models, existing methods mainly rely on intra-class prompt aggrega…