PulseAugur
EN
LIVE 10:15:08

New research enhances vision-language models for medical, retrieval, and robotics tasks

Researchers are developing new methods to improve vision-language models (VLMs) across various domains. One paper introduces CoT-Mediate, a framework to assess how generated reasoning influences VLM predictions in medical contexts, finding that the position of reasoning is more critical than its stated source. Another study presents ReToken, a single learnable embedding that enhances visual retrieval by selecting relevant visual tokens, showing significant gains on benchmarks like Visual Haystacks and LVBench. For robotics, BioVLN is a new simulation platform designed for visual-language navigation in biomedical labs, addressing the need for precise instrument interaction. Additionally, a method called SeGP-CL is proposed for continual learning in VLMs, aiming to preserve cross-modal semantic geometry and combat catastrophic forgetting. Finally, S-GRPO offers a unified post-training framework for large VLMs, integrating supervised fine-tuning and reinforcement learning to improve adaptation and preserve general capabilities. AI

IMPACT These advancements in VLM reasoning, retrieval, robotics navigation, continual learning, and unified post-training frameworks are pushing the boundaries of AI capabilities in specialized and general applications.

RANK_REASON Cluster contains multiple research papers on advancements in vision-language models and related technologies.

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 406 sources. How we write summaries →

New research enhances vision-language models for medical, retrieval, and robotics tasks

COVERAGE [406]

  1. arXiv cs.AI TIER_1 English(EN) · Jieyu Zhang, Ziqi Gao, Luke Zettlemoyer, Ranjay Krishna ·

    Vision-Language Grounding as Bidirectional Concept Correspondence

    arXiv:2608.07886v1 Announce Type: cross Abstract: Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the corresponding…

  2. arXiv cs.AI TIER_1 English(EN) · R\'ois\'in Luo ·

    ZetaGPT: A Reference Implementation of Positional--Encoding--Free State--Space--Attention Language Models

    arXiv:2608.09432v1 Announce Type: cross Abstract: Transformer-based language models rely on self-attention, whose computation is permutation-equivariant and therefore lacks an intrinsic mechanism for representing token order. Existing architectures address this limitation by expl…

  3. arXiv cs.LG TIER_1 English(EN) · Ao Zhou, Zhiwei Jiang, Zifeng Cheng, Cong Wang, Shufan Yang, Haoru Chen, Qing Gu ·

    Dynamic Distribution-Aware Uncertainty Tracking in Vision-Language Representation Learning

    arXiv:2608.09011v1 Announce Type: new Abstract: Uncertainty Quantification (UQ) aims to measure the reliability of model predictions, serving as a critical safeguard for deploying Vision-Language Models (VLMs) in safety-critical scenarios. Post-hoc approaches are widely adopted d…

  4. arXiv cs.CL TIER_1 English(EN) · Zhanna Mukhametsharip (Saarland University, Germany), Vera Demberg (Saarland University, Germany, Max Planck Institute for Informatics, Germany), Varsha Suresh (Max Planck Institute for Informatics, Germany) ·

    PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models

    arXiv:2608.09772v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason about relationships between images and text or rely on superficial correlations…

  5. arXiv cs.AI TIER_1 English(EN) · Gesina Schwalbe, Mert Keser, Moritz Bayerkuhnlein, Edgar Heinert, Annika M\"utze, Marvin Keller, Sparsh Tiwari, Georgii Mikriukov, Diedrich Wolter, Jae Hee Lee, Matthias Rottmann ·

    Explaining, Verifying, and Aligning Semantic Hierarchies in Vision-Language Model Embeddings

    arXiv:2603.26798v2 Announce Type: replace-cross Abstract: Vision-language model (VLM) encoders such as CLIP enable strong retrieval and zero-shot classification in a shared image-text embedding space, yet the semantic organization of this space is rarely inspected. We present a p…

  6. arXiv cs.AI TIER_1 English(EN) · Bhavika Jalli, Nikhil Korati Prasanna, Jayanta Choudhury ·

    A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy

    arXiv:2608.07427v1 Announce Type: new Abstract: LLM inference accounts for over 90% of AI operational energy, scaling directly with input token count---a critical inefficiency for telecom network analytics and numerical time-series data analysis (NTSDA), where raw multivariate KP…

  7. arXiv cs.AI TIER_1 English(EN) · Ying Chen, Weizhen Li, Zhe Hu, Zhenjiang Li, Rui Jiang, Zhifeng Gu, Lihuang Fang, Jiangping Liu, Lei Yi, Jie Chen ·

    Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

    arXiv:2608.06756v1 Announce Type: new Abstract: Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reaso…

  8. arXiv cs.CL TIER_1 English(EN) · Rahul Murali Shankar, Titus von der Malsburg, Sebastian Pad\'o ·

    Gaze Behavior in Visual World Experiments Can be Modeled With Off-the-shelf Language-Vision Encoders

    arXiv:2608.07282v1 Announce Type: new Abstract: The recent advances in neural language models have also spurred much work in computational psycholinguistics, asking whether neural LMs are also promising models of human language processing. However, work has been overwhelmingly fo…

  9. arXiv cs.LG TIER_1 English(EN) · Minseok Kang, Hyunwoo Kim, Chanyoung Kim, Minwoo Kim, Jaekoo Lee, Dahuin Jung ·

    Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models

    arXiv:2608.06901v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal tasks through large-scale pre-training, yet their rapidly increasing computational and memory requirements pose significant challenges…

  10. arXiv cs.AI TIER_1 English(EN) · Nikos Theodoridis, Reenu Mohandas, Ganesh Sistu, Anthony Scanlan, Ciar\'an Eising, Tim Brophy ·

    Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving

    arXiv:2603.06054v2 Announce Type: replace-cross Abstract: The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios. However,…

  11. arXiv cs.LG TIER_1 English(EN) · Rasul Khanbayov, Hasan Kurban ·

    Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading

    arXiv:2608.05675v1 Announce Type: new Abstract: Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change. This has a known blind spot, a systematic misreading survives the perturbation and gets certi…

  12. arXiv cs.AI TIER_1 English(EN) · J. de Curt\`o, Dayani Plasencia, Diego S\'anchez, I. de Zarz\`a ·

    Visual Grounding in Zero-Shot Vision-Language Control

    arXiv:2608.06154v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that decisions are grounded in visual input: simulator dynamics and conservative action priors can p…

  13. arXiv cs.AI TIER_1 English(EN) · Qinwu Xu ·

    Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift

    arXiv:2605.16411v2 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likel…

  14. Hugging Face Daily Papers TIER_1 English(EN) ·

    Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

    Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reasoned about, and verified. Meeting these demands r…

  15. Hugging Face Daily Papers TIER_1 English(EN) ·

    Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading

    Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change. This has a known blind spot, a systematic misreading survives the perturbation and gets certified wrong, which we show is computable, not jus…

  16. arXiv cs.CL TIER_1 English(EN) · Sirun Li, Minghao Liu, Ling Dai, Yong Li, Haoxin Lyu, Junting Zhou, Fan Zhang ·

    SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

    arXiv:2608.04244v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they confli…

  17. arXiv cs.LG TIER_1 English(EN) · Chen Zhong, Xiao An, Zijie Wang, Jiepan Li, Guangyi Yang, Wei He ·

    DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models

    arXiv:2608.04496v1 Announce Type: cross Abstract: Visual inputs in vision-language models (VLMs) are often encoded into substantially longer token sequences than text, making visual tokens a major bottleneck for efficient inference. Abundant recent methods address this bottleneck…

  18. arXiv cs.AI TIER_1 English(EN) · De Jiang, Zhengyang Zhang, Kehong Yuan, Shaohua Ma ·

    CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models

    arXiv:2608.04509v1 Announce Type: new Abstract: Vision-language systems combine images with retrieved text, but these sources can disagree or jointly fail to support an answer. Reliable models must identify the trustworthy source and abstain when neither is adequate. Existing pos…

  19. arXiv cs.AI TIER_1 English(EN) · Houze Xu, Jizhong Li, Ziyi Ye ·

    Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models

    arXiv:2608.04765v1 Announce Type: cross Abstract: Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language understanding, and robotic control. However, existing VLA models still face major challenges in long-horizon tasks: sparse e…

  20. arXiv cs.AI TIER_1 English(EN) · Cheng Yin, Yankai Lin, Wang Xu, Sikyuen Tam, Xiangrui Zeng, Zhiyuan Liu, Zhouping Yin ·

    DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models

    arXiv:2511.15669v3 Announce Type: replace-cross Abstract: Does Chain-of-Thought (CoT) reasoning genuinely improve Vision Language Action (VLA) models, or does it merely add overhead? Existing CoT-VLA systems report limited and inconsistent gains, yet no prior work has rigorously …

  21. arXiv cs.CL TIER_1 English(EN) · Ali Khoramfar, Mohammad Javad Dousti, Alireza Mohamadian, Heshaam Faili ·

    Evaluating the Diagnostic Robustness of Vision-Language Models Under Visual and Textual Perturbations

    arXiv:2608.04885v1 Announce Type: cross Abstract: Standard accuracy metrics for VLMs often mask significant reliability failures in sensitive domains. In this work, we utilize a histopathology-validated brain MRI dataset to systematically assess the diagnostic robustness of four …

  22. arXiv cs.LG TIER_1 English(EN) · Grzegorz Gruszczynski, Pawel Olszowiec, Michal Byra, Grzegorz Stefanski, Alberto Presta ·

    Training Crossroads for Recurrent Vision Transformers: Recurrence, Neural ODEs, and Deep Supervision

    arXiv:2608.04879v1 Announce Type: new Abstract: Vision Transformers (ViTs) achieve strong image-recognition performance, but their parameter count grows linearly with depth when each block is independently parameterized. Single-block recurrent ViTs (bViT) remove this growth by re…

  23. arXiv cs.LG TIER_1 English(EN) · Hongyu Zhang, Cheng Yan, Xiang Xia, Wuyang Zhang ·

    Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models

    arXiv:2608.04454v1 Announce Type: cross Abstract: Mixture-of-experts vision-language models (MoE-VLMs) increase model capacity with sparse expert activation, yet deployment requires storing the full expert pool. Training-free expert merging reduces this burden, and many routing-b…

  24. Hugging Face Daily Papers TIER_1 English(EN) ·

    Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models

    Mixture-of-experts vision-language models (MoE-VLMs) increase model capacity with sparse expert activation, yet deployment requires storing the full expert pool. Training-free expert merging reduces this burden, and many routing-based methods aggregate routing statistics across a…

  25. arXiv cs.AI TIER_1 English(EN) · Jinquan Zhang, Dongfu Yin, Run Yang, Yufeng Yan, Zhen Tian, F. Richard Yu ·

    Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking

    arXiv:2608.03231v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies promise general robotic manipulation, but their robustness against physical-world attacks remains fragile. In particular, we show that physically realizable adversarial patches can reliably in…

  26. arXiv cs.AI TIER_1 English(EN) · Tianbao Jiang, Weicong Ni, Gerard de Melo, Linlin Wang ·

    Aligning Large Vision-Language Models at Test Time: A Trajectory-Guided Structured Sampling Approach

    arXiv:2608.03204v1 Announce Type: cross Abstract: Post-training reinforcement learning (RL) algorithms are commonly used to align large vision-language models (LVLMs) with human intent and the requirements of visual reasoning tasks. However, existing RL-based alignment methods ar…

  27. arXiv cs.CL TIER_1 English(EN) · Jiashu Yao, Haoyu Wen, Siyuan Gao, Yuhang Guo, Zeming Liu, Heyan Huang ·

    HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection

    arXiv:2509.23690v2 Announce Type: replace-cross Abstract: Safety hazards in the home are a leading cause of preventable domestic injuries, motivating an automated inspector that actively explores a home and reports hazards before they cause harm. We introduce HomeSafeBench, the f…

  28. arXiv cs.CL TIER_1 English(EN) · Haz Sameen Shahgir, Xiaofu Chen, Yu Fu, Erfan Shayegani, Nael Abu-Ghazaleh, Yova Kementchedjhieva, Yue Dong ·

    VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors

    arXiv:2604.02486v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have achieved impressive performance across a wide range of multimodal tasks. However, they often fail on tasks that require fine-grained visual perception, even when the required information …

  29. arXiv cs.AI TIER_1 English(EN) · Chunyang Jiang, Pingping Zhang, Yuzhi Zhao, Wenao Ma, Zhijian Hou, Mengyang Wu, Yiyang Cai, Senkang Hu, Sitong Cheng, Chi-Min Chan, Wei Xue, Yike Guo ·

    Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement

    arXiv:2608.03733v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimodal data that are costly to annotate. Self-augmentati…

  30. arXiv cs.AI TIER_1 English(EN) · Mohammad Rostami ·

    In-Context Collapse in Vision-Language Models and How to Mitigate it?

    arXiv:2608.02830v1 Announce Type: cross Abstract: Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely assumed to improve as more demonstrations are supplied. We show the opposite: as d…

  31. arXiv cs.AI TIER_1 English(EN) · Inkyu Sa, Konstantin Stulov, Rajat Bhageria ·

    ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies

    arXiv:2608.02958v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies trained by behavior cloning fail silently: from the action stream alone, a collapsing rollout looks much like one making clean progress, because imitation supplies no notion of progress. Reinf…

  32. arXiv cs.AI TIER_1 English(EN) · Paribesh Regmi, Qingshuang Chen, Chi Zhang, Heba Aly, Yelin Kim, Hongda Mao ·

    Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models

    arXiv:2608.03112v1 Announce Type: cross Abstract: Vision-language models excel at image and video understanding but suffer from high inference latency due to the need to process thousands of tokens per image, limiting their deployment on resource-constrained edge devices and in r…

  33. Hugging Face Daily Papers TIER_1 English(EN) ·

    BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

    Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution shifts, and …

  34. Hugging Face Daily Papers TIER_1 English(EN) ·

    MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization

    Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-time adaptation (TTA) via Low-Rank Adaptation offers a parameter-efficient solution, we identify a fundamental bottleneck in curre…

  35. Hugging Face Daily Papers TIER_1 English(EN) ·

    Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement

    Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimodal data that are costly to annotate. Self-augmentation offers a promising alternative by enabling mo…

  36. Hugging Face Daily Papers TIER_1 English(EN) ·

    When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware

    Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead, shared work, and the operators each policy can avoid. A stage-level decomposition reconciles these components with measured end-…

  37. arXiv cs.CL TIER_1 English(EN) · Jing Wu, Jianhua Wu, Jiayi Guan, Jiahong Chen, Jinghui Lu, Hangjun Ye, Bingzhao Gao, Long Chen ·

    SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

    arXiv:2608.01899v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or external spatial encoders, which increase complexity …

  38. arXiv cs.CL TIER_1 English(EN) · Shalom Kachko, Raz Lapid, Margarita Vald, Almog Dubin, Moshe Sipper ·

    Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations

    arXiv:2608.00561v1 Announce Type: cross Abstract: Vision-language models (VLMs) process image patches and text tokens in a shared residual stream, but the local geometry through which the two modalities interact remains poorly understood. Most interpretability methods identify gl…

  39. arXiv cs.CL TIER_1 English(EN) · Eugene Lee, Ting-Yu Chang, Jui-Huang Tsai, Jiajie Diao, Chen-Yi Lee ·

    Hierarchical Pre-Training of Vision Encoders with Large Language Model

    arXiv:2604.00086v2 Announce Type: replace-cross Abstract: The field of computer vision has experienced significant advancements through scalable vision encoders and multimodal pre-training frameworks. However, existing approaches often treat vision encoders and large language mod…

  40. arXiv cs.LG TIER_1 English(EN) · Mayank Nautiyal, Li Ju, Andreas Hellander, Ekta Vats, Prashant Singh ·

    GeoFlowVLM: Geometry-Aware Joint Uncertainty for Frozen Vision-Language Embedding

    arXiv:2605.13352v2 Announce Type: replace Abstract: Standard dual-encoder vision-language models that map images and text to deterministic points on a shared unit hypersphere through $\ell_2$ normalization typically expose neither \emph{aleatoric} uncertainty (cross-modal ambigui…

  41. arXiv cs.CL TIER_1 English(EN) · Jin Cui, Chuanchang Su, Jiayi Lu, Xinyue Long, Boran Zhao, Pengju Ren ·

    HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models

    arXiv:2608.02124v1 Announce Type: cross Abstract: Vision-language models (VLMs) remain unreliable when predictions require fine-grained visual evidence. We identify a previously overlooked cause: spectral response rigidity. Despite substantial frequency variation across images an…

  42. arXiv cs.LG TIER_1 English(EN) · Leyan Xue, Feng Xiong, Mingjun Ma, Changqing Zhang ·

    Distill What the Student Can See: Fisher-Projected On-Policy Distillation for Vision-Language Models

    arXiv:2608.01263v1 Announce Type: new Abstract: On-policy distillation (OPD) samples trajectories from the current student policy and minimizes token-level divergence between student and teacher next-token distributions at prefixes along those trajectories. This aligns the distil…

  43. Hugging Face Daily Papers TIER_1 English(EN) ·

    SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

    Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled coun…

  44. Hugging Face Daily Papers TIER_1 English(EN) ·

    SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models

    Vision-language models (VLMs) achieve strong semantic understanding but remain unreliable in metric spatial reasoning, particularly when queries require comparing multiple instances of the same object category. We study this problem through the Closest-Instance Distance Query (CI…

  45. arXiv cs.CL TIER_1 English(EN) · Senyu Fei, Xiaopeng Yu, Siyin Wang, Xianzhong Zhao, Jingjing Gong, Xipeng Qiu ·

    WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

    arXiv:2607.29613v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on…

  46. arXiv cs.AI TIER_1 English(EN) · Md Ashikur Rahman, Md Arifur Rahman, Niamul Hassan Samin, Abdullah Ibne Hanif Arean, Juena Ahmed Noshin ·

    Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models

    arXiv:2603.06828v2 Announce Type: replace-cross Abstract: We uncover a behavioral law of long-horizon vision-language models: models that maintain temporally grounded beliefs generalize better. Standard benchmarks measure only final-answer accuracy, which obscures how models use …

  47. arXiv cs.LG TIER_1 English(EN) · Mushir Akhtar, M. Tanveer ·

    Mitigating Class-Tail Undercoverage in Medical Vision-Language Models under Clinical Shift

    arXiv:2607.28696v1 Announce Type: new Abstract: Medical vision-language models (VLMs) can retain high observed marginal coverage after clinical shift while substantially under-covering an individual disease class. The affected class varies with acquisition protocol and backbone g…

  48. Hugging Face Daily Papers TIER_1 English(EN) ·

    CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models

    In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the prefill stage to rise sharply. Such visual sequences are highly redundant along the spatio-temporal dimension, yet a high compressi…

  49. Hugging Face Daily Papers TIER_1 English(EN) ·

    Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts

    Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this token mix. We call the standard token-level Switch auxiliary loss Std-Aux. Std-Aux balances only the mixed load, so large image…

  50. arXiv cs.LG TIER_1 English(EN) · Chiyuan He, Zihuan Qiu, Fanman Meng, Runtong Zhang, Linfeng Xu, Qingbo Wu, Hongliang Li ·

    Continual Learning with Vision-Language Models via Semantic-Geometry Preservation

    arXiv:2603.12055v3 Announce Type: replace-cross Abstract: Continual learning of pretrained vision-language models (VLMs) is prone to catastrophic forgetting, yet current approaches adapt to new tasks without explicitly preserving the cross-modal semantic geometry inherited from p…

  51. arXiv cs.LG TIER_1 English(EN) · Yao Xiao, Reuben Tan, Zhen Zhu, Yuqun Wu, Jianfeng Gao, Derek Hoiem ·

    ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

    arXiv:2607.28627v1 Announce Type: cross Abstract: Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present Re…

  52. arXiv cs.CL TIER_1 English(EN) · Yuming Yan, Kai Tang, Sihong Chen, Ke Xu, Dan Hu, Qun Yu, Pengfei Hu ·

    S-GRPO: Unified Post-Training for Large Vision-Language Models

    arXiv:2604.16557v2 Announce Type: replace-cross Abstract: Current post-training methodologies for adapting Large Vision-Language Models (LVLMs) generally fall into two paradigms: Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). Despite their prevalence, both approach…

  53. arXiv cs.AI TIER_1 English(EN) · Zhe Liu, Quan Lu, Zhaohui Du, Zhe Wang, Huanbo Jin, Jiaming Gu, Qi Wang, Ting Xiao, Minting Pan, Dongzhan Zhou ·

    BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories

    arXiv:2607.26914v1 Announce Type: cross Abstract: Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms are designed for household environments and treat a target as an object center or an arbit…

  54. arXiv cs.LG TIER_1 English(EN) · Supratik Bhowal, Subhrajyoti Basu, Aritra Gir Mahanta, Anik Pal Chowdhury ·

    Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models

    arXiv:2607.27304v1 Announce Type: new Abstract: Medical vision-language models (VLMs) generate chain-of-thought (CoT) reasoning before answering clinical questions, but whether this reasoning causally influences predictions remains unclear. We present CoT-Mediate, a behavioral fr…

  55. Hugging Face Daily Papers TIER_1 English(EN) ·

    WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

    Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or single-frame VLM bac…

  56. Hugging Face Daily Papers TIER_1 English(EN) ·

    Scaling Vision-Language Models Is Not Enough to Mitigate Bias

    Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly understood at scale. We present the first large-scale empirical study of 194 publicly available VLMs, including 16 model families, c…

  57. Hugging Face Daily Papers TIER_1 English(EN) ·

    ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

    Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding trained as an …

  58. Hugging Face Daily Papers TIER_1 English(EN) ·

    RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

    Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance often degrades on challenging and out-of-domain tasks. Recent test-time steering and scaling methods improve performance without extensive data collection and retraini…

  59. Hugging Face Daily Papers TIER_1 English(EN) ·

    MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models

    Integrating 3D medical images with vision-language models (VLMs) holds substantial promise for computer-aided diagnosis. However, volumetric images generate prohibitively long visual-token sequences with considerable spatial and inter-slice redundancy. Existing token compression …

  60. arXiv cs.AI TIER_1 English(EN) · Minhyeok Lee, Chiyoung Kim, Chanhoe Gu, Seongrok Kim, Sanghyuk Roy Choi, Donghwan Hwang, Donghun Ryu, Seokhyun Kim ·

    CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model

    arXiv:2607.25487v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness benchmark use three- to seven-billion-parameter backbones whose memory demands ca…

  61. arXiv cs.AI TIER_1 English(EN) · Zonghe Liu (University of Hong Kong), Shanyuan Jie (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences), Xiaoquan Sun (Huazhong University of Science and Technology), Chen Cao (University of Hong Kong), Zetian Xu (University of Hong … ·

    SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models

    arXiv:2607.25912v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, especially und…

  62. Hugging Face Daily Papers TIER_1 English(EN) ·

    HumanCLAW: Can Vision-Language Models Act Through a Body?

    Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed t…

  63. Hugging Face Daily Papers TIER_1 English(EN) ·

    TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

    Vision-language-action (VLA) models commonly adopt an LLM-centric V to L to A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial compu…

  64. arXiv cs.AI TIER_1 English(EN) · Ridwan Mahbub, Mohammed Saidul Islam, Md Tahmid Rahman Laskar, Mizanur Rahman, Mir Tafseer Nayeem, Enamul Hoque ·

    Chart Deception in Vision-Language Models: From Vulnerability to Mitigation

    arXiv:2607.22600v1 Announce Type: new Abstract: Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or inverted axes, distorted aspect ratios, inappropriate encodings, and misleading color mappin…

  65. arXiv cs.AI TIER_1 English(EN) · Ali Ansari, Yasmin Mohammadi, Farnoush Nili, Parsa Esmaeilkhani, Longin Jan Latecki, Eduard Dragut ·

    ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams

    arXiv:2607.24707v1 Announce Type: new Abstract: Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiting AI-assisted database engineering. We introduce ER…

  66. arXiv cs.AI TIER_1 English(EN) · Renjie Liang ·

    Cheap Probes Predict Expensive Training in 3D-CT Vision--Language Models

    arXiv:2607.22771v1 Announce Type: cross Abstract: Picking the frozen image encoder for a 3D~CT vision--language model (VLM), together with the token-compression scheme on top of it, is a search over many candidates. There are several encoders, several ways to compress their token…

  67. arXiv cs.AI TIER_1 English(EN) · Qing Yang, Xun Wang, Ziguan Wang, Zhenjiang Li, Hongqiang Wang, Dongdong Weng ·

    Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline

    arXiv:2607.22997v1 Announce Type: cross Abstract: Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the real world -- has emerged as the next major frontier for AI, echoed by industry leaders such as Jensen Huang (``the n…

  68. arXiv cs.AI TIER_1 English(EN) · Khang Nhat Hoang Vo ·

    From Pixels to Prompts: Vision-Language Models

    arXiv:2605.07544v3 Announce Type: replace Abstract: When you read a paper about a new Vision-Language Model today, it can be easy to forget how strange this idea would have sounded not so long ago. Teaching machines to see was already hard. Teaching them to read and generate lang…

  69. arXiv cs.AI TIER_1 English(EN) · Alexander G. Ororbia, Ankur Mali, Mary Alexandria Kelly, David Reitter ·

    Like a Baby: Visually Situated Neural Language Acquisition

    arXiv:1805.11546v3 Announce Type: replace-cross Abstract: We examine the benefits of visual context in training neural language models to perform next-word prediction. A multi-modal neural architecture is introduced that outperform its equivalent trained on language alone with a …

  70. arXiv cs.CL TIER_1 English(EN) · Sultan Alshehri, Zhantao Yang, Han Zhang, Marios Savvides ·

    Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models

    arXiv:2607.23052v1 Announce Type: cross Abstract: Dual-encoder vision-language models (VLMs) expose a similarity interface that enables zero-shot retrieval but fails compositional constraints: queries like "umbrella and no person" retrieve images containing both, even when concep…

  71. arXiv cs.CL TIER_1 English(EN) · M M Asif Ferdous ·

    Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation

    arXiv:2607.24440v1 Announce Type: cross Abstract: Vision-language models (VLMs) deployed on consumer hardware must decide when to answer and when to defer, and that decision depends on having a confidence signal that tracks correctness. A practitioner with a fixed memory budget f…

  72. arXiv cs.CL TIER_1 English(EN) · Khai-Nguyen Nguyen, Zixin Tang, Ankur Mali, Mary ALexandria Kelly ·

    Like a bilingual baby: The advantage of visually grounding a bilingual language model

    arXiv:2210.05487v3 Announce Type: replace Abstract: Unlike most neural language models, humans learn language in a rich, multi-sensory and, often, multi-lingual environment. Current language models typically fail to fully capture the complexities of multilingual language use. We …

  73. arXiv cs.LG TIER_1 English(EN) · Shuai Wang, Daoan Zhang, Zhe Tang, Hao Cheng, Jiaheng Wei ·

    Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation

    arXiv:2607.23125v1 Announce Type: new Abstract: Post-training enables vision-language models (VLMs) to understand human instructions and perform various downstream tasks. Current post-training methods usually rely on human-annotated data, distillation from external models, reinfo…

  74. arXiv cs.AI TIER_1 English(EN) · Jun Ling, Tao Huang, Junzhuo Liu, Bowen Tang, Peng Wang ·

    GOTS: Greedy Orthogonal Token Selection for High-Resolution Vision-Language Models

    arXiv:2607.23913v1 Announce Type: new Abstract: Modern vision-language models (VLMs) increasingly rely on dynamic or high-resolution visual encoding, producing thousands of visual tokens that substantially increase downstream language-model inference cost. Existing token-reductio…

  75. Hugging Face Daily Papers TIER_1 English(EN) ·

    Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

    Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. However, they often fail on vision-centric tasks, especially when visual evidence conflicts with pretrained knowledge. We explore t…

  76. arXiv cs.CL TIER_1 English(EN) · Gwang Gook Lee, Kenan Emir Ak, Jay Mohta, Yan Xu, Dimitrios Dimitriadis ·

    Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models

    arXiv:2607.21617v1 Announce Type: cross Abstract: Vision Language Models (VLMs) are increasingly used in place of traditional OCR pipelines for document understanding. In this paper, we show they do not always act as faithful transcribers: when text is imperfect, they often tend …

  77. Hugging Face Daily Papers TIER_1 English(EN) ·

    PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

    We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures i…

  78. Hugging Face Daily Papers TIER_1 English(EN) ·

    N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

    We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current vision-bas…

  79. Hugging Face Daily Papers TIER_1 English(EN) ·

    UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models

    Large Vision-Language Models (LVLMs) remain bottlenecked by massive computational footprints, precluding their deployment on resource-constrained edge devices. While efforts to compress LVLMs focus heavily on vision token reduction or smaller language models, the vision encoder i…

  80. arXiv cs.AI TIER_1 English(EN) · Dongbin Na ·

    When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation

    arXiv:2607.21401v1 Announce Type: cross Abstract: A vision-language AI assistant returns its answer as a stream of generated tokens. Therefore, a safety guard that watches that answer has to keep up with the stream and stop a harmful reply before a user reads it. Recent vision-la…

  81. arXiv cs.AI TIER_1 English(EN) · Karan Goyal, Afreen Hossain, Debojyoti Das, Vishal Bhutani ·

    ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

    arXiv:2607.20092v1 Announce Type: cross Abstract: Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mec…

  82. arXiv cs.CL TIER_1 English(EN) · Aditi Gupta, Yossi Gandelsman ·

    Test-Time Training for Modality Order Consistency in Vision-Language Models

    arXiv:2607.20351v1 Announce Type: cross Abstract: We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently …

  83. Hugging Face Daily Papers TIER_1 English(EN) ·

    Test-Time Training for Modality Order Consistency in Vision-Language Models

    We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently outperforms question-first prompting, revealing a …

  84. arXiv cs.AI TIER_1 English(EN) · Sibo Wang, Jie Zhang, Shiguang Shan, Xilin Chen, Wen Gao ·

    Dual Adversarial Fine-tuning for Enhancing Robustness of Large Vision Language Model

    arXiv:2607.18958v1 Announce Type: cross Abstract: While Large Vision-Language Models (LVLMs), represented by LLaVA and GPT-4V, have demonstrated remarkable capabilities, their visual inputs remain vulnerable to adversarial attacks, posing significant security risks. Existing defe…

  85. arXiv cs.AI TIER_1 English(EN) · Gautam Rajendrakumar Gare, Jia Shi, Zhiqiu Lin, Deepak Pathak, John Galeotti, Deva Ramanan ·

    Attributes Should Come from Images, Not Class Names: Distribution-Conditioned Attribute Selection for Vision-Language Models

    arXiv:2607.18695v1 Announce Type: cross Abstract: A popular route to interpretable zero-shot classification asks a large language model (LLM) to describe each class name and prompts CLIP with the resulting descriptors. We show that these descriptors carry little visual evidence o…

  86. Hugging Face Daily Papers TIER_1 English(EN) ·

    ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

    Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whet…

  87. arXiv cs.AI TIER_1 English(EN) · Jianing Guo, Zhenhong Wu, Chang Tu, Yiyao Ma, Xiangqi Kong, Zhiqian Liu, Jiaming Ji, Shuning Zhang, Yuanpei Chen, Kai Chen, Qi Dou, Yaodong Yang, Xianglong Liu, Huijie Zhao, Weifeng Lv, Simin Li ·

    RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

    arXiv:2510.00037v5 Announce Type: replace-cross Abstract: In Vision-Language-Actionf(VLA) models, robustness to real-world perturbations is critical for deployment. Existing methods target simple visual disturbances, overlooking the broader multi-modal perturbations that arise in…

  88. arXiv cs.CL TIER_1 English(EN) · Wei Chen, Liangmin Wu, Yunhai Hu, Zhiyuan Li, Zhiyuan Cheng, Yicheng Qian, Lingyue Zhu, Zhipeng Hu, Luoyi Liang, Qiang Tang, Zhen Liu, Han Yang ·

    AutoNeural: Co-Designing Vision-Language Models for NPU Inference

    arXiv:2512.02924v3 Announce Type: replace Abstract: While Neural Processing Units (NPUs) offer high theoretical efficiency for edge AI, state-of-the-art Vision--Language Models (VLMs) tailored for GPUs often falter on these substrates. We attribute this hardware-model mismatch to…

  89. arXiv cs.CL TIER_1 English(EN) · Tomer Ullman ·

    The Illusion-Illusion: Vision Language Models See Illusions Where There Are None

    arXiv:2412.18613v2 Announce Type: replace-cross Abstract: Illusions are entertaining, but they are also a useful diagnostic tool in cognitive science, philosophy, and neuroscience. A typical illusion shows a gap between how something `really is' and how something `appears to be',…

  90. arXiv cs.AI TIER_1 English(EN) · Tuan Duong Trinh, Naveed Akhtar, Basim Azam ·

    Reasoning as a Double-Edged Sword: Architecture and Cross-Stage Robustness in Vision-Language-Action Models

    arXiv:2607.17786v1 Announce Type: cross Abstract: Does adding a reasoning step make a Vision-Language-Action (VLA) model more robust to perturbation? Intuitively, a policy that reasons before acting should absorb a perturbed input better than one that maps observations directly t…

  91. arXiv cs.CL TIER_1 English(EN) · Sudharshan Balaji, Yili Ren, Guangjing Wang, Yimin Chen, Ning Wang ·

    One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models

    arXiv:2607.16442v1 Announce Type: cross Abstract: Machine unlearning is widely used to remove hazardous knowledge from large language models. Modern Vision-Language Models (VLMs), however, process both text and visual inputs, raising a fundamental security question: does unlearni…

  92. arXiv cs.AI TIER_1 English(EN) · Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman ·

    SceneBind: Binding What and Where Across Vision, Audio and Language

    arXiv:2607.15265v1 Announce Type: cross Abstract: We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what …

  93. arXiv cs.LG TIER_1 English(EN) · Pegah Khayatan, Sara Meziane, Jayneel Parekh, Matthieu Cord ·

    DiMaS: Distribution Matching for Steering Vision-Language-Action Models

    arXiv:2607.14280v1 Announce Type: cross Abstract: Flow-matching-based vision-language-action (VLA) models have emerged as powerful policies for robotic manipulation, yet a critical capability remains underexplored: fine-grained behavioral control, the ability to govern how a robo…

  94. arXiv cs.AI TIER_1 English(EN) · Wei Li, Peijin Jia, Yuan Ma, Xuefeng Jiang, Titong Jiang, Sheng Sun, Yujian Li, Xin Wen, Han Hong, Zhikang Liu, Bailin Li, Kun Zhan ·

    FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models

    arXiv:2607.14739v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have achieved impressive results in visuomotor policy learning, yet remain fundamentally reactive, mapping current observations and language to actions without explicit forward prediction of wor…

  95. arXiv cs.AI TIER_1 English(EN) · Yijiang Li, Huiqi Zou, Bingyang Wang, Ziang Xiao ·

    Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions

    arXiv:2607.14499v1 Announce Type: new Abstract: Multi-modal Large Language Models (MLLMs) have made substantial advances on benchmarks, yet their real-world effectiveness remains uncertain. This gap stems from the fundamental misalignment between benchmarks in controlled, static …

  96. arXiv cs.AI TIER_1 English(EN) · Eli Shlizerman ·

    SceneBind: Binding What and Where Across Vision, Audio and Language

    We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial struc…

  97. arXiv cs.AI TIER_1 English(EN) · Kun Zhan ·

    FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models

    Vision-Language-Action (VLA) models have achieved impressive results in visuomotor policy learning, yet remain fundamentally reactive, mapping current observations and language to actions without explicit forward prediction of world dynamics. Existing visual foresight methods pre…

  98. arXiv cs.AI TIER_1 English(EN) · Jose Mart\'inez-Fajardo, Pablo Pueyo, Fernando Caballero, Luis Merino ·

    From Language to Navigation Goals: A Vision-Language Approach for Semantic Navigation of Mobile Robots Using RGB-D Perception

    arXiv:2607.13624v1 Announce Type: cross Abstract: Natural language interaction provides an intuitive way for non-expert users to communicate with robotic platforms. However, transforming user requests into executable navigation actions remains a challenging task, requiring the in…

  99. arXiv cs.LG TIER_1 English(EN) · Charlotte Morissette, Amin Abyaneh, Wei-Di Chang, Anas Houssaini, David Meger, Hsiu-Chin Lin, Jonathan Tremblay, Gregory Dudek ·

    Tactile Modality Fusion for Vision-Language-Action Models

    arXiv:2603.14604v2 Announce Type: replace-cross Abstract: We propose TacFiLM, a lightweight modality-fusion approach that integrates visual-tactile signals into vision-language-action (VLA) models. While advances in VLAs have introduced robot policies that are both generalizable …

  100. arXiv cs.AI TIER_1 English(EN) · Luyuan Jia, Yinfeng Yu ·

    HRO: Hierarchical Room-to-Object Framework for Zero-Shot Object Goal Navigation with Large Language Models

    arXiv:2607.13072v1 Announce Type: cross Abstract: Zero-shot object-goal navigation aims to enable an intelligent agent to explore and navigate to objects of unknown categories in an unfamiliar environment without specific target training. In zero-shot navigation tasks, pre-traine…

  101. Hugging Face Daily Papers TIER_1 English(EN) ·

    From Language to Navigation Goals: A Vision-Language Approach for Semantic Navigation of Mobile Robots Using RGB-D Perception

    Natural language interaction provides an intuitive way for non-expert users to communicate with robotic platforms. However, transforming user requests into executable navigation actions remains a challenging task, requiring the integration of language understanding, environment p…

  102. arXiv cs.AI TIER_1 English(EN) · Luis Merino ·

    From Language to Navigation Goals: A Vision-Language Approach for Semantic Navigation of Mobile Robots Using RGB-D Perception

    Natural language interaction provides an intuitive way for non-expert users to communicate with robotic platforms. However, transforming user requests into executable navigation actions remains a challenging task, requiring the integration of language understanding, environment p…

  103. arXiv cs.AI TIER_1 English(EN) · Jiawei Liang, Jianjie Huang, Ruoyu Chen, Xianghao Jiao, Siyuan Liang, Shiming Liu, Xiaochun Cao ·

    Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models

    arXiv:2509.22415v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains difficult to inspect. Recent logit-lens attribution methods project each visual-token…

  104. arXiv cs.AI TIER_1 English(EN) · Peiyu Li, Xiaobao Huang, Ting Hua, Nitesh V. Chawla ·

    CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?

    arXiv:2511.09483v3 Announce Type: replace Abstract: While multimodal large language models can describe visual content, their ability to generate executable procedures remains underexplored. CrochetBench presented in this paper evaluates this shift from describing to doing throug…

  105. arXiv cs.LG TIER_1 English(EN) · Yuanjie Lu, Beichen Wang, Zhengqi Wu, Yang Li, Xiaomin Lin, Chengzhi Mao, Xuesu Xiao ·

    APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action Model

    arXiv:2603.08862v2 Announce Type: replace-cross Abstract: Autonomous navigation in highly constrained environments remains challenging for mobile robots. Classical navigation approaches offer safety assurances but require environment-specific parameter tuning; end-to-end learning…

  106. arXiv cs.AI TIER_1 English(EN) · Maijunxian Wang, Yijiang Li, Bingyang Wang, Tianwei Zhao, Ran Ji, Qingying Gao, Emmy Liu, Hokin Deng, Dezhi Luo ·

    Egocentric Bias in Vision-Language Models

    arXiv:2602.15892v2 Announce Type: replace-cross Abstract: Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce FlipSet, a diagnostic benchmark for Level-2 visual perspective taking (L2 VPT) in visio…

  107. arXiv cs.AI TIER_1 English(EN) · Hiroto Osaka, Shohei Taniguchi, Gouki Minegishi, Kai Yamashita, Masahiro Suzuki, Yutaka Matsuo ·

    Visual Access Boundaries in Vision-Language Model Reasoning

    arXiv:2607.12815v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires conti…

  108. Hugging Face Daily Papers TIER_1 English(EN) ·

    Visual Access Boundaries in Vision-Language Model Reasoning

    Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires continued access to image tokens, or whether it mainl…

  109. arXiv cs.AI TIER_1 English(EN) · Yutaka Matsuo ·

    Visual Access Boundaries in Vision-Language Model Reasoning

    Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires continued access to image tokens, or whether it mainl…

  110. Hugging Face Daily Papers TIER_1 English(EN) ·

    CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models

    Cross-image comparative reasoning remains challenging for vision-language models (VLMs), especially when correct prediction requires fine-grained attribute grounding and globally consistent reasoning. We present CoRe, a unified framework for this problem. CoRe includes: (i) CoRe-…

  111. arXiv cs.AI TIER_1 English(EN) · Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, … ·

    ABot-N1: Toward a General Visual Language Navigation Foundation Model

    arXiv:2607.10383v1 Announce Type: cross Abstract: Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic polici…

  112. arXiv cs.AI TIER_1 English(EN) · Hayeon Kim, Ji Ha Jang, Junghun James Kim, Se Young Chun ·

    Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models

    arXiv:2603.22042v3 Announce Type: replace-cross Abstract: While Vision-Language Models (VLMs) have achieved remarkable performance, their Euclidean embeddings remain limited in capturing hierarchical relationships such as part-to-whole or parent-child structures, and often face c…

  113. arXiv cs.AI TIER_1 English(EN) · Murad Farzulla ·

    Context-Dependent Affordance Computation in Vision-Language Models

    arXiv:2603.04419v2 Announce Type: replace-cross Abstract: We characterize the phenomenon of context-dependent affordance computation in vision-language models (VLMs). Our primary study uses Qwen3-VL-30B-A3B ($n = 3{,}213$ scene-context pairs from COCO-2017: 479 images under 7 age…

  114. arXiv cs.AI TIER_1 English(EN) · Byungkun Lee, Dongyoon Hwang, Dongjin Kim, Hojoon Lee, Minho Park, Jaegul Choo ·

    See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

    arXiv:2607.11498v1 Announce Type: cross Abstract: Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, c…

  115. arXiv cs.LG TIER_1 English(EN) · Finn Ferchau, Daniel Pommer, Cristian Axenie ·

    On the Efficiency of LoRA Fine-Tuning for Vision-Language-Action Models in Industrial Robotic Manipulation

    arXiv:2607.10172v1 Announce Type: cross Abstract: Deploying billion-parameter Vision-Language-Action (VLA) models on industrial hardware requires fine-tuning to bridge the embodiment gap. Full Fine-Tuning (FFT) provides maximal plasticity but requires data centre-grade GPUs. We p…

  116. arXiv cs.AI TIER_1 English(EN) · Mingjie Xie, Guangjun He, Dongli Xu, Youtian Lin, Hongjue Li, Pengming Feng, Jian Guan, Yue Deng ·

    SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception

    arXiv:2607.11008v1 Announce Type: cross Abstract: Open-vocabulary dense perception (OVDP) aims to localize objects unseen during training by leveraging textual knowledge. Despite the remarkable progress of recent CLIP-based approaches, we identify a critical limitation: synonym-i…

  117. arXiv cs.AI TIER_1 English(EN) · Liuyi Wang, Kai Sheng, Zongtao He, Jinlong Li, Yongrui Qin, Haojie Dai, Xiangyi Wang, Jingwei Yang, Qingqing Yan, Chengju Liu, Qijun Chen ·

    A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation

    arXiv:2607.09792v1 Announce Type: cross Abstract: Navigation is a fundamental capability of autonomous systems, yet most existing approaches rely on highly structured models and strong prior assumptions, limiting their robustness in open and uncertain real-world environments. Vis…

  118. arXiv cs.AI TIER_1 English(EN) · Shengzhuo Yang, Ronghao Yu, Chuanjie Lv, Linpeng Peng, Hang Yu, Jie Ren, Jiajun Lv, Yong Liu ·

    TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging

    arXiv:2607.09818v1 Announce Type: cross Abstract: Vision-language-action (VLA) models aim to understand natural-language instructions and visual observations, and to generate and execute corresponding actions as embodied agents. Recently, autoregressive token-based action generat…

  119. arXiv cs.AI TIER_1 English(EN) · Jaegul Choo ·

    See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

    Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene i…

  120. arXiv cs.AI TIER_1 English(EN) · Shravan Murlidaran, Miguel P. Eckstein ·

    Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models

    arXiv:2607.09654v1 Announce Type: cross Abstract: Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactions or behaviors, only a handfu…

  121. arXiv cs.LG TIER_1 English(EN) · Xingyu Zhu, Huanshen Wu, Shuo Wang, Beier Zhu, Jiannan Ge, Jiaheng Zhang, Long Chen ·

    Robustifying Vision-Language Models via Test-Time Prompt Adaptation

    arXiv:2607.09450v1 Announce Type: cross Abstract: Pre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations. Existing test-time adaptation methods typically rely on sample…

  122. arXiv cs.AI TIER_1 English(EN) · Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen ·

    Scalable Visual Pretraining for Language Intelligence

    arXiv:2607.09657v1 Announce Type: cross Abstract: The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations…

  123. Hugging Face Daily Papers TIER_1 English(EN) ·

    ABot-N1: Toward a General Visual Language Navigation Foundation Model

    Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map observations directly to actions, yet …

  124. arXiv cs.AI TIER_1 English(EN) · Kai Chen ·

    Scalable Visual Pretraining for Language Intelligence

    The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that can…

  125. arXiv cs.AI TIER_1 English(EN) · Miguel P. Eckstein ·

    Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models

    Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactions or behaviors, only a handful of non-curated human descriptions as a benchmark…

  126. arXiv cs.LG TIER_1 English(EN) · Long Chen ·

    Robustifying Vision-Language Models via Test-Time Prompt Adaptation

    Pre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations. Existing test-time adaptation methods typically rely on sample-level confidence heuristics, overlooking the intr…

  127. arXiv cs.AI TIER_1 English(EN) · Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu, Vaidehi Patil ·

    Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

    arXiv:2607.07907v1 Announce Type: cross Abstract: With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associations that originate from their training data. Retrai…

  128. arXiv cs.AI TIER_1 English(EN) · Qi Lyu, Baicheng Liu, Xudong Wang, Jiahua Dong, Lianqing Liu, Zhi Han ·

    LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action

    arXiv:2607.08182v1 Announce Type: cross Abstract: Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all visual tokens uniformly and reasoning with human-…

  129. arXiv cs.AI TIER_1 English(EN) · Emily Jin, Joy Hsu, Yiqing Xu, Weiyu Liu, Nick Haber, Jiajun Wu ·

    APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts

    arXiv:2607.08024v1 Announce Type: cross Abstract: Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility. To successfully execute a task, a robot must decompose goals, select task-relevant objects, and sequence actions, while…

  130. Hugging Face Daily Papers TIER_1 English(EN) ·

    Scalable Visual Pretraining for Language Intelligence

    The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that can…

  131. Hugging Face Daily Papers TIER_1 English(EN) ·

    LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action

    Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all visual tokens uniformly and reasoning with human-selected factors, which lack mechanisms to emphasi…

  132. arXiv cs.AI TIER_1 English(EN) · Zhi Han ·

    LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action

    Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all visual tokens uniformly and reasoning with human-selected factors, which lack mechanisms to emphasi…

  133. arXiv cs.AI TIER_1 English(EN) · Inkyu Sa, Chanoh Park, Hea-Min Lee, Donghee Noh, Ho Seok Ahn ·

    Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review

    arXiv:2607.06706v1 Announce Type: cross Abstract: Vision Language Action (VLA) models unify visual perception, natural-language understanding, and action generation within a single foundation model, allowing a robot to follow instructions such as fold the towel or fly to the red …

  134. arXiv cs.AI TIER_1 English(EN) · Chethan Krishnamurthy Ramanaik, Tobias Callies, Michael Hecht, Eirini Ntoutsi ·

    On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces

    arXiv:2607.07375v1 Announce Type: cross Abstract: Adversarial vulnerability in deep neural networks (DNNs) has been studied from the perspectives of decision-boundary geometry, feature robustness, input-output Jacobians, and the instability of inverse problems. Here, we focus on …

  135. arXiv cs.AI TIER_1 English(EN) · Peter Bohm, Saimunur Rahman, Abdelwahed Khamis, Sagun Man Singh Shrestha, Chris McCool, Peyman Moghadam ·

    GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model

    arXiv:2607.06882v1 Announce Type: cross Abstract: Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke action head, and training on thousands of hours of cross-embodiment datasets. We ask whether t…

  136. arXiv cs.AI TIER_1 English(EN) · Juyi Lin, Amir Taherin, Arash Akbari, Arman Akbari, Lei Lu, Guangyu Chen, Taskin Padir, Xiaomeng Yang, Weiwei Chen, Yiqian Li, Xue Lin, David Kaeli, Pu Zhao, Yanzhi Wang ·

    VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

    arXiv:2507.05116v5 Announce Type: replace-cross Abstract: Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic manipulation tasks guided by natural language. However, current VLA models suffer from two drawbacks: (i) generation of mass…

  137. arXiv cs.CL TIER_1 English(EN) · Kaito Watanabe, Taisei Yamamoto, Tomoki Doi, Hitomi Yanaka ·

    Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models

    arXiv:2607.07251v1 Announce Type: new Abstract: One of the expected abilities of vision-language models (VLMs) is spatial reasoning ability based on a given text and image. To evaluate the spatial reasoning abilities of VLMs, we focus on the use of spatial deictic expressions, wh…

  138. arXiv cs.AI TIER_1 English(EN) · Zhiying Du, Bei Liu, Yaobo Liang, Yichao Shen, Haidong Cao, Xiangyu Zheng, Zhiyuan Feng, Zuxuan Wu, Jiaolong Yang, Yu-Gang Jiang ·

    HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies

    arXiv:2512.05693v2 Announce Type: replace-cross Abstract: Generalist vision--language--action (VLA) policies are typically trained on heterogeneous mixtures of robot demonstrations spanning diverse embodiments, action spaces, and observation configurations. Modeling such heteroge…

  139. arXiv cs.CL TIER_1 English(EN) · Vaidehi Patil ·

    Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

    With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associations that originate from their training data. Retraining after deletion requests or policy updates is …

  140. arXiv cs.AI TIER_1 English(EN) · Eirini Ntoutsi ·

    On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces

    Adversarial vulnerability in deep neural networks (DNNs) has been studied from the perspectives of decision-boundary geometry, feature robustness, input-output Jacobians, and the instability of inverse problems. Here, we focus on the spectral structure of intermediate linear tran…

  141. arXiv cs.CL TIER_1 English(EN) · Hitomi Yanaka ·

    Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models

    One of the expected abilities of vision-language models (VLMs) is spatial reasoning ability based on a given text and image. To evaluate the spatial reasoning abilities of VLMs, we focus on the use of spatial deictic expressions, which are defined as spatial expressions whose ref…

  142. arXiv cs.LG TIER_1 English(EN) · Haitao Lin, Hanyang Yu, Jingshun Huang, He Zhang, Yonggen Ling, Ping Tan, Xiangyang Xue, Yanwei Fu ·

    PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

    arXiv:2602.19710v3 Announce Type: replace-cross Abstract: Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision. Since these model…

  143. arXiv cs.AI TIER_1 English(EN) · Sohwi Lim, Lee Hyoseok, Jungjoon Park, Tae-Hyun Oh ·

    CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space

    arXiv:2604.11539v2 Announce Type: replace-cross Abstract: Human perception of visual similarity is inherently adaptive and subjective, depending on the users' interests and focus. However, most image retrieval systems fail to reflect this flexibility, relying on a fixed, monolith…

  144. arXiv cs.CL TIER_1 English(EN) · Liang Chen, Weichu Xie, Yiyan Liang, Hongfeng He, Hans Zhao, Zhibo Yang, Zhiqi Huang, Haoning Wu, Haoyu Lu, Y. charles, Yiping Bao, Yuantao Fan, Guopeng Li, Haiyang Shen, Xuanzhong Chen, Wendong Xu, Shuzheng Si, Zefan Cai, Wenhao Chai, Ziqi Huang, Fangfu… ·

    BabyVision: Visual Reasoning Beyond Language

    arXiv:2601.06521v2 Announce Type: replace-cross Abstract: While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a cruc…

  145. arXiv cs.LG TIER_1 English(EN) · Ryuji Oi, Hikari Otsuka, Kosuke Matsushima, Yuki Ichikawa, Masato Motomura, Tatsuya Kaneko, Daichi Fujiki ·

    Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement

    arXiv:2607.06370v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow matching-based VLA models have shown remarkable success due to their capability to generate prec…

  146. Hugging Face Daily Papers TIER_1 English(EN) ·

    GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model

    Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke action head, and training on thousands of hours of cross-embodiment datasets. We ask whether this recipe is necessary. In this paper, we introdu…

  147. Hugging Face Daily Papers TIER_1 English(EN) ·

    Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

    LaMem-VLA introduces a latent-memory-native framework that integrates historical experience into vision-language-action reasoning through coordinated memory components operating in the same latent space.

  148. arXiv cs.LG TIER_1 English(EN) · Daichi Fujiki ·

    Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement

    Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow matching-based VLA models have shown remarkable success due to their capability to generate precise and smooth action sequences and capture multim…

  149. Hugging Face Daily Papers TIER_1 English(EN) ·

    Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement

    Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow matching-based VLA models have shown remarkable success due to their capability to generate precise and smooth action sequences and capture multim…

  150. arXiv cs.LG TIER_1 English(EN) · Harsh Goel, S P Sharan, Sahil Shah, Minkyu Choi, Joungbin An, Kristen Grauman, Sandeep P. Chinchali ·

    Incentivizing Vision Language Models to Search for Long Video Question Answering

    arXiv:2607.02959v1 Announce Type: cross Abstract: We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process. VSeek utilizes a natural language-driven search to iden…

  151. arXiv cs.AI TIER_1 English(EN) · Guli Zhu, Chenwei Wu, Liyue Shen ·

    Evaluating and Understanding Model Editing for Medical Vision Language Models

    arXiv:2607.05310v1 Announce Type: new Abstract: Model editing promises a fast, targeted way to correct post-deployment mistakes in medical vision-language models (VLMs) without costly retraining. However, existing multimodal model editing benchmarks focus on general-purpose tasks…

  152. arXiv cs.AI TIER_1 English(EN) · Kaiyun Yang, Ruilin Yang, Zhimin Yao, J. Wang, Wei Ge ·

    Criterion-Conditional In-Context Learning: Evaluating Criterion-Shift Adaptation in Vision-Language Models

    arXiv:2607.02575v1 Announce Type: cross Abstract: Vision-language models can perform new tasks without parameter updates through in-context learning (ICL), whose core mechanism is utilizing the support set for task induction. In the standard ICL setting, once the task is induced,…

  153. arXiv cs.AI TIER_1 English(EN) · Zikai Zhang, Rui Hu, Olivera Kotevska, Jiahao Xu ·

    Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models

    arXiv:2607.02819v1 Announce Type: cross Abstract: Cloud-edge Large Vision-Language Model (LVLM) inference enables efficient deployment by splitting computation between edge devices and cloud servers. In this process, intermediate vision tokens are transmitted from the edge to the…

  154. arXiv cs.AI TIER_1 English(EN) · Junchi Liao, Jiawen Deng, Fuji Ren ·

    VISTA: Auditing Semantic Divergence in Vision-Language Models

    arXiv:2607.02995v1 Announce Type: cross Abstract: Vision-language models can exhibit visual concept-conditioned divergence: given images containing demographic features, corporate logos, or ideological symbols, some models produce unusually uniform responses that differ from what…

  155. arXiv cs.AI TIER_1 English(EN) · Chengzhen Yu, Canran Xiao, Siyuan Ma, Yang Liu ·

    Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning

    arXiv:2607.03143v1 Announce Type: cross Abstract: Vision-language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified, making similarity brittle and overly confident under paraphrase and omitted details. We ai…

  156. arXiv cs.AI TIER_1 English(EN) · Qi Liu, Yabei Li, Hongsong Wang, Heng Zhang, Lei He ·

    AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning

    arXiv:2607.03182v1 Announce Type: cross Abstract: Autonomous driving planning requires translating navigation intent, traffic rules, dynamic interactions, and language instructions into executable continuous trajectories. Vision-Language-Action models have been introduced into dr…

  157. arXiv cs.AI TIER_1 English(EN) · Li Ji, Siyin Wang, Pengfang Qian, Xiaopeng Yu, Yihai Tian, Zhaoye Fei, Jingjing Gong, Xipeng Qiu ·

    HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control

    arXiv:2607.03449v1 Announce Type: cross Abstract: Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian tasks requiring long-term memory and reasoning due to their reliance on immediate observations. Existing solutions fac…

  158. arXiv cs.AI TIER_1 English(EN) · Seung Il Lee, Qinqian Lei, Daguang Xu, Dong Yang, Robby T. Tan, Yixin Chen, Bo Wang ·

    Token-Based Affordance Grounding with Large Vision-Language Models

    arXiv:2607.03595v1 Announce Type: cross Abstract: Affordance grounding aims to localize image regions that support a specific action, serving as a core capability for physical intelligence and embodied perception. Previous studies have primarily relied on weakly supervised learni…

  159. arXiv cs.AI TIER_1 English(EN) · Kai Tang, Jinhao You, Bohua Zhang, Yichen Guo, Yiding Sun, Dongxu Zhang, Chenxi Li, Xiande Huang, Shanghang Zhang ·

    SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering

    arXiv:2607.04163v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual understanding tasks such as image captioning and visual question answering. However, they remain susceptible to hallucinations, generating content th…

  160. arXiv cs.AI TIER_1 English(EN) · Xinchuan Qiu, Yi Yu ·

    Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning

    arXiv:2607.04591v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation by integrating visual perception, language understanding, and robot action generation. Existing research has primarily focused on im…

  161. arXiv cs.AI TIER_1 English(EN) · Riccardo Renzulli, Gabriele Spadaro, Shruthi Gowda, Alaa Eddine Mazouz, Van-Tam Nguyen ·

    TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models

    arXiv:2607.04593v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated impressive capabilities across different tasks, but their computational cost is dominated by the large number of visual tokens fed to the language model. Existing token reduction met…

  162. arXiv cs.AI TIER_1 English(EN) · Suhyeong Park, Junha Jung, Jungwoo Park, Jaewoo Kang ·

    Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval

    arXiv:2607.04605v1 Announce Type: cross Abstract: Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce t…

  163. arXiv cs.AI TIER_1 English(EN) · Mansi Phute, Ravikumar Balakrishnan ·

    VISOR: Visual Input-based Steering for Output Redirection in Vision-Language Models

    arXiv:2508.08521v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) are increasingly being used in a broad range of applications, bringing their security and behavioral control to the forefront. While existing approaches for behavioral control or output redire…

  164. arXiv cs.AI TIER_1 English(EN) · Ravikumar Balakrishnan, Mansi Phute ·

    VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models

    arXiv:2509.25533v2 Announce Type: replace-cross Abstract: As Vision Language Models (VLMs) are deployed across safety-critical applications, understanding and controlling their behavioral patterns has become increasingly important. Existing behavioral control methods face signifi…

  165. arXiv cs.AI TIER_1 English(EN) · Suhyeok Jang, Dongyoung Kim, Changyeon Kim, Youngsuk Kim, Jinwoo Shin ·

    Verifier-free Test-Time Sampling for Vision-Language-Action Models

    arXiv:2510.05681v2 Announce Type: replace-cross Abstract: Vision-Language-Action models (VLAs) have demonstrated remarkable performance in robot control. However, they remain fundamentally limited in tasks that require high precision due to their single-inference paradigm. While …

  166. arXiv cs.AI TIER_1 English(EN) · George-Andrei Dima, R\u{a}zvan-Alexandru Sm\u{a}du, Dumitru-Clementin Cercel ·

    Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models

    arXiv:2512.14926v2 Announce Type: replace-cross Abstract: Focusing on low-resource languages is an essential step toward democratizing generative AI. In this work, we contribute to reducing the multimodal NLP resource gap for Romanian. We translate the widely known Flickr30K data…

  167. arXiv cs.AI TIER_1 English(EN) · Qingqian Yang, Hao Wang, Sai Qian Zhang, Jian Li, Yang Hua, Miao Pan, Tao Song, Zhengwei Qi, Haibing Guan ·

    pFedNavi: Structure-Aware Personalized Federated Vision-Language Navigation for Embodied AI

    arXiv:2602.14401v2 Announce Type: replace-cross Abstract: Vision-Language Navigation VLN requires large-scale trajectory instruction data from private indoor environments, raising significant privacy concerns. Federated Learning FL mitigates this by keeping data on-device, but va…

  168. arXiv cs.CL TIER_1 English(EN) · Israfel Salazar, Stella Frank, Dan Oneata, Desmond Elliott, Constanza Fierro ·

    Pathways of Visual Information Flow in Vision-Language Models

    arXiv:2607.03358v1 Announce Type: cross Abstract: We study how visual information is routed in vision-language models (VLMs). Using causal patching on controlled synthetic and natural datasets, we find that models rely on two distinct pathways to solve visual tasks: A direct path…

  169. arXiv cs.CL TIER_1 English(EN) · Khang Nhat Hoang Vo, Artem Vazhentsev, Artem Shelmanov, Timothy Baldwin, Yova Kementchedjhieva ·

    Does It Fail to See or Fail to Know? Attributing Errors in Vision-Language Models

    arXiv:2607.04683v1 Announce Type: cross Abstract: Vision-language models (VLMs) perform well on visual question answering with high-quality images but struggle when questions require knowledge beyond what is clearly and directly visible. In such settings, uncertainty quantificati…

  170. arXiv cs.LG TIER_1 English(EN) · Zelin Zhao, Min Shi, Bo Yuan, Haotian Xue, Jialuo Li, Lama Moukheiber, Humphrey Shi, Yongxin Chen ·

    WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling

    arXiv:2607.03461v1 Announce Type: cross Abstract: World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in Vision-Language-Action-World (VLAW) modeling. Meanwhile, unified vision-langu…

  171. arXiv cs.LG TIER_1 English(EN) · Jaeyoung Kim, Eunseok Kim, Dongsuk Jang ·

    Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models

    arXiv:2607.05268v1 Announce Type: cross Abstract: Whether a hyperbolic representation model uses its geometry cannot be read off its curvature parameter: what matters is the dimensionless operating point $\sqrt{c}\rho$ and whether the radial and cone machinery is active there. We…

  172. arXiv cs.AI TIER_1 English(EN) · Liyue Shen ·

    Evaluating and Understanding Model Editing for Medical Vision Language Models

    Model editing promises a fast, targeted way to correct post-deployment mistakes in medical vision-language models (VLMs) without costly retraining. However, existing multimodal model editing benchmarks focus on general-purpose tasks and do not reflect realistic clinical domain re…

  173. arXiv cs.LG TIER_1 English(EN) · Dongsuk Jang ·

    Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models

    Whether a hyperbolic representation model uses its geometry cannot be read off its curvature parameter: what matters is the dimensionless operating point $\sqrt{c}ρ$ and whether the radial and cone machinery is active there. We develop a battery of necessary-condition diagnostics…

  174. Hugging Face Daily Papers TIER_1 English(EN) ·

    SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments

    Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test vision-language models under the visual and procedural conditions of real industrial CCTV, where workers appear as distant figur…

  175. arXiv cs.CL TIER_1 English(EN) · Yova Kementchedjhieva ·

    Does It Fail to See or Fail to Know? Attributing Errors in Vision-Language Models

    Vision-language models (VLMs) perform well on visual question answering with high-quality images but struggle when questions require knowledge beyond what is clearly and directly visible. In such settings, uncertainty quantification should not only indicate whether the model is l…

  176. Hugging Face Daily Papers TIER_1 English(EN) ·

    Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval

    Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce this cost, yet they can remove or collapse object- …

  177. arXiv cs.CL TIER_1 English(EN) · Jaewoo Kang ·

    Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval

    Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce this cost, yet they can remove or collapse object- …

  178. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Jaewoo Kang ·

    Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval

    Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce this cost, yet they can remove or collapse object- …

  179. Hugging Face Daily Papers TIER_1 English(EN) ·

    Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning

    Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation by integrating visual perception, language understanding, and robot action generation. Existing research has primarily focused on improving model architectures, training strategies, …

  180. Hugging Face Daily Papers TIER_1 English(EN) ·

    Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval

    Object-aware token merging framework SaMer compresses image-side tokens while preserving query-selectable visual evidence, achieving significant storage reduction and improved retrieval performance.

  181. arXiv cs.AI TIER_1 English(EN) · Ninghao Zhang, Bin Zhu, Shijie Zhou, Jingjing Chen ·

    Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration

    arXiv:2603.06001v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models enable robots to perform manipulation tasks directly from natural language instructions and are increasingly viewed as a foundation for generalist robotic policies. However, their reliab…

  182. arXiv cs.AI TIER_1 English(EN) · Tien-Huy Nguyen, Minh-Nhat Nguyen, Nguyen Nhat Huy, Hung Viet Nguyen, Huy Nguyen Minh Nhat, Thanh-Huy Nguyen, Cuong Tuan Nguyen, Hoang M. Le, Dat Nguyen, Phat Kim Huynh, Min Xu, Ulas Bagci ·

    ESC: Emotional Self-Correction for Reliable Vision-Language Models

    arXiv:2607.02089v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved strong performance across diverse multimodal tasks, yet they remain vulnerable to unreliable reasoning. Existing self-correction methods mitigate these issues but typically rely on post-…

  183. arXiv cs.AI TIER_1 English(EN) · Qi Lyu, Jiahua Dong, Baichen Liu, Xudong Wang, Mingfei Han, Yulun Zhang, Fahad Shahbaz Khan, Salman Khan, Lianqing Liu, Zhi Han ·

    SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models

    arXiv:2607.01876v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal understanding, yet their enormous parameter scale and cross-modal computation incur substantial memory and latency overhead, severely limiting re…

  184. arXiv cs.AI TIER_1 English(EN) · Sung June Kim, Sangpil Kim, Honglak Lee ·

    Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation

    arXiv:2607.01754v1 Announce Type: new Abstract: On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader state distribution. However, such exploration inevitably leads to trajectories that deviate fr…

  185. arXiv cs.AI TIER_1 English(EN) · Phillip Howard, Xin Su, Kathleen C. Fraser ·

    Cross-Cultural Value Attribution in Large Vision-Language Models

    arXiv:2604.09945v2 Announce Type: replace-cross Abstract: The rapid adoption of large vision-language models (LVLMs) in recent years has been accompanied by growing fairness concerns due to their propensity to reinforce harmful societal stereotypes. While significant attention ha…

  186. arXiv cs.AI TIER_1 English(EN) · Rintaro Otsubo, Ryo Fujii, Reina Ishikawa, Taiki Kanaya, Kanta Sawafuji, Hiroki Kajita, Shigeki Sakai, Hideo Saito, Ryo Hachiuma ·

    AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

    arXiv:2607.02269v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation protocols are largely confined to zero-shot assessments on general, daily-life benchmarks. This…

  187. arXiv cs.AI TIER_1 English(EN) · Ryo Hachiuma ·

    AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

    Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation protocols are largely confined to zero-shot assessments on general, daily-life benchmarks. This creates a critical disconnect from real-world app…

  188. Hugging Face Daily Papers TIER_1 English(EN) ·

    SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models

    Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal understanding, yet their enormous parameter scale and cross-modal computation incur substantial memory and latency overhead, severely limiting real-world deployment on resource-constrained device…

  189. Hugging Face Daily Papers TIER_1 English(EN) ·

    Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation

    On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader state distribution. However, such exploration inevitably leads to trajectories that deviate from expert demonstrations, resulting in a semanti…

  190. arXiv cs.LG TIER_1 English(EN) · Sajjad Ghiasvand, Haniyeh Ehsani Oskouie, Mahnoosh Alizadeh, Ramtin Pedarsani ·

    MMLoP: Multi-Modal Low-Rank Prompting for Efficient Vision-Language Adaptation

    arXiv:2602.21397v2 Announce Type: replace-cross Abstract: Prompt learning has become a dominant paradigm for adapting vision-language models (VLMs) such as CLIP to downstream tasks without modifying pretrained weights. While extending prompts to both vision and text encoders acro…

  191. arXiv cs.AI TIER_1 English(EN) · Lukas Kuhn, Giuseppe Serra, Randall Balestriero, Florian Buettner ·

    LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives

    arXiv:2607.00784v1 Announce Type: cross Abstract: Vision-language pretraining remains dominated by contrastive objectives, whereas vision-only self-supervised learning has largely adopted non-contrastive methods. At the same time, the role of vision-language encoders has shifted:…

  192. arXiv cs.AI TIER_1 English(EN) · Shaoheng Zhang, Zhichen Li, Jie Mei ·

    DART-VLN: Test-Time Memory Decay and Anti-Loop Regularization for Discrete Vision-Language Navigation

    arXiv:2607.01043v1 Announce Type: cross Abstract: Memory-based discrete vision-language navigation (VLN) agents must act under partial observability, yet even strong frozen backbones remain vulnerable at test time. Two common failure modes are stale historical evidence at memory …

  193. arXiv cs.LG TIER_1 English(EN) · Kai Hu, Akash Bharadwaj, Weichen Yu, Matt Fredrikson ·

    Steal the Patch Size: Adversarially Manipulate Vision-Language Models

    arXiv:2607.00174v1 Announce Type: cross Abstract: We present a black-box model-stealing attack that recovers private vision-tokenizer configurations of deployed vision-language models (VLMs), including the visual patch size and input preprocessing pipeline. The key idea is a task…

  194. arXiv cs.AI TIER_1 English(EN) · Arpita Nema, Hanwei Zhu, Xi Zhang, Weisi Lin ·

    LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models

    arXiv:2607.01086v1 Announce Type: cross Abstract: The evaluation of long-term video quality understanding remains an open challenge for large vision-language models (LVLMs). Existing video quality benchmarks predominantly focus on short clips and isolated distortions, overlooking…

  195. arXiv cs.LG TIER_1 English(EN) · Seokhee Jin, Changhwan Sung, Sunung Mun, Hoyoung Kim, Jungseul Ok ·

    AdaBoosting Text Prompts for Vision-Language Models

    arXiv:2607.00684v1 Announce Type: new Abstract: The classification accuracy of pretrained Vision-Language Models (VLMs) relies on the quality of the text prompts. Handcrafted templates and Large Language Model (LLM)-generated descriptions not only make predictions more interpreta…

  196. Hugging Face Daily Papers TIER_1 English(EN) ·

    AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

    Vision-Language Models struggle with domain adaptation in specialized spatio-temporal video grounding tasks, highlighting limitations in zero-shot generalization and in-context learning capabilities.

  197. arXiv cs.AI TIER_1 English(EN) · Weisi Lin ·

    LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models

    The evaluation of long-term video quality understanding remains an open challenge for large vision-language models (LVLMs). Existing video quality benchmarks predominantly focus on short clips and isolated distortions, overlooking the temporal continuity, cumulative degradation, …

  198. arXiv cs.AI TIER_1 English(EN) · Jie Mei ·

    DART-VLN: Test-Time Memory Decay and Anti-Loop Regularization for Discrete Vision-Language Navigation

    Memory-based discrete vision-language navigation (VLN) agents must act under partial observability, yet even strong frozen backbones remain vulnerable at test time. Two common failure modes are stale historical evidence at memory readout and inefficient local backtracking during …

  199. Hugging Face Daily Papers TIER_1 English(EN) ·

    DART-VLN: Test-Time Memory Decay and Anti-Loop Regularization for Discrete Vision-Language Navigation

    Memory-based discrete vision-language navigation (VLN) agents must act under partial observability, yet even strong frozen backbones remain vulnerable at test time. Two common failure modes are stale historical evidence at memory readout and inefficient local backtracking during …

  200. arXiv cs.AI TIER_1 English(EN) · Florian Buettner ·

    LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives

    Vision-language pretraining remains dominated by contrastive objectives, whereas vision-only self-supervised learning has largely adopted non-contrastive methods. At the same time, the role of vision-language encoders has shifted: they are increasingly deployed not as zero-shot c…

  201. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Xiangxiang Chu ·

    M2Note: Continual Evolution of Vision Language Models via Mistake Notebook Learning

    Vision Language Models (VLMs) have demonstrated remarkable capabilities in multimodal reasoning tasks, yet they still suffer from recurring failures, such as skipping key visual checks, misapplying domain rules, and hallucinating unsupported concepts. Most existing solutions rely…

  202. arXiv cs.LG TIER_1 English(EN) · Jungseul Ok ·

    AdaBoosting Text Prompts for Vision-Language Models

    The classification accuracy of pretrained Vision-Language Models (VLMs) relies on the quality of the text prompts. Handcrafted templates and Large Language Model (LLM)-generated descriptions not only make predictions more interpretable, but also enable reuse of the same prompts a…

  203. arXiv cs.LG TIER_1 English(EN) · Aditi Naiknaware, Salimeh Sekeh ·

    T-QPM: Enabling Temporal Out-Of-Distribution Detection and Domain Generalization for Vision-Language Models in Open-World

    arXiv:2603.18481v2 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection remains a critical challenge in open-world learning, where models must adapt to evolving data distributions. While recent vision-language models (VLMS) like CLIP enable multimodal OOD de…

  204. arXiv cs.AI TIER_1 English(EN) · Ta Duc Huy, Trang Nguyen, Townim Chowdhury, Ankit Yadav, Minh-Son To, Zhibin Liao, Johan W. Verjans, Vu Minh Hieu Phan ·

    Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?

    arXiv:2606.31407v1 Announce Type: cross Abstract: Vision-language models can produce confident answers on visually ambiguous inputs, resulting in biased predictions. Common entropy-based methods, such as Semantic Entropy (SE), rely on output diversity. Yet our analysis shows that…

  205. arXiv cs.AI TIER_1 English(EN) · Nan Li, Albert Gatt, Massimo Poesio ·

    Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue

    arXiv:2606.31719v1 Announce Type: cross Abstract: In collaborative dialogue, shared perception does not guarantee shared interpretation. Mutual understanding must be established through interaction. We investigate whether vision-language models (VLMs) can distinguish what could b…

  206. arXiv cs.CL TIER_1 English(EN) · Kaier Liang, Hengde Dai, Cristian-Ioan Vasile ·

    ViTL: Temporal Logic-Guided Zero-Shot Natural Language Navigation via Vision-Language Models

    arXiv:2606.30696v1 Announce Type: cross Abstract: Enabling robots to follow natural language commands to complete zero-shot long-horizon tasks remains challenging. It requires extracting implicit temporal and logical constraints from natural language commands and executing multip…

  207. arXiv cs.LG TIER_1 English(EN) · Cl\'ement Fuchs, Tim Bary, Beno\^it Macq ·

    Localized Conformal Prediction for Image Classification with Vision-Language Models

    arXiv:2606.31577v1 Announce Type: cross Abstract: Conformal predictions have attracted significant attention in the field of uncertainty quantification, mainly because of their strong marginal coverage guarantees. Full conditional guarantee is not an attainable goal, a well known…

  208. arXiv cs.AI TIER_1 English(EN) · Massimo Poesio ·

    Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue

    In collaborative dialogue, shared perception does not guarantee shared interpretation. Mutual understanding must be established through interaction. We investigate whether vision-language models (VLMs) can distinguish what could be shared from what has been shared between dialogu…

  209. arXiv cs.CL TIER_1 English(EN) · Vu Minh Hieu Phan ·

    Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?

    Vision-language models can produce confident answers on visually ambiguous inputs, resulting in biased predictions. Common entropy-based methods, such as Semantic Entropy (SE), rely on output diversity. Yet our analysis shows that overconfident visual embeddings suppress output d…

  210. arXiv cs.AI TIER_1 English(EN) · Minh-Loi Nguyen, Nghiem Tuong Diep, Hung Khang Nguyen, Minh Le, Doanh Le Thien, Hoang H. Tran, Dung D. Le, Vu N. Duong, Daniel Sonntag, An Thai Le, Duy Minh Ho Nguyen, Vien Anh Ngo, Tran Van Nhiem ·

    RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

    arXiv:2606.28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning. However, evaluating these videos is challenging: visually realistic outputs often violate physical laws, temporal consis…

  211. arXiv cs.LG TIER_1 English(EN) · Kwun Ho Ngan, Saman Sadeghi Afgeh, Joe Townsend, Artur d'Avila Garcez ·

    Contrastive vision-language learning with paraphrasing and negation

    arXiv:2511.16527v2 Announce Type: replace-cross Abstract: Contrastive vision-language models continue to be the dominant approach for image-text retrieval. Contrastive Language-Image Pre-training (CLIP) trains two neural networks to align their image and text embeddings in a shar…

  212. arXiv cs.LG TIER_1 English(EN) · Haitao Wu, Qirui Zhang, Zhouheng Yao, Shangquan Sun, Qihao Zheng, Mianxin Liu, Chi Zhang, Wanli Ouyang, Chunfeng Song, Changqing Zhang, Jiamin Wu ·

    BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language

    arXiv:2606.30319v1 Announce Type: cross Abstract: Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged as a critical frontier in neuroscience. However, existing approaches predominantly treat brain encoding and decodi…

  213. arXiv cs.CL TIER_1 English(EN) · Hee-Seon Kim, Minbeom Kim, Seokil Ham, Changick Kim ·

    Exploiting Vision Encoder Vulnerabilities for Universal Adversarial Perturbations on Large Vision-Language Models

    arXiv:2412.08108v3 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable performance on multimodal tasks but remain highly vulnerable to small adversarial perturbations in input images. Existing attacks typically target the vision en…

  214. arXiv cs.CL TIER_1 English(EN) · Matteo Farina, Vishaal Udandarao, Thao Nguyen, Selim Kuzucu, Maximilian B\"other, Andreas Hochlehnert, Adhiraj Ghosh, Marianna Nezhurina, Karsten Roth, Joschka Struber, Yuhui Zhang, Sebastian Dziadzio, Elaine Sui, Soumya Jahagirdar, Dhruba Ghosh, Hasan H… ·

    DataComp-VLM: Improved Open Datasets for Vision-Language Models

    arXiv:2606.28551v1 Announce Type: cross Abstract: Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curation strategies. We introduce DataComp for VLMs (DC…

  215. arXiv cs.AI TIER_1 English(EN) · Simon Roschmann, Paul Krzakala, Sonia Mazelet, Quentin Bouniot, Zeynep Akata ·

    SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport

    arXiv:2602.23353v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical model of the world. Recent work exploits this convergence by aligning frozen pretraine…

  216. arXiv cs.AI TIER_1 English(EN) · Chen Yang, Yuhao Wei, Ze Xu, Ziheng Zou, Shuang Liang, Delin Ouyang, Lingfeng Qi, Jie Li, Guofa Li ·

    LWDrive: Layer-Wise World-Model-Guided Vision-Language Model Planning for Autonomous Driving

    arXiv:2606.29879v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning. However, trajectories directly generated by VLMs often encode only coarse driving…

  217. arXiv cs.AI TIER_1 English(EN) · Rahul Chowdhury, Timothy A Rupprecht, Xuan Shen, Pu Zhao, Yanzhi Wang ·

    ScAle: Attention Head Scaling as a Minimal Adapter for Spatial Reasoning in Vision Language Models

    arXiv:2606.29579v1 Announce Type: cross Abstract: Spatial reasoning remains a persistent challenge for many vision language models (VLMs), and improving it typically requires fine-tuning with substantial additional parameters. Our preliminary analysis reveals that rescaling activ…

  218. arXiv cs.AI TIER_1 English(EN) · Jongoh Jeong, Sun-Kyung Lee, Kuk-Jin Yoon ·

    Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation

    arXiv:2606.29464v1 Announce Type: cross Abstract: Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can efficiently train contrastive vision-language models under strict data and compute budgets. Most…

  219. arXiv cs.AI TIER_1 English(EN) · Xiao Wang, Liye Jin, Dan Xu, Yuehang Li, Lan Chen, Yaowei Wang, Yonghong Tian, Jin Tang ·

    Dynamic Parsing and Updating Natural Language Specification using VLMs for Robust Vision-Language Tracking

    arXiv:2606.29357v1 Announce Type: cross Abstract: Vision-language tracking guided by natural language specifications leverages high-level semantic cues of target objects to substantially boost tracking accuracy and robustness. Existing studies have verified that adaptively optimi…

  220. arXiv cs.AI TIER_1 English(EN) · Yichen Guo, Kai Tang, Fenglai Lin, Yiding Sun, Dongshuo Zhang, Wenya Wang, Lin William Cong, Shanghang Zhang ·

    FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models

    arXiv:2606.29431v1 Announce Type: new Abstract: Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucination, generating content inconsistent with the input image. Recent studies attribute this to the dominance of language …

  221. arXiv cs.AI TIER_1 English(EN) · Guanglong Sun, Shuang Cui, Bo Lei, Liyuan Wang, Zihan Zhai, Hongwei Yan, Hang Su, Jun Zhu, Yi Zhong ·

    ComMem: Complementary Memory Systems for Test-Time Adaptation of Vision-Language Models

    arXiv:2606.28719v1 Announce Type: new Abstract: Test-time adaptation (TTA) of vision-language models (VLMs) is essential for their robust deployment in dynamic, real-world environments. However, existing TTA methods often adapt locally without accumulating knowledge over time, or…

  222. Hugging Face Daily Papers TIER_1 English(EN) ·

    Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue

    Vision-language models struggle to distinguish between shared and interpreted visual information in dialogue, relying on static map cues rather than dynamic grounding processes.

  223. Hugging Face Daily Papers TIER_1 English(EN) ·

    3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance

    3D HAMSTER framework enhances robot manipulation by integrating a vision-language model with depth encoding to generate metrically accurate 3D trajectories for point cloud-based control policies.

  224. arXiv cs.LG TIER_1 English(EN) · Jiamin Wu ·

    BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language

    Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged as a critical frontier in neuroscience. However, existing approaches predominantly treat brain encoding and decoding as isolated tasks, relying heavily on unimodal …

  225. Hugging Face Daily Papers TIER_1 English(EN) ·

    Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents

    Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dynamo, a training-free framework that adapts a frozen VLM without any weight updates. On a small labeled training subset, the agent inspects i…

  226. arXiv cs.AI TIER_1 Deutsch(DE) · Dong-Jae Lee, Sunghyun Baek, Junmo Kim ·

    IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models

    arXiv:2604.00757v2 Announce Type: replace-cross Abstract: Large Vision Language Models show impressive performance across image and video understanding tasks, yet their computational cost grows rapidly with the number of visual tokens. Existing token pruning methods mitigate this…

  227. arXiv cs.CL TIER_1 English(EN) · Jaume Guasch-Mart\'i, Enrique Lopez-Cuena, Mart\'in Su\'arez-Fern\'andez, Jordi Bayarri-Planas, Anna Arias-Duart, Dario Garcia-Gasulla ·

    Aloe-Vision: Robust Vision-Language Models for Healthcare

    arXiv:2606.27500v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) specialized in healthcare are emerging as a promising research direction due to their potential impact in clinical and biomedical applications. However, progress is constrained by the scarcity …

  228. arXiv cs.CL TIER_1 English(EN) · Niclas Lietzow, Danielle Bitterman, Carsten Eickhoff, William Rudman, Michal Golovanevsky ·

    Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models

    arXiv:2606.28273v1 Announce Type: new Abstract: Vision-language models must reconcile visual evidence with memorized world knowledge when the two conflict. How they resolve this conflict shapes the reliability of multimodal systems, yet prior work characterizes it behaviorally wi…

  229. Hugging Face Daily Papers TIER_1 English(EN) ·

    BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language

    BrainJanus represents the first unified brain model integrating brain, vision, and language through a shared Omni space, enabling bidirectional mapping between neural activity and sensory stimuli via a tokenized representation and autoregressive architecture.

  230. Hugging Face Daily Papers TIER_1 English(EN) ·

    Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation

    Vision-language dataset distillation method using rank-aware hyperbolic alignment to optimize synthetic image-text pairs for efficient contrastive model training while preserving modality-specific diversity.

  231. arXiv cs.CL TIER_1 English(EN) · Michal Golovanevsky ·

    Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models

    Vision-language models must reconcile visual evidence with memorized world knowledge when the two conflict. How they resolve this conflict shapes the reliability of multimodal systems, yet prior work characterizes it behaviorally without a component-level causal account. We combi…

  232. arXiv cs.AI TIER_1 English(EN) · Chenyang Zhang, Anqi Dong, Guangming Zhu, Nuoye Xiong, Siyuan Wang, Lin Mei, Liang Zhang ·

    Bridging Vision and Language Concepts through Optimal Transport Semantic Flow

    arXiv:2606.26891v1 Announce Type: cross Abstract: Concept Bottleneck Models (CBMs) promise transparent reasoning by predicting through human-interpretable concepts, yet their effectiveness fundamentally depends on how well visual and textual representations are aligned or matched…

  233. arXiv cs.CL TIER_1 English(EN) · Byung-Kwan Lee, Ryo Hachiuma, Yong Man Ro, Yu-Chiang Frank Wang, Yueh-Hua Wu ·

    GenRecal: Generation after Recalibration from Large to Small Vision-Language Models

    arXiv:2506.15681v4 Announce Type: replace Abstract: Recent advancements in vision-language models (VLMs) have leveraged large language models (LLMs) to achieve performance on par with closed-source systems like GPT-4V. However, deploying these models in real-world scenarios, part…

  234. arXiv cs.AI TIER_1 English(EN) · Kai Tang, Jinhao You, Yichen Guo, Yiding Sun, Dongxu Zhang, Wenya Wang, Hanze Li, Tao Luo, Renyuan Li, Xiande Huang ·

    Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models

    arXiv:2505.12343v2 Announce Type: replace-cross Abstract: Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucinations, where generated content is inconsistent with the input image. Existing training-free hallucination mit…

  235. arXiv cs.AI TIER_1 English(EN) · Haoxiang Sun, Tao Wang, Li Yuan, Jian Zhao, Jiancheng Lv ·

    From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

    arXiv:2606.26196v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, especially following the introduction of models such as OpenAI's O-series and DeepSeek's R-se…

  236. Hugging Face Daily Papers TIER_1 English(EN) ·

    DataComp-VLM: Improved Open Datasets for Vision-Language Models

    DataComp for VLMs (DCVLM) establishes a comprehensive benchmark for evaluating data curation strategies in vision-language models, demonstrating that data mixing rather than filtering significantly improves model performance at scale.

  237. arXiv cs.CL TIER_1 English(EN) · Dario Garcia-Gasulla ·

    Aloe-Vision: Robust Vision-Language Models for Healthcare

    Large Vision-Language Models (LVLMs) specialized in healthcare are emerging as a promising research direction due to their potential impact in clinical and biomedical applications. However, progress is constrained by the scarcity of high-quality medical multimodal data, concerns …

  238. arXiv cs.AI TIER_1 English(EN) · Hui Li ·

    Steering Vision-Language Models with Joint Sparse Autoencoders

    Sparse Autoencoders (SAEs) have shown promise for analyzing language models, but applying them to vision-language models (VLMs) often yields representations that are difficult to use as controllable cross-modal steering directions. We introduce the Joint Sparse Autoencoder (JSAE)…

  239. arXiv cs.AI TIER_1 English(EN) · Ahmad Algadhi, Ahmed Alzuhair, Omar Alkhulaif, Muzammil Behzad ·

    The Professor: Multi-Teacher Unsupervised Prompt Distillation for Vision-Language Models

    arXiv:2606.23897v1 Announce Type: cross Abstract: Prompt distillation compresses large vision-language models (VLMs) such as CLIP into lightweight student models by matching teacher predictions on unlabeled domain images. PromptKD (CVPR 2024) established this paradigm with a sing…

  240. Hugging Face Daily Papers TIER_1 English(EN) ·

    4DVLT: Dynamic Scene Understanding with Worldline-Centered Vision-Language Tracking

    4D dynamic scene understanding requires grounding language to a persistent worldline that binds identity, metric 3D motion, and synchronized multi-view 2D projections. Existing paradigms capture only part of this structure: large multimodal models reason over rich visual evidence…

  241. arXiv cs.CL TIER_1 English(EN) · Yusuf Salcan (Computer Vision Group, University of Freiburg, Germany, CRIION-AI Lab, Freiburg, Germany), Simon Ging (Computer Vision Group, University of Freiburg, Germany, Adaptive & Agentic AI), Robin Schirrmeister (Department of Radiology, Medical Cen… ·

    Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology

    arXiv:2606.20477v1 Announce Type: cross Abstract: We study how to train visually grounded vision-language models (VLMs) for radiology without manual spatial annotations. We introduce RefRad2D, a large-scale bilingual (German/English) dataset of 1.2M CT and MR image-text pairs der…

  242. arXiv cs.AI TIER_1 English(EN) · Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh ·

    Vision-language models for chest radiography do not always need the image

    arXiv:2606.17710v1 Announce Type: cross Abstract: Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image. That inference is unsafe: a model exploiting finding-name priors scores like one that reads…

  243. arXiv cs.CL TIER_1 English(EN) · Soroosh Tayebi Arasteh ·

    Vision-language models for chest radiography do not always need the image

    Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image. That inference is unsafe: a model exploiting finding-name priors scores like one that reads the scan, and no standard benchmark separates the…

  244. arXiv cs.CV TIER_1 English(EN) · Joseph Bingham ·

    A Dynamic-Semantics Framework for Grounding Human Referring Expressions in Visual Perceptual Data

    arXiv:2608.08663v1 Announce Type: cross Abstract: Humans converge on shared names for novel, hard-to-describe objects through repeated interaction, a process psycholinguists call lexical entrainment. Leading vision-language models fail at this: recent empirical work documents tha…

  245. arXiv cs.CV TIER_1 English(EN) · Hongjin Ji, Guoyang Xia, Luoyang Sun, Fangxiang Feng, Lei Ren ·

    VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction

    arXiv:2608.09448v1 Announce Type: cross Abstract: Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains difficult to use reliably in closed-loop manipulation. A shared adaptation space …

  246. arXiv cs.CV TIER_1 English(EN) · Ningxin Pan, Hanyu Li, Yehui Tang ·

    ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making

    arXiv:2608.07584v1 Announce Type: new Abstract: Vision-language models (VLMs) have made rapid progress in visual perception and increasingly support real-world tasks that depend on images. Many such tasks, however, require more than rec- ognizing what an image contains: a model m…

  247. arXiv cs.CV TIER_1 English(EN) · Sajjad Ghiasvand, Yifan Yang, Mahnoosh Alizadeh, Ramtin Pedarsani ·

    ZOMP: Zeroth-Order Multi-Modal Prompt Tuning for Vision-Language Models

    arXiv:2608.08060v1 Announce Type: new Abstract: Fine-tuning vision-language models such as CLIP typically requires backpropagation (BP) through the full model, which is infeasible when only forward-pass access is available, as is common for memory-constrained edge devices and pro…

  248. arXiv cs.CV TIER_1 English(EN) · Issa Sugiura, Keito Sasagawa, Keisuke Nakao, Koki Maeda, Ziqi Yin, Zhishen Yang, Shuhei Kurita, Yusuke Oda, Ryoko Tokuhisa, Daisuke Kawahara, Naoaki Okazaki ·

    Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models

    arXiv:2604.02048v2 Announce Type: replace Abstract: Developing vision-language models (VLMs) that generalize across diverse tasks requires large-scale training datasets with diverse content. In English, such datasets are typically constructed by aggregating and curating numerous …

  249. arXiv cs.CV TIER_1 English(EN) · Xuan Yao, Yuze Zhu, Junyu Gao, Zongmeng Wang, Changsheng Xu ·

    SC$^{2}$-WM: A Self-Correcting World Model with Closed-Loop Feedback for Vision-and-Language Navigation in Continuous Environments

    arXiv:2608.07548v1 Announce Type: cross Abstract: Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to make fine-grained navigation decisions under partial observability. However, most existing methods rely on open-loop execution, lacking mechanis…

  250. arXiv cs.CV TIER_1 English(EN) · Zhewei Zhang, Puyue Wang, Guanren Qiao, Yijie Weng, Jiawei Hu, Guo Li, Lujia Wang, Junyan Wang, Tao Gu, Hongliang Lu, Guiliang Liu, Hong Jia, Xinhu Zheng ·

    LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding

    arXiv:2608.07596v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models transform representations from pretrained vision-language models (VLMs) into robot actions, yet the interface that routes intermediate VLM features into action decoders remains underexplored. Ex…

  251. arXiv cs.CV TIER_1 English(EN) · Ziheng Liu, Quantao Yang ·

    TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models

    arXiv:2608.07314v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are commonly adapted to downstream manipulation tasks via supervised fine-tuning (SFT) or online reinforcement learning (RL) post-training. SFT is prone to distribution mismatch, and existing RL…

  252. arXiv cs.CV TIER_1 English(EN) · Harisankar Babu, Benjamin Coors, Christopher Lang, Hendrik Berkemeyer, Tamim Asfour, Simon Foell ·

    Depth-Wise Probing and Pruning of the Planning Token in a Driving Vision-Language-Action Model

    arXiv:2608.07361v1 Announce Type: cross Abstract: Vision-language-action (VLA) models route driving decisions through a deep language model, but it is unclear how much of that depth the action itself requires. We study a representative driving VLA whose entire plan is carried by …

  253. arXiv cs.CV TIER_1 English(EN) · Guiyu Zhao, Longteng Guo, Yanghong Mei, Zilin Zhu, Yu Zhang, Bin Cao, Mingming Yu, Xingjian He, Jie Jiang, Jing Liu ·

    AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

    arXiv:2608.06729v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camer…

  254. arXiv cs.CV TIER_1 English(EN) · Kexin Ma, Jing Xiao, Chaofeng Chen, Geyong Min, Guibo Zhu, Jinqiao Wang, Liang Liao ·

    Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models

    arXiv:2604.11240v2 Announce Type: replace Abstract: Token pruning has emerged as an effective approach to reduce the substantial computational overhead of Large Vision-Language Models (LVLMs) by discarding less informative visual tokens while preserving performance. However, exis…

  255. arXiv cs.CV TIER_1 English(EN) · Thong Nguyen, Vinh-Hien Do, Quynh Vo, Cong-Duy Nguyen, See-Kiong Ng ·

    DynaPix: Can Vision-Language Models Identify the Exact Future?

    arXiv:2608.05505v1 Announce Type: new Abstract: Acting in a physical scene requires knowing its real later state, not a plausible one. Current evaluations often accept words or a realistic-looking image, so the predicted state is never checked against the true one. We introduce D…

  256. arXiv cs.CV TIER_1 English(EN) · Mahyar Ghazanfari, Amin Tabrizian, Arsyi Aziz, Binshuai Wang, Peng Wei ·

    A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval

    arXiv:2608.05260v1 Announce Type: new Abstract: Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from detailed textual descriptions. While methods such as Long-CLIP extend the token l…

  257. arXiv cs.CV TIER_1 English(EN) · Xi Xiao, Xingjian Li, Cheng Han, Tianyang Wang, Lin Zhao, Yunbei Zhang, Guosheng Hu, Runmin Jiang, Xi Li, Xiao Wang, Min Xu ·

    Adapting Vision Foundation Models with Cascaded Semantics

    arXiv:2608.05393v1 Announce Type: new Abstract: Prompt tuning, a leading parameter-efficient adaptation paradigm in NLP, has recently been extended to computer vision. Visual prompt tuning (VPT) adapts pre-trained vision transformers (ViTs) by updating a small set of additional p…

  258. arXiv cs.CV TIER_1 English(EN) · Jihoon Oh, Kento Kawaharazuka, Kei Okada ·

    VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances

    arXiv:2608.05215v1 Announce Type: cross Abstract: Learning manipulation skills from human videos is promising for scalable robot learning. However, the embodiment mismatch between humans and robots makes this challenging. One promising solution is to learn object-centric actionab…

  259. arXiv cs.CV TIER_1 English(EN) · Jingyan Jiang, Yaru Sun, Xiao Chen, Jiazhen Huang, Caiting Li, Zhijian He, Yin Chen, Pingting Hao ·

    Respect Your Zero-Shot Uncertainty: Conservative Calibration for Test-Time-Adapted Vision-Language Models

    arXiv:2608.05945v1 Announce Type: new Abstract: Test-time adaptation (TTA) can improve the recognition accuracy of vision-language models under distribution shift, but often degrades calibration, making predictive confidence unreliable for downstream decision-making. Many existin…

  260. arXiv cs.CV TIER_1 English(EN) · Behnam Raoufi, Hossein Sharify, Mohamad Mahdee Ramezanee, Khosrow Hajsadeghi, Saeed Bagheri Shouraki ·

    CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision

    arXiv:2512.22969v2 Announce Type: replace Abstract: Conventional object detectors rely on cross-entropy classification, which can be vulnerable to class imbalance and label noise. We propose CLIP-Joint-Detect, a simple and detector-agnostic framework that integrates CLIP-style co…

  261. arXiv cs.CV TIER_1 English(EN) · Yan Zhang, Yinan Wu, Haoran Duan, Jungong Han ·

    CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention

    arXiv:2608.04396v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have driven significant progress in robotic manipulation, yet they fundamentally struggle with the vision-override phenomenon. Driven by the severe modality imbalance between dense visual streams …

  262. arXiv cs.CV TIER_1 English(EN) · Weihan Cai, Hao Tan, Zichang Tan, Jun Wan, Xinping Gao ·

    Unleashing the Potential of Vision-Language Models for Generalizable AI-Generated Image Detection

    arXiv:2608.04935v1 Announce Type: new Abstract: Recent work has shown that a simple linear probe on frozen representations from modern vision foundation models (VFMs) can achieve state-of-the-art AIGI detection performance, substantially outperforming specialized detectors in cha…

  263. arXiv cs.CV TIER_1 English(EN) · Hao Dou, Ruiwen Tian ·

    When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware

    arXiv:2608.03649v1 Announce Type: new Abstract: Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead, shared work, and the operators each policy can avoid. A stage-level decompositio…

  264. arXiv cs.CV TIER_1 English(EN) · Yaozhi Wen, Jialong Guo, Zhenliang Ni, Han Shu, Xinghao Chen ·

    SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models

    arXiv:2608.03580v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have demonstrated remarkable performance in processing and understanding both text and images, their large parameter sizes lead to significant computational overhead, limiting their deployment on …

  265. arXiv cs.CV TIER_1 English(EN) · Zihan Wang, Tong Liu, Zhiwei Wang, Tao Huang, Wentao Jiang, Sihan Ma, Shanshan Ye, Xiaohui Yang, Jing Zhang ·

    LocAnyMed: Vision-Language Grounding for Multimodal Medical Images

    arXiv:2608.03322v1 Announce Type: new Abstract: Medical visual grounding connects free-form clinical queries to spatial evidence in medical images and is an important component of interpretable medical artificial intelligence. However, general-purpose grounding models are predomi…

  266. arXiv cs.CV TIER_1 English(EN) · Xiangyun Huang, Xiangchen Wang, Runfeng Lin, Yihao Xu, Kangyu Huang, Jiang Hengchen, Xiwang Dong, Lin Jiarong ·

    From Routes to Steps: Separating Semantic Progress from Local Execution in Vision-and-Language Navigation

    arXiv:2608.03143v1 Announce Type: new Abstract: Vision-and-Language Navigation (VLN) requires an agent to follow a route-level instruction by executing its constituent steps from egocentric visual observations. Existing VLM-based navigators typically supervise both capabilities t…

  267. arXiv cs.CV TIER_1 English(EN) · Lucy Lin, Ayush Jain, Yifan Liu, Katerina Fragkiadaki ·

    Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

    arXiv:2608.02980v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to frame-centric tokenization and limited context windows. 3D geometry provides a na…

  268. arXiv cs.CV TIER_1 English(EN) · Amir Sabbaghziarani, Mohammadsajad Abavisani, Sergey Plis ·

    Confident but Unreliable: A Behavioral Safety Audit of Vision-Language Models on Brain MRI

    arXiv:2608.02790v1 Announce Type: new Abstract: Vision-language models (VLMs), including medical specialists, are increasingly proposed for medical imaging, yet their stated confidence is rarely evaluated separately from correctness. We use brain MRI as a controlled, high-stakes …

  269. arXiv cs.CV TIER_1 English(EN) · Gengyuan Liu, Nanzhou Wang, Chang Liu, Qinwen Wu, Zhenhao Wang, Jiacong Wang, Bokui Chen, Xiangyang Ji ·

    MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization

    arXiv:2608.03885v1 Announce Type: new Abstract: Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-time adaptation (TTA) via Low-Rank Adaptation offers a parameter-efficient solutio…

  270. arXiv cs.CV TIER_1 English(EN) · Guoliang You, Haifan Gong, Xiaomeng Chu ·

    Semantically Calibrated Evidence Composition for CT Vision-Language Learning

    arXiv:2608.00239v1 Announce Type: new Abstract: Learning transferable representations from CT-report pairs requires combining whole-volume context with anatomy-specific evidence. Existing methods typically emphasize either global CT-report alignment or fine-grained anatomy-level …

  271. arXiv cs.CV TIER_1 English(EN) · Yanbin Hu, Jin Cui, Jun Ye, Jiepeng Zhou, Jiangcheng Song, Boran Zhao, Pengju Ren ·

    Distill What RGB Can Recover: Privileged 3D Evidence for RGB-Only Vision-Language Models

    arXiv:2608.00110v1 Announce Type: new Abstract: 3D scene understanding requires reasoning about entity existence, spatial layout, and object relations, yet RGB images alone often provide insufficient 3D cues. Existing 3D-VLMs commonly rely on depth or 3D-position-aware inputs at …

  272. arXiv cs.CV TIER_1 English(EN) · Qi Lv, Jianming Xing, Zhao Yang, Mingyuan Yao, Yinan Shi, Yawei Jueluo, Mike Zheng Shou, Xiang Deng ·

    Hermite Curves as Trajectory Priors for Vision-Language-Action Models

    arXiv:2608.01265v1 Announce Type: cross Abstract: Despite recent progress in Vision-Language-Action (VLA) models for robotic manipulation, the action chunk remains a weakly structured interface. Existing work typically flatten each chunk into per-timestep controls, relying on imp…

  273. arXiv cs.CV TIER_1 English(EN) · Brian Song, Michael A. Lepori, Ellie Pavlick ·

    Linguistic Context Recodes Visual Representations in Vision-Language Models

    arXiv:2608.00035v1 Announce Type: cross Abstract: Goal-directed visual processing is a hallmark of human visual intelligence, resulting in representations that support downstream tasks such as categorization or search. Though vision-language models (VLMs) are often faced with the…

  274. arXiv cs.CV TIER_1 English(EN) · Yan Huang, Guowei Wang, Xu Wang, Kangjun Liu, Xin Lin ·

    Local Margin Restoration for Test-Time Adaptation of Vision-Language Models

    arXiv:2608.02216v1 Announce Type: new Abstract: Vision-language models (VLMs) such as CLIP exhibit remarkable zero-shot capabilities, yet their performance frequently degrades sharply under unexpected test-time distribution shifts. While Test-Time Adaptation (TTA) offers a promis…

  275. arXiv cs.CV TIER_1 English(EN) · Xuanhui Lin, Junhao Dong, Mingrong Gong, Yucheng Chen, Xinghua Qu, Yew-Soon Ong ·

    Two Sides of the Same Coin: Co-Evolving Search for Cross-Task Attacks on Vision-Language Models

    arXiv:2608.02137v1 Announce Type: new Abstract: Vision-language models (VLMs) exhibit strong generalization across multimodal tasks but remain vulnerable to adversarial perturbations. Existing attacks typically follow single-trajectory gradient optimization or task-specific objec…

  276. arXiv cs.CV TIER_1 English(EN) · Hongjie Zhou, Shiqin Wang, Haoyang Chen, Haonan Guo, Di Wang, Juhua Liu, Fu Lin, Yong Luo ·

    RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos?

    arXiv:2608.02039v1 Announce Type: new Abstract: Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, actions, interactions, and scene changes that cannot be captured by isolated images.…

  277. arXiv cs.CV TIER_1 English(EN) · Landi He, Mingde Yao, Shawn Young, Lijian Xu ·

    DiffPrune: differentiable information throttling for token pruning in vision-language models

    arXiv:2608.01985v1 Announce Type: new Abstract: Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. The key is to learn a score that measures whether a token is useful. Existing methods typically rely on Gumbel…

  278. arXiv cs.CV TIER_1 English(EN) · Hai Nguyen, Tung Vu, Cong Tran ·

    SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models

    arXiv:2608.01709v1 Announce Type: new Abstract: Vision-language models (VLMs) achieve strong semantic understanding but remain unreliable in metric spatial reasoning, particularly when queries require comparing multiple instances of the same object category. We study this problem…

  279. arXiv cs.CV TIER_1 English(EN) · Yu Chen, Xiaohong Li, Xiaole Wang, Jianjin Zhang, Jun Sun, Yafeng Deng ·

    CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models

    arXiv:2608.01644v1 Announce Type: new Abstract: In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the prefill stage to rise sharply. Such visual sequences are highly redundant along the…

  280. arXiv cs.CV TIER_1 English(EN) · Ashfak Yeafi, Mehedi Hasan, Md Khairul Islam ·

    Linear Multi-Timescale Retention as a Memory-Efficient Vision-Language Bridge

    arXiv:2608.01614v1 Announce Type: new Abstract: Vision-Language Models (VLMs) face a critical computational bottleneck when processing high-resolution imagery due to the $O(N^2)$ memory complexity of Softmax Multi-Head Attention (MHA). While substituting MHA with independent Mult…

  281. arXiv cs.CV TIER_1 English(EN) · Omid Nejati Manzari, Guillaume Lajoie, Hassan Rivaz ·

    Recursive Vision Language Models for General Symbolic Reasoning

    arXiv:2608.01534v1 Announce Type: new Abstract: Hard symbolic-reasoning tasks such as Sudoku, maze pathfinding, and ARC remain challenging for LLMs due to their fixed-depth autoregressive reasoning, which limits systematic search, refinement, and backtracking. While recursive mod…

  282. arXiv cs.CV TIER_1 English(EN) · Puzhuo Zheng, Hasan Kurban ·

    It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling

    arXiv:2608.01207v1 Announce Type: new Abstract: Test-time scaling lifts large language model reasoning by sampling many candidate solutions and selecting among them, yet the same recipe transfers poorly to vision-language models (VLMs): recent work shows that simple majority voti…

  283. arXiv cs.CV TIER_1 English(EN) · Hashmat Shadab Malik, Toluwani Aremu, Samuele Poppi, Muzammal Naseer, Salman Khan ·

    ReACT-CLIP: Response-Aware Test-Time Defense for Vision--Language Models

    arXiv:2608.01067v1 Announce Type: new Abstract: Training-free test-time defenses offer a practical way to improve the adversarial robustness of CLIP-style vision--language models without modifying the pretrained model. However, their correction strength is typically fixed for a n…

  284. arXiv cs.CV TIER_1 English(EN) · Ziang Wu, Peng Jin, Qishen Yin, Munan Ning, Hao Li, Peizhen Zhang, Li Yuan ·

    Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts

    arXiv:2608.00574v1 Announce Type: new Abstract: Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this token mix. We call the standard token-level Switch auxiliary loss Std-Aux. Std-…

  285. arXiv cs.CV TIER_1 English(EN) · Meibo Hu, Guohao Sun, Annemarie D. Ross, Sheng Li, Zhiqiang Tao ·

    Attention-Steered Vision-Language Models for Sign Language Translation

    arXiv:2608.00235v1 Announce Type: new Abstract: Vision-language models (VLMs) have emerged as a powerful framework for multimodal video understanding. However, they remain limited in the sign language translation task, where we identify a key failure mode of existing VLMbased tra…

  286. arXiv cs.CV TIER_1 English(EN) · Li Lin, Wujun Xu, Weiwei Meng, Kaiwen Xia, Kang Hao Cheong, Shuai Wang ·

    FibVLA: An Efficient Temporal Vision-Language-Action Model with Fibonacci Sampling

    arXiv:2607.29596v1 Announce Type: cross Abstract: Vision-language-action models (VLAs), which leverage the cognition of multimodal information to infer physical-world actions, provide a generalized solution for embodied AI applications. Conventional VLAs usually concentrate on cu…

  287. arXiv cs.CV TIER_1 English(EN) · Enyi Shi, Fei Shen, Chuancheng Shi, Shuyi Miao, Linxia Zhu, Pengyang Shao, Jinhui Tang, Tat-Seng Chua ·

    Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models

    arXiv:2604.08881v2 Announce Type: replace Abstract: With the widespread deployment of vision-language large models (VLLMs), their safety alignment faces dual challenges across languages and modalities. Existing methods model multilingual and multimodal safety separately, overlook…

  288. arXiv cs.CV TIER_1 English(EN) · Ioannis Sarridis, Ioannis Kompatsiaris, Symeon Papadopoulos ·

    Scaling Vision-Language Models Is Not Enough to Mitigate Bias

    arXiv:2607.28211v1 Announce Type: new Abstract: Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly understood at scale. We present the first large-scale empirical study of 194 publicl…

  289. arXiv cs.CV TIER_1 English(EN) · Darsha Udayanga, Pin-Yu Chen, Payel Das, Qiang Ji ·

    Can Vision-Language Models Reason about AI Edits in Images?

    arXiv:2607.28464v1 Announce Type: new Abstract: Detection and localization of AI-tampered images are critical for trustworthy AI, yet modern generative models have made such manipulations increasingly difficult to identify. While traditional binary classifiers can detect image ta…

  290. arXiv cs.CV TIER_1 English(EN) · Jialuo He, Huangxun Chen ·

    Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models

    arXiv:2603.05950v2 Announce Type: replace Abstract: Visual token reduction is critical for accelerating Vision-Language Models (VLMs), since visual inputs are represented as token sequences that introduce substantial computational overhead in the LLM backbone. However, most pruni…

  291. arXiv cs.CV TIER_1 English(EN) · Nguyen Duc Thai, Junhao Dong, Sua Qi Rong, Hua Yu, Yew-Soon Ong ·

    Unifying Adversarially Robust Model Experts in Vision-Language Models

    arXiv:2607.27897v1 Announce Type: new Abstract: Vision-language models (VLMs), such as CLIP, are vulnerable to adversarial attacks, posing a serious problem for real-life applications and deployment. Adversarial fine-tuning emerges as a prominent defense method; however, differen…

  292. arXiv cs.CV TIER_1 English(EN) · Jiasheng Li, Zhong Ji, Yan Zhang, Huihui Li ·

    Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs

    arXiv:2607.27700v1 Announce Type: new Abstract: Large Vision-Language Models (VLMs) suffer from prohibitive inference overhead due to long sequences of visual tokens. However, existing visual token reduction methods mainly improve efficiency by pruning or compressing redundant to…

  293. arXiv cs.CV TIER_1 English(EN) · Yitao Zhu, Mengjun Liu, Yingji Fu, Haowen Pang, Anqi Qiu ·

    MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models

    arXiv:2607.26554v1 Announce Type: new Abstract: Integrating 3D medical images with vision-language models (VLMs) holds substantial promise for computer-aided diagnosis. However, volumetric images generate prohibitively long visual-token sequences with considerable spatial and int…

  294. arXiv cs.CV TIER_1 English(EN) · Yunzhan Fu, Enyu Bao, Xiangyu Shen, Yihao Wu, Chunbo Jiang, Fangli Guan, Liqi Yan ·

    SCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation

    arXiv:2607.26885v1 Announce Type: new Abstract: Vision-language pre-training (VLP) serves as a cornerstone for medical multimodal representation learning. However, existing medical VLP frameworks are often constrained by the limited context windows and shallow representational ca…

  295. arXiv cs.CV TIER_1 English(EN) · Siyao Li, Jiawei Gu, Shuai Liu, Kairui Hu, Zekun Li, Linjie Li, Chengcheng Tang, Po-Chen Wu, Ivan Shugurov, Lingni Ma, Michael Zollhoefer, Sizhe An, Abhay Mittal, Amy Zhao, Ranjay Krishna, Manling Li, Ziwei Liu, Chuan Guo ·

    HumanCLAW: Can Vision-Language Models Act Through a Body?

    arXiv:2607.27180v1 Announce Type: new Abstract: Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a ba…

  296. arXiv cs.CV TIER_1 English(EN) · Hengyi Xie, Chenfei Yao, Xianjin Wu, Xuanyang Xi, Yiping Tang, Di Xu, Yingying Zhu, Dingkang Liang, Xiang Bai, Han Ding ·

    TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

    arXiv:2607.27205v1 Announce Type: new Abstract: Vision-language-action (VLA) models commonly adopt an LLM-centric $V \to L \to A$ pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Alth…

  297. arXiv cs.CV TIER_1 English(EN) · Zhongbin Guo, Jiahe Liu, Wenyu Gao, Yushan Li, Xiaomin He, Chengzhi Li, Ping Jian ·

    LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency

    arXiv:2512.01008v2 Announce Type: replace Abstract: Text-driven 3D reconstruction requires masks that understand free-form instructions and remain stable under viewpoint changes. We present LISA-3D, a two-stage framework that adapts the instruction-following segmenter LISA with g…

  298. arXiv cs.CV TIER_1 English(EN) · Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu, Serge Belongie, V\'esteinn Sn{\ae}bjarnarson ·

    Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

    arXiv:2607.26326v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. However, they often fail on vision-centric tasks, especially when visual evidence c…

  299. arXiv cs.CV TIER_1 English(EN) · Jinkun Zhao, Kui Zhang, Wenjun Wu ·

    TPCD: Tone-Pressure Contrastive Decoding and the Label-Free Gating Bottleneck in Vision-Language Models

    arXiv:2607.26536v1 Announce Type: new Abstract: High-pressure prompts can push vision-language models (VLMs) into unsupported commitments, such as reading illegible text, reporting indeterminate times, or affirming absent objects. This paper asks whether the pressure-induced dist…

  300. arXiv cs.CV TIER_1 English(EN) · Mingkuan Feng, Zhengqi Wen, Jianhua Tao ·

    Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution

    arXiv:2607.26596v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual understanding within a unified transformer architecture. However, fine-tuning all parameters of these models for vi…

  301. arXiv cs.CV TIER_1 English(EN) · Zichao Lin, Yifeng Xie, Bowen Qu, Haiming Wang, Jia Li, Haoning Wu, Yuhao Dong, Zuhao Yang, Jinguo Zhu, Haoyu Lu, Zijia Zhao, Tongtian Yue, Zhangyang Qi, Junwei Yang, Mengfan Dong, Peizhou Cao, Chenzhuang Du, Zaida Zhou, Haotian Yao, Hao Yang, Hongcheng … ·

    PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

    arXiv:2607.24957v1 Announce Type: new Abstract: We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evalua…

  302. arXiv cs.CV TIER_1 English(EN) · Moshiur Farazi, Sameera Ramasinghe, Bekir Sait Ciftler, Mahbub Ahmed Turza, Shafin Rahman ·

    The Gate Always Closes: On Injecting Auxiliary Signals into Frozen Vision-Language Models

    arXiv:2607.23335v1 Announce Type: new Abstract: Auxiliary signal pathways in VLMs are routinely fitted with learnable gates so the optimiser can decide how much of the signal to admit. We find that the optimiser almost always decides on zero: across five injection designs, every …

  303. arXiv cs.CV TIER_1 English(EN) · Ioannis Maniadis Metaxas, Adrian Bulat, Alberto Baldrati, Anestis Zaganidis, Yassine Ouali, Hyeonuk Kim, Georgios Tzimiropoulos ·

    UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models

    arXiv:2607.23373v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) remain bottlenecked by massive computational footprints, precluding their deployment on resource-constrained edge devices. While efforts to compress LVLMs focus heavily on vision token reduction …

  304. arXiv cs.CV TIER_1 English(EN) · Yuqi Liu, Shengju Qian, Tianyuan Qu, Mingxian Lin, Zixuan Wang, Xin Wang, Bei Yu, Jiaya Jia ·

    MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation

    arXiv:2607.23504v1 Announce Type: new Abstract: Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to maintain long-horizon visual history for trajectory consistency while executing actions with low latency. Existing video-based VLN approaches typi…

  305. arXiv cs.CV TIER_1 English(EN) · Shaofei Lei ·

    MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning

    arXiv:2607.24424v1 Announce Type: new Abstract: Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features can discard text, local attributes, and spatial relationships, while high-resoluti…

  306. arXiv cs.CV TIER_1 English(EN) · Riccardo Andrea Izzo, Gianluca Bardaro, Matteo Matteucci ·

    Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models

    arXiv:2603.05147v2 Announce Type: replace Abstract: Current research on Vision-Language-Action (VLA) models predominantly focuses on enhancing generalization through reasoning techniques. While effective, these improvements increase computational complexity and inference latency.…

  307. arXiv cs.CV TIER_1 English(EN) · Daojie Peng, Fulong Ma, Jun Ma ·

    Structured Observation Language for Efficient and Generalizable Vision-Language Navigation

    arXiv:2603.27577v2 Announce Type: replace Abstract: Vision-Language Navigation (VLN) requires an embodied agent to navigate complex environments by following natural language instructions, which typically demands tight fusion of visual and language modalities. Existing VLN method…

  308. arXiv cs.CV TIER_1 English(EN) · Yihao Wu, Chenyi Xu, Liqi Yan, Chenhuan Cai, Geyong Min, Bin Lin, Fangli Guan, Jianhui Zhang, Pan Li ·

    Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation

    arXiv:2607.23181v1 Announce Type: new Abstract: Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multimodal large models and world-model-based methods have…

  309. arXiv cs.CV TIER_1 English(EN) · Futa Waseda, Saku Sugawara, Isao Echizen ·

    Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models

    arXiv:2507.16257v2 Announce Type: replace Abstract: Defending pre-trained vision-language models (VLMs), such as CLIP, against adversarial attacks is crucial, as these models are widely used in diverse zero-shot tasks, including image classification. However, existing adversarial…

  310. arXiv cs.CV TIER_1 English(EN) · Pau de Jorge, C\'esar Roberto de Souza, Bj\"orn Michele, Mert B\"ulent Sar{\i}y{\i}ld{\i}z, Philippe Weinzaepfel, Florent Perronnin, Diane Larlus, Yannis Kalantidis ·

    Task Alignment: A Simple Proxy for Practical Model Merging Across Diverse Vision Tasks

    arXiv:2604.12935v2 Announce Type: replace Abstract: Efficiently merging several models fine-tuned for different tasks, but stemming from the same pretrained base model, is of great practical interest. Despite extensive prior work, most evaluations of model merging in computer vis…

  311. arXiv cs.CV TIER_1 English(EN) · Wenqi Marshall Guo, Qingyun Qian, Shiyu Zhou, Guoping Luo, Shan Du ·

    MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts

    arXiv:2607.18673v1 Announce Type: new Abstract: Vision Language Models (VLMs) are well known for hallucinating non-existent objects in images. Objects with missing parts present a unique challenge for VLMs, stemming from both real-world knowledge bias and the scarcity of such ima…

  312. arXiv cs.CV TIER_1 English(EN) · Futa Waseda, Shojiro Yamabe, Daiki Shiono, Kento Sasaki, Tsubasa Takahashi ·

    Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models

    arXiv:2512.11899v2 Announce Type: replace Abstract: Large vision-language models (LVLMs) are vulnerable to typographic attacks, where misleading text inserted into an image can override visual understanding. However, existing evaluation protocols and defenses are largely focused …

  313. arXiv cs.CV TIER_1 English(EN) · Xiping Li, Jianghong Ma ·

    AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning

    arXiv:2509.25699v4 Announce Type: replace Abstract: Interleaved-Modal Chain-of-Thought (I-MCoT) advances vision-language reasoning, such as Visual Question Answering (VQA). This paradigm integrates specially selected visual evidence from the input image into the context of Vision…

  314. arXiv cs.CV TIER_1 English(EN) · Tarun Tomar ·

    Searching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language Models

    arXiv:2607.17052v1 Announce Type: new Abstract: Vision-language models normally execute the same complete vision encoder for every question, even when OCR, counting, object, attribute, and spatial queries may not require identical computation. We study whether fixed-budget combin…

  315. arXiv cs.CV TIER_1 English(EN) · Wei Chen, Zhiyuan Li, Shuo Xin ·

    OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference

    arXiv:2412.11475v3 Announce Type: replace Abstract: We present OmniVLM, a sub-billion-parameter vision-language model for efficient on-device inference. OmniVLM introduces a token compression mechanism that reduces visual token sequence length from 729 to 81 tokens, significantly…

  316. arXiv cs.CV TIER_1 English(EN) · Zhifang Zhang, Qiqi Tao, Jiaqi Lv, Na Zhao, Lei Feng, Joey Tianyi Zhou ·

    TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models

    arXiv:2509.24566v2 Announce Type: replace Abstract: Large vision-language models (LVLMs) have achieved impressive performance across a wide range of vision-language tasks, while they remain vulnerable to backdoor attacks. Existing backdoor attacks on LVLMs aim to force the victim…

  317. arXiv cs.CV TIER_1 English(EN) · Biao Chen, Yunqian Yu, Xiangxu Zhao, Zhongshu Chen, Mengmeng Jing, Lin Zuo ·

    U-shaped Multi-granularity Learning for Vision-Language Models

    arXiv:2607.14966v1 Announce Type: new Abstract: The prompt learning paradigm for vision-language models is effective yet faces a granularity dilemma: global prompts lack fine-grained semantic awareness, while local prompts ignore contextual associations, limiting cross-task gener…

  318. arXiv cs.CV TIER_1 English(EN) · Lin Zuo ·

    U-shaped Multi-granularity Learning for Vision-Language Models

    The prompt learning paradigm for vision-language models is effective yet faces a granularity dilemma: global prompts lack fine-grained semantic awareness, while local prompts ignore contextual associations, limiting cross-task generalization. This dilemma exists in dense predicti…

  319. arXiv cs.CV TIER_1 English(EN) · Yu Fang, Yuchun Feng, Dong Jing, Jiaqi Liu, Yue Yang, Zhenyu Wei, Daniel Szafir, Mingyu Ding ·

    When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs

    arXiv:2602.17659v2 Announce Type: replace Abstract: Vision-Language-Action models (VLAs) promise to ground language instructions in robot control, yet in practice often fail to faithfully follow language. When presented with instructions that lack strong scene-specific supervisio…

  320. arXiv cs.CV TIER_1 English(EN) · Xuanyi Hao, Zuoyuan Zhang, Zhibo Wang, Xiaoyi Pang, Jiahui Hu, Jiacheng Du, Shuguo Zhuo ·

    Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models

    arXiv:2607.13500v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have achieved strong performance in multimodal understanding, yet remain challenging to deploy on resource-constrained edge devices due to the substantial computational overhead of processing numerous v…

  321. arXiv cs.CV TIER_1 English(EN) · Shuguo Zhuo ·

    Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models

    Vision-Language Models (VLMs) have achieved strong performance in multimodal understanding, yet remain challenging to deploy on resource-constrained edge devices due to the substantial computational overhead of processing numerous visual tokens. Token reduction is a promising dir…

  322. arXiv cs.CV TIER_1 English(EN) · Jiahang Wang, Yirong Yang, Yanqing Zhu, Minghua Luo, Shichao Xie, Fei Liu, Mu Xu ·

    ReflectVLN: Training Vision-Language Navigation Agents with Reflective Reasoning

    arXiv:2607.12680v1 Announce Type: new Abstract: Existing vision-language navigation methods often couple a VLM with waypoint decoders to produce multi-step action plans, but they typically lack an explicit closed-loop mechanism for tracking semantic progress, diagnosing execution…

  323. arXiv cs.CV TIER_1 English(EN) · Xinyue Xu, Zheng Zhang, Kunyang Ma, Ge Zhu, Lianshuai Cao, Lei Wang, Zixuan Li, Yi Cheng ·

    DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery

    arXiv:2607.12319v1 Announce Type: new Abstract: As vision-language models (VLMs) are increasingly deployed in geospatial question answering and visual scene understanding, improving their spatial cognition capability on street view imagery for complex logical reasoning has emerge…

  324. arXiv cs.CV TIER_1 English(EN) · Mao Chen, Xiangkai Zhang, Zhiyong Liu, Chuankai Liu, Xu Yang ·

    More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization

    arXiv:2607.12429v1 Announce Type: new Abstract: Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same scene across extreme viewpoint gaps. Cross-view localization naturally provides a …

  325. arXiv cs.CV TIER_1 English(EN) · Lin Peng, Cong Wan, Zeyu Guo, SongLin Dong, Yihong Gong ·

    CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models

    arXiv:2607.12786v1 Announce Type: new Abstract: Cross-image comparative reasoning remains challenging for vision-language models (VLMs), especially when correct prediction requires fine-grained attribute grounding and globally consistent reasoning. We present CoRe, a unified fram…

  326. arXiv cs.CV TIER_1 English(EN) · Sania Waheed, Michael Milford, Sarvapali D. Ramchurn, Shoaib Ehsan ·

    Breaking D\'ej\`a Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning

    arXiv:2607.12818v1 Announce Type: new Abstract: Visual place recognition (VPR) is a key enabler of accurate localization and long-term autonomous navigation in robotics applications, such as loop closure detection for simultaneous localisation and mapping (SLAM). However, real-wo…

  327. arXiv cs.CV TIER_1 English(EN) · Jiho Hong, Eunae Kang, Sanghyun Kim, Young-Sik Shin ·

    Instance-Enriched Semantic Maps for Visual Language Navigation

    arXiv:2607.12630v1 Announce Type: cross Abstract: Visual Language Navigation (VLN) aims to enable an embodied agent to navigate complex environments by following natural language instructions. Recent approaches build semantic spatial maps and leverage Large Language Models (LLMs)…

  328. arXiv cs.CV TIER_1 English(EN) · Shoaib Ehsan ·

    Breaking Déjà Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning

    Visual place recognition (VPR) is a key enabler of accurate localization and long-term autonomous navigation in robotics applications, such as loop closure detection for simultaneous localisation and mapping (SLAM). However, real-world VPR deployment relies on selecting an image …

  329. arXiv cs.CV TIER_1 English(EN) · Yihong Gong ·

    CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models

    Cross-image comparative reasoning remains challenging for vision-language models (VLMs), especially when correct prediction requires fine-grained attribute grounding and globally consistent reasoning. We present CoRe, a unified framework for this problem. CoRe includes: (i) CoRe-…

  330. arXiv cs.CV TIER_1 English(EN) · Mu Xu ·

    ReflectVLN: Training Vision-Language Navigation Agents with Reflective Reasoning

    Existing vision-language navigation methods often couple a VLM with waypoint decoders to produce multi-step action plans, but they typically lack an explicit closed-loop mechanism for tracking semantic progress, diagnosing execution failures, and recovering from error accumulatio…

  331. arXiv cs.CV TIER_1 English(EN) · Young-Sik Shin ·

    Instance-Enriched Semantic Maps for Visual Language Navigation

    Visual Language Navigation (VLN) aims to enable an embodied agent to navigate complex environments by following natural language instructions. Recent approaches build semantic spatial maps and leverage Large Language Models (LLMs) for reasoning and decision making. Despite these …

  332. arXiv cs.CV TIER_1 English(EN) · Xu Yang ·

    More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization

    Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same scene across extreme viewpoint gaps. Cross-view localization naturally provides a promising pathway toward this ability, as it req…

  333. arXiv cs.CV TIER_1 English(EN) · Zhaoyang Li, Yanjun Li, Wangkai Li, Yujia Chen, Tianzhu Zhang ·

    Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models

    arXiv:2607.10640v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are costly at inference time because they must process long sequences of visual tokens. Existing token pruning methods often degrade under high compression by blindly discarding information, breaking sp…

  334. arXiv cs.CV TIER_1 English(EN) · Robert Wijaya, Ngai-Man Cheung ·

    Mixture of Cognitive Experts in Large Vision-Language Models

    arXiv:2607.10796v1 Announce Type: new Abstract: Large Vision Language Models (LVLMs) require strong reasoning over both visual and textual input. Recent work suggests that cognitive elements, especially diverse representations and metacognition, correlate with better performance.…

  335. arXiv cs.CV TIER_1 English(EN) · Hao Zheng, Jinyi Huang, Tiantian Zheng, Xun Xu, Tuka Alhanai ·

    Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos

    arXiv:2607.10797v1 Announce Type: new Abstract: Assembly action understanding is a key enabler for effective human-robot collaborative assembly, yet it remains challenging due to subtle motions and fine-grained hand-object interactions. We adapt vision-language models (VLMs) to t…

  336. arXiv cs.CV TIER_1 English(EN) · Quynh Vo, Phuc Dao, Cong-Duy Nguyen, Thong Nguyen ·

    When Depth Is Better Told Than Shown: Depth-Ordinal Prompting for Vision-Language Spatial Reasoning

    arXiv:2607.11173v1 Announce Type: new Abstract: Vision-language models (VLMs) are expected to reason about physical space -- which object is closer, what lies behind what, and how objects are arranged in 3D -- yet they still struggle with such spatial judgments. A natural remedy …

  337. arXiv cs.CV TIER_1 English(EN) · Yi Cheng ·

    DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery

    As vision-language models (VLMs) are increasingly deployed in geospatial question answering and visual scene understanding, improving their spatial cognition capability on street view imagery for complex logical reasoning has emerged as a key research priority. However, existing …

  338. arXiv cs.CV TIER_1 English(EN) · Thong Nguyen ·

    When Depth Is Better Told Than Shown: Depth-Ordinal Prompting for Vision-Language Spatial Reasoning

    Vision-language models (VLMs) are expected to reason about physical space -- which object is closer, what lies behind what, and how objects are arranged in 3D -- yet they still struggle with such spatial judgments. A natural remedy is to show the model a depth map, but we find th…

  339. arXiv cs.CV TIER_1 English(EN) · Yuncheng Yang, Feiyang Ye, Shixian Luo, Yinna Zhu, Lianlei Shan, Wangcai Zhao, Kuo Zhang, Yan Chen, Yong Wu, Yan Xie ·

    MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models

    arXiv:2607.09029v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have achieved success using homogeneous Transformers to process multimedia data. Recent studies show that heterogeneous structures interleaving efficient mechanisms, like linear attention, improve both …

  340. arXiv cs.CV TIER_1 English(EN) · Francis Fernandez, Arash Jahangiri, Salimeh Sekeh ·

    C-GAP: Class-Aware and Online Prompting Improves Vision-Language Models on Imbalanced Classes

    arXiv:2607.09008v1 Announce Type: new Abstract: Safety-critical perception systems must reliably detect rare object classes within small label spaces, a setting that long-tailed detection methods, designed for hundreds of classes with dense annotation, fundamentally do not addres…

  341. arXiv cs.CV TIER_1 English(EN) · Meng Wei, Chenyang Wan, Xiqian Yu, Tai Wang, Yuqiang Yang, Xiaohan Mao, Chenming Zhu, Wenzhe Cai, Hanqing Wang, Yilun Chen, Xihui Liu, Jiangmiao Pang ·

    StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling

    arXiv:2507.05240v2 Announce Type: replace-cross Abstract: Vision-and-Language Navigation (VLN) in real-world settings requires agents to process continuous visual streams and generate actions with low latency grounded in language instructions. While Video-based Large Language Mod…

  342. arXiv cs.CV TIER_1 English(EN) · Zuntao Liu, Yi Du, Taimeng Fu, Shaoshu Su, Cherie Ho, Chen Wang ·

    Vision-Language Memory for Spatial Reasoning

    arXiv:2511.20644v2 Announce Type: replace Abstract: Spatial reasoning is a critical capability for intelligent robots, yet current vision-language models (VLMs) still fall short of human-level performance in video-based spatial reasoning. This gap mainly stems from two challenges…

  343. arXiv cs.CV TIER_1 English(EN) · Mingjia Shi, Shuo Wang, Xiaobo Wang, Sifan Zhou, Kai Wang, Tianyu Fu, Chenxu Zhao, Anyang Su, Ping Jiang, Minghui Wu ·

    Dive Into the Implicit Biases of Low-rank Vision-language Alignment

    arXiv:2607.08194v1 Announce Type: new Abstract: Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requiring full-parameter updates. We challenge this view and investigate what happens…

  344. arXiv cs.CV TIER_1 English(EN) · Yan Xie ·

    MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models

    Vision-Language Models (VLMs) have achieved success using homogeneous Transformers to process multimedia data. Recent studies show that heterogeneous structures interleaving efficient mechanisms, like linear attention, improve both performance and inference latency over homogeneo…

  345. arXiv cs.CV TIER_1 English(EN) · Salimeh Sekeh ·

    C-GAP: Class-Aware and Online Prompting Improves Vision-Language Models on Imbalanced Classes

    Safety-critical perception systems must reliably detect rare object classes within small label spaces, a setting that long-tailed detection methods, designed for hundreds of classes with dense annotation, fundamentally do not address. Open-vocabulary detectors offer a promising a…

  346. arXiv cs.CV TIER_1 English(EN) · Minghui Wu ·

    Dive Into the Implicit Biases of Low-rank Vision-language Alignment

    Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requiring full-parameter updates. We challenge this view and investigate what happens when low-rank adaptation is applied to the LLM …

  347. arXiv cs.CV TIER_1 English(EN) · Hongyu Qu, Jianzhe Gao, Xiaobin Hu, Shaohuan Yang, Xinlei Yu, Rui Yan, Wenguan Wang, Xiangbo Shu, Shuicheng Yan ·

    Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

    arXiv:2607.07608v1 Announce Type: cross Abstract: Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Existing memory-augmented VLAs eith…

  348. arXiv cs.CV TIER_1 English(EN) · Jiajun Wu ·

    APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts

    Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility. To successfully execute a task, a robot must decompose goals, select task-relevant objects, and sequence actions, while ensuring that plans satisfy spatial constraints s…

  349. arXiv cs.CV TIER_1 English(EN) · Shuicheng Yan ·

    Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

    Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Existing memory-augmented VLAs either expand the observation window or retrieve histo…

  350. arXiv cs.CV TIER_1 English(EN) · Dominick Reilly, Manish Kumar Govind, Le Xue, Srijan Das ·

    VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models

    arXiv:2510.13808v2 Announce Type: replace Abstract: Large Vision Language Models (VLMs) excel at general visual reasoning but experience significant performance degradation when deployed in novel domains that exhibit substantial distribution shifts from their pretraining data. Ex…

  351. arXiv cs.CV TIER_1 English(EN) · Jiho Choi, Jaemin Kim, Sanghwan Kim, Seunghoon Hong, Jin-Hwi Park ·

    When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models

    arXiv:2604.03316v3 Announce Type: replace Abstract: Attention sinks are defined as tokens that attract disproportionate attention. While these have been studied in single modality transformers, their cross-modal impact in Large Vision-Language Models (LVLM) remains largely unexpl…

  352. arXiv cs.CV TIER_1 English(EN) · Xinda Liu, Qinyu Zhang, Weiqing Min, Guohua Geng, Shuqiang Jiang ·

    Structured-Condensed Prompt Tuning in Vision-Language Models for Fine-grained Image Recognition

    arXiv:2607.06185v1 Announce Type: new Abstract: Fine-grained image recognition poses a significant challenge due to the substantial expertise and effort required for manual annotation. Vision-language models (VLMs) like CLIP provide a compelling zero-shot alternative, reducing re…

  353. arXiv cs.CV TIER_1 English(EN) · Zhiwei Yang, Yuanchen Wu, Nan Zhang, Yucong Meng, Ke Yan, Shouhong Ding ·

    Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models

    arXiv:2607.05716v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong perception and reasoning capabilities. However, most existing models focus on isolated objects and neglect structured relationships for efficient target navigation, l…

  354. arXiv cs.CV TIER_1 English(EN) · Shuqiang Jiang ·

    Structured-Condensed Prompt Tuning in Vision-Language Models for Fine-grained Image Recognition

    Fine-grained image recognition poses a significant challenge due to the substantial expertise and effort required for manual annotation. Vision-language models (VLMs) like CLIP provide a compelling zero-shot alternative, reducing reliance on extensive labeled data. However, their…

  355. arXiv cs.CV TIER_1 English(EN) · Jaehyun Kwak, Nam Cao, Boryeong Cho, Segyu Lee, Sumyeong Ahn, Se-Young Yun ·

    Stage-wise Attention-Guided Region Sequencing for Adversarial Attacks on Large Vision-Language Models

    arXiv:2602.04356v2 Announce Type: replace Abstract: Targeted adversarial attacks on Large Vision-Language Models (LVLMs) test whether small image perturbations can steer model responses toward attacker-specified content. Under the standard L-infinity constraint, targeted attacks …

  356. arXiv cs.CV TIER_1 English(EN) · Xiangyu Shi, Ruoxi Yang, Wei Tao, Jiwen Zhang, Yanyuan Qiao, Qi Wu ·

    From Region Arrival to Instance-Level Grounding in Vision-and-Language Navigation

    arXiv:2607.03792v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) agents may satisfy conventional success criteria while still failing to establish reliable object-level grounding, because current evaluation protocols mainly reward stopping within a 3-meter r…

  357. arXiv cs.CV TIER_1 English(EN) · Anas Zafar, Leema Krishna Murali, Siddhant Bharadwaj, Ashish Vashist, Jia Wu ·

    Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs

    arXiv:2607.03647v1 Announce Type: new Abstract: Large vision language models (VLMs) report strong accuracy on medical question-answering, yet it remains unclear whether they reason from visual evidence or exploit textual shortcuts. We introduce a counterfactual evaluation framewo…

  358. arXiv cs.CV TIER_1 English(EN) · Amanda Adkins, Tarunvidyut Ravisankar, Joydeep Biswas ·

    PreSIST: Vision-Language-Informed Object Persistence Prediction in Open-World Scenes

    arXiv:2607.04057v1 Announce Type: new Abstract: Robots deployed over long periods must reason about environments that change over time. Existing long-term perception systems often address object change reactively, updating their maps only after revisiting a scene and observing th…

  359. arXiv cs.CV TIER_1 English(EN) · Zhaoxu Li, Chenqi Kong, Yi Yu, Qiangqiang Wu, Xinghao Jiang, Ngai-Man Cheung, Bihan Wen, Alex Kot, Xudong Jiang ·

    SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision

    arXiv:2508.03177v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their real-world applicability. Although previous mitiga…

  360. arXiv cs.CV TIER_1 English(EN) · Marwane Hariat, David Filliat, Antoine Manzanera ·

    VLRC: Vision-Language Reprojection Consistency as a scalable signal for better feed-forward 3D pretraining

    arXiv:2607.02707v1 Announce Type: new Abstract: Feed-forward 3D models are commonly trained using either expensive geometric supervision or self-supervised photometric objectives, both of which provide incomplete learning signals. We introduce Vision-Language Reprojection Consist…

  361. arXiv cs.CV TIER_1 English(EN) · Zaiyu Cheng, Khai-Nguyen Nguyen, Antonio Mastropaolo ·

    Prior Bias in Vision Language Models on UML Diagram Interpretation

    arXiv:2607.02853v1 Announce Type: new Abstract: Vision Language Models (VLMs) are increasingly applied to software engineering artifacts, especially UML class diagrams whose meaning depends on visual notation. Yet, it is unclear whether VLMs actually read such diagrams or instead…

  362. arXiv cs.CV TIER_1 English(EN) · Haoyu Zhang, Yangyang Guo, Mohan Kankanhalli ·

    Overloading Large Vision-Language Models for Jailbreaking

    arXiv:2607.02961v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) exhibit remarkable vision-language capabilities and are increasingly deployed in real-world applications such as personal assistants, document analysis systems, and embodied agents. However, thei…

  363. arXiv cs.CV TIER_1 English(EN) · Suryanarayana Reddy Yarrabothula, Manisha Chawla, Kunal Sinha, Gagan Raj Gupta, Sashank Lekkala, Ashirvadhan Dosapati, Saikamal Nannuri, Katragadda Ajay RamaSwamy Chowdary Gowtham ·

    SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments

    arXiv:2607.05264v1 Announce Type: new Abstract: Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test vision-language models under the visual and procedural conditions of real indust…

  364. arXiv cs.CV TIER_1 English(EN) · Pin Tang, Guoqing Wang, Xiangxuan Ren, Zhongdao Wang, Guodongfang Zhao, Bailan, Chao Ma ·

    PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving

    arXiv:2607.04637v1 Announce Type: new Abstract: Vision-Language-Action Models (VLAs), which leverage the advanced reasoning capabilities of Vision-Language Models (VLMs), show promising generalization in complex autonomous driving scenarios. Existing VLAs typically predict and op…

  365. arXiv cs.CV TIER_1 English(EN) · Geng Li, Yuxin Peng ·

    BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception

    arXiv:2607.03184v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) demonstrate impressive general capabilities, they struggle with fine-grained perception in ultra-high-resolution (UHR) images, particularly for tiny objects in cluttered scenes. Existin…

  366. arXiv cs.CV TIER_1 English(EN) · Katragadda Ajay RamaSwamy Chowdary Gowtham ·

    SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments

    Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test vision-language models under the visual and procedural conditions of real industrial CCTV, where workers appear as distant figur…

  367. arXiv cs.CV TIER_1 English(EN) · Lev Sorokin, Chen Yang, Ken E. Friedl, Andrea Stocco ·

    Search-based Testing of Vision Language Models for In-Car Scene Understanding

    arXiv:2607.02300v1 Announce Type: new Abstract: In the automotive domain, in-car scene understanding (ISU) enables the detection of safety-critical events, such as driver distraction, and supports drivers or passengers by analyzing the in-car scene and adapting the environment (e…

  368. arXiv cs.CV TIER_1 English(EN) · Cheng Chen, Yuyu Guo, Pengpeng Zeng, Jingkuan Song, Peng Di, Hang Yu, Lianli Gao ·

    From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion

    arXiv:2601.10710v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) create a severe visual feature bottleneck by using a crude, asymmetric connection that links only the output of the vision encoder to the input of the large language model (LLM). This static archite…

  369. arXiv cs.CV TIER_1 English(EN) · Zelin Peng, Yichen Zhao, Yu Huang, Piao Yang, Feilong Tang, Zhengqin Xu, Xiaokang Yang, Wei Shen ·

    NEARL: Interacted Query Adaptation with Orthogonal Regularization for Medical Vision-Language Understanding

    arXiv:2508.04101v2 Announce Type: replace Abstract: Computer-aided medical image analysis is crucial for disease diagnosis and treatment planning. While vision-language models (VLMs) such as CLIP exhibit strong generalization ability, their direct application to medical imaging r…

  370. arXiv cs.CV TIER_1 English(EN) · Yiqian Liu, Iuliia Kotseruba, John K. Tsotsos ·

    Disentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task

    arXiv:2607.01503v1 Announce Type: new Abstract: In this paper, we study depth perception of vision-language models (VLMs) to isolate the effects of pictorial depth cues and disentangle vision and language influences on model performance. To this end, we combine depth-ordering and…

  371. arXiv cs.CV TIER_1 English(EN) · Andrea Stocco ·

    Search-based Testing of Vision Language Models for In-Car Scene Understanding

    In the automotive domain, in-car scene understanding (ISU) enables the detection of safety-critical events, such as driver distraction, and supports drivers or passengers by analyzing the in-car scene and adapting the environment (e.g., ambient lighting). The industry is increasi…

  372. arXiv cs.CV TIER_1 English(EN) · Ke Wu, Yanan Zhang, Yingjie Gao, Wenhao Li, Chenyu Zhou, XinZhu Ma, Jiaxin Chen, Di Huang ·

    DroneFINE: Domain-Aware Parameter-Efficient Fine-Tuning of Vision-Language Detectors for Drone Images

    arXiv:2607.00338v1 Announce Type: new Abstract: Object detection for Unmanned Aerial Vehicles (UAVs) working in open and dynamic environments is a highly challenging task. While Vision-Language Models (VLMs) have offered a powerful solution for universal object detection, adaptin…

  373. arXiv cs.CV TIER_1 English(EN) · Diogo Gl\'oria-Silva, Jo\~ao Cardeira, Manuel Letras da Luz, Afonso Simpl\'icio, Gon\c{c}alo Vinagre, Diogo Tavares, Rafael Ferreira, In\^es Calvo, In\^es Vieira, David Semedo, Jo\~ao Magalh\~aes ·

    AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model

    arXiv:2606.19100v3 Announce Type: replace Abstract: Large Vision and Language Models (LVLMs) have advanced rapidly, yet European Portuguese (pt-PT) remains systematically underserved by existing open-source multimodal models, which either conflate it with Brazilian Portuguese or …

  374. arXiv cs.CV TIER_1 English(EN) · Sangyun Chung, Youngjoon Yu, Se Yeon Kim, Youngchae Chee, Yong Man Ro ·

    Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking

    arXiv:2412.20750v3 Announce Type: replace Abstract: Large-scale Vision-Language Models (VLMs) have achieved notable progress in aligning visual inputs with text. However, their ability to deeply understand the unique physical properties of non-RGB vision sensor images remains lim…

  375. arXiv cs.CV TIER_1 English(EN) · Kyan Mahajan, Mohammad Saqlain ·

    SpiralFovea: Input-Adaptive Foveated Tokenization as a Third Lever of Resource-Adaptive Inference

    arXiv:2607.00780v1 Announce Type: new Abstract: Most adaptive-inference techniques for foundation models change what the model does - early exit, MoE routing, KV-cache compression, dynamic attention sparsity. The input that hits the backbone, however, remains a fixed-grid tokenis…

  376. arXiv cs.CV TIER_1 English(EN) · Kensuke Nakamura, Byung-Woo Hong ·

    Personalized Object Identification and Localization via In-Context Inference with Vision-Language Models

    arXiv:2607.00357v1 Announce Type: new Abstract: Personalized object localization (POL) localizes an object instance in a query image based on a few reference images with bounding-box annotations and a target object label. The pioneering method, IPLoc, solves this task through in-…

  377. arXiv cs.CV TIER_1 English(EN) · Mohammad Saqlain ·

    SpiralFovea: Input-Adaptive Foveated Tokenization as a Third Lever of Resource-Adaptive Inference

    Most adaptive-inference techniques for foundation models change what the model does - early exit, MoE routing, KV-cache compression, dynamic attention sparsity. The input that hits the backbone, however, remains a fixed-grid tokenisation indifferent to image content. We argue tha…

  378. arXiv cs.CV TIER_1 English(EN) · Benoît Macq ·

    Localized Conformal Prediction for Image Classification with Vision-Language Models

    Conformal predictions have attracted significant attention in the field of uncertainty quantification, mainly because of their strong marginal coverage guarantees. Full conditional guarantee is not an attainable goal, a well known fact in conformal predictions literature. As a re…

  379. arXiv cs.CV TIER_1 English(EN) · Dieter Schmalstieg ·

    Think While You Map: Asynchronous Vision-Language Agents for Incremental 3D Scene Graphs

    Open-vocabulary 3D scene graph methods typically operate in two stages: first reconstruct, then enrich with vision-language models, leaving the graph unqueryable during exploration. We argue that this sequential coupling is unnecessary and propose an asynchronous architecture in …

  380. arXiv cs.CV TIER_1 English(EN) · Kai Jiang, Ruishu Zhu, Siqi Huang, Hongyuan Zhang, Xuelong Li ·

    Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models

    arXiv:2606.30168v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) often fail in fine-grained visual reasoning, as question-relevant visual cues are diluted by dense and redundant image tokens. Recent multimodal reasoning methods usually extend chain-of-thou…

  381. arXiv cs.CV TIER_1 English(EN) · Atif Belal, Heitor R. Medeiros, Marco Pedersoli, Eric Granger ·

    VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors

    arXiv:2510.00458v3 Announce Type: replace Abstract: Vision-language object detectors (VLODs) such as YOLO-World and Grounding DINO exhibit strong zero-shot generalization, but their performance degrades under distribution shift. Test-time adaptation (TTA) offers a practical way t…

  382. arXiv cs.CV TIER_1 English(EN) · Enshen Zhou, Yibo Li, Jingkun An, Jiayuan Zhang, Shanyu Rong, Mengzhen Liu, Yi Han, Yuheng Ji, Huajie Tan, Jiawei He, Pengwei Wang, Zhongyuan Wang, Cheng Chi, Lu Sheng, Shanghang Zhang ·

    Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics

    arXiv:2512.13660v3 Announce Type: replace-cross Abstract: Spatial tracing, as a fundamental embodied interaction ability for robots, is inherently challenging as it requires multi-step metric-grounded reasoning compounded with complex spatial referring and real-world metric measu…

  383. arXiv cs.CV TIER_1 English(EN) · Duo Zheng, Shijia Huang, Yanyang Li, Liwei Wang ·

    Efficient-VLN: A Simple yet Strong Baseline for Efficient Vision-Language Navigation

    arXiv:2512.10310v2 Announce Type: replace Abstract: While Multimodal Large Language Models (MLLMs) have demonstrated significant promise in Vision-Language Navigation (VLN), existing agents remain heavily constrained by systemic bottlenecks across inference, training, and data co…

  384. arXiv cs.CV TIER_1 English(EN) · Wenhui Liao, Hongliang Li, Pengyu Xie, Xinyu Cai, Yufan Shen, Yi Xin, Qi Qin, Shenglong Ye, Tianbin Li, Ming Hu, Junjun He, Yihao Liu, Wenhai Wang, Min Dou, Bin Fu, Botian Shi, Yu Qiao, Lianwen Jin ·

    HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding

    arXiv:2602.12957v3 Announce Type: replace Abstract: Document parsing is a fundamental task in multimodal understanding, supporting a wide range of downstream applications such as information extraction and intelligent document analysis. Benefiting from strong semantic modeling an…

  385. arXiv cs.CV TIER_1 English(EN) · Fawaz Sammani, Tzoulio Chamiti, Nikos Deligiannis ·

    On Test-Time Scaling for Vision-Language Models

    arXiv:2606.28864v1 Announce Type: new Abstract: Test-time scaling is a paradigm where large models use additional compute at inference to achieve better performance, without changing model weights. While it has been widely studied for Large Language Models (LLMs), its applicabili…

  386. arXiv cs.CV TIER_1 English(EN) · Hong-Han Wang, Yuntao Wang, Hu Ding ·

    MIRROR: Aligning Semantic Relations from Language to Image via Gromov--Wasserstein

    arXiv:2606.29462v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) inherit rich relational priors from their language backbones, yet often fail when asked to apply these relationships in visual contexts. We trace this failure to a structural blind spot: proj…

  387. arXiv cs.CV TIER_1 English(EN) · Xuelong Li ·

    Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models

    Multimodal large language models (MLLMs) often fail in fine-grained visual reasoning, as question-relevant visual cues are diluted by dense and redundant image tokens. Recent multimodal reasoning methods usually extend chain-of-thought from language models into visual or latent s…

  388. arXiv cs.CV TIER_1 English(EN) · Jiajun Cheng, Xianwu Zhao, Sainan Liu, Xiaofan Yu, Ravi Prakash, Patrick J. Codd, Jonathan Elliott Katz, Shan Lin ·

    SurgXBench: Explainable Vision-Language Model Benchmark for Surgery

    arXiv:2505.10764v4 Announce Type: replace Abstract: Innovations in digital intelligence are transforming robotic surgery with more informed decision-making. Real-time awareness of surgical instrument presence and actions (e.g., cutting tissue) is essential for such systems. Yet, …

  389. arXiv cs.CV TIER_1 English(EN) · Chanik Kang, Rapha\"el Pestourie, Haejun Chung ·

    VLM-Aware Meta-Optic Front-End Design for Frozen Vision-Language Models

    arXiv:2606.27646v1 Announce Type: new Abstract: Conventional machine-vision pipelines typically rely on high-quality optics that produce clean, human-interpretable images, and optical design has therefore been driven by image-level criteria such as resolution, aberration correcti…

  390. arXiv cs.CV TIER_1 English(EN) · Nan Yang, Zhanwen Liu, Linfeng Zhang, Shangyu Xie, Yang Wang, Wenzhuo Zhou, Xiangmo Zhao ·

    MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving

    arXiv:2606.27660v1 Announce Type: new Abstract: Vision-Language Models (VLMs) improve generalization and interpretability in autonomous driving but suffer from efficiency issues due to long visual token sequences, particularly in standard multi-view settings. Existing token pruni…

  391. arXiv cs.CV TIER_1 English(EN) · Chaoxiang Cai, Minghe Weng, Jie Li, Yibo Jiang, Longrong Yang, Zequn Qin, Xi Li ·

    GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models

    arXiv:2606.26287v1 Announce Type: new Abstract: With the increase in model parameters and training data, the instruction following and generalization capabilities of Large VisionLanguage Models (LVLMs) have been significantly improved. Based on the Mixture of Experts (MoE) archit…

  392. arXiv cs.CV TIER_1 English(EN) · Sourabh Sharma, Sonam Gupta, Sadbhawna ·

    Improving Reasoning in Vision-Language Models via Perception Verified Self-Training

    arXiv:2606.22158v2 Announce Type: replace Abstract: Achieving human-like reasoning in Vision-Language Models (VLMs) remains a long-standing challenge. Recent approaches leverage Chain-of-Thought (CoT) rationales generated by human annotators or proprietary models to improve reaso…

  393. arXiv cs.CV TIER_1 English(EN) · Xiangmo Zhao ·

    MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving

    Vision-Language Models (VLMs) improve generalization and interpretability in autonomous driving but suffer from efficiency issues due to long visual token sequences, particularly in standard multi-view settings. Existing token pruning methods employ fixed pruning rate allocation …

  394. arXiv cs.CV TIER_1 English(EN) · Haejun Chung ·

    VLM-Aware Meta-Optic Front-End Design for Frozen Vision-Language Models

    Conventional machine-vision pipelines typically rely on high-quality optics that produce clean, human-interpretable images, and optical design has therefore been driven by image-level criteria such as resolution, aberration correction, and pixel fidelity. However, such optics are…

  395. arXiv cs.CV TIER_1 English(EN) · Liang Zhang ·

    Bridging Vision and Language Concepts through Optimal Transport Semantic Flow

    Concept Bottleneck Models (CBMs) promise transparent reasoning by predicting through human-interpretable concepts, yet their effectiveness fundamentally depends on how well visual and textual representations are aligned or matched. Existing vision-language CBMs often rely on pre-…

  396. arXiv cs.CV TIER_1 English(EN) · Huizhen Shu, Xuying Li, Hongxu Lin, Wenjie Sun, Hui Li ·

    Steering Vision-Language Models with Joint Sparse Autoencoders

    arXiv:2606.25657v1 Announce Type: new Abstract: Sparse Autoencoders (SAEs) have shown promise for analyzing language models, but applying them to vision-language models (VLMs) often yields representations that are difficult to use as controllable cross-modal steering directions. …

  397. arXiv cs.CV TIER_1 English(EN) · Qitong Wang, Fan Du, Pranav Maneriker, Jihui Jin, Christopher Rasmussen ·

    Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding

    arXiv:2606.25160v1 Announce Type: cross Abstract: The rapid rise of Vision-Language Models (VLMs) in egocentric visual understanding has made low-latency inference in human-robot collaborative (HRC) tasks increasingly critical. Weight pruning techniques developed for VLMs to shri…

  398. arXiv cs.CV TIER_1 English(EN) · Weiran Huang ·

    Black-Box Continual Learning for Vision-Language Models

    The rapid deployment of Vision-Language Models (VLMs) in dynamic environments necessitates the ability to learn continuously without forgetting. However, traditional continual learning (CL) settings often rely on white-box paradigms, which is increasingly invalidated by the shift…

  399. arXiv cs.CV TIER_1 English(EN) · Thomas Brox ·

    Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology

    We study how to train visually grounded vision-language models (VLMs) for radiology without manual spatial annotations. We introduce RefRad2D, a large-scale bilingual (German/English) dataset of 1.2M CT and MR image-text pairs derived from clinical practice, with task-specific VQ…

  400. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    ICML 2026 | Semantic Robustness Certification for Vision-Language Models

    <h1 style="font-size: 15px; line-height: 1.85; margin: 24px 0 14px; font-weight: 700; color: #111827;"><span><br /></span></h1><h1 style="font-size: 15px; line-height: 1.85; margin: 24px 0 14px; font-weight: 700; color: #111827;">原文作者:公众号“专知”</h1><p>原文链接:<a href="https://mp.weixi…

  401. Medium — fine-tuning tag TIER_1 English(EN) · Arpit Neewaliya ·

    How I Fine-Tuned a Vision Language Model for Drone Image Understanding

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@arpitneewaliya/how-i-fine-tuned-a-vision-language-model-for-drone-image-understanding-ab3329d6c210?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1536/1*kzeM1R7oQ…

  402. Towards AI TIER_1 English(EN) · Sanjana Dubey ·

    Vision Language Grounding: How AI Connects “Dog” to Pixels, and Where It Falls Apart

    <h4><em>A research grounded deep dive</em></h4><h3>1. The Hook</h3><p>Picture an ordinary photo like this one. A vision language model will confidently caption it “a dog in a park,” and it will be right, almost every time [1].</p><figure><img alt="A golden retriever sits on green…

  403. dev.to — LLM tag TIER_1 English(EN) · pixelbank dev ·

    Vision-Language Models — Deep Dive + Problem: Polynomial Regression Error

    <p><em>A daily deep dive into llm topics, coding problems, and platform features from <a href="https://pixelbank.dev" rel="noopener noreferrer">PixelBank</a>.</em></p> <h2> Topic Deep Dive: Vision-Language Models </h2> <p><em>From the Multimodal LLMs chapter</em></p> <h2> Introdu…

  404. dev.to — LLM tag TIER_1 English(EN) · Vahid Aghajani ·

    Vision Language Models — When AI Learns to See and Talk (Part 3 of 3)

    <blockquote> <p>Originally published on <a href="https://software-engineer-blog.com/content/vision-language-models-when-ai-learns-to-see-and-talk-part-3-of-3?id=51" rel="noopener noreferrer">my blog</a>. Cross-posted here with a canonical link.</p> </blockquote> <p> </p> <p>This …

  405. Mastodon — mastodon.social TIER_1 English(EN) · appassionato ·

    2/ Generative AI: Powered by Chery's integration with large language and vision models, allowing the robot to understand natural speech, process visual environm

    2/ Generative AI: Powered by Chery's integration with large language and vision models, allowing the robot to understand natural speech, process visual environments, and carry on human-like conversations. Combined with Ai (Artificial Intelligence), the name essentially translates…

  406. Mastodon — mastodon.social TIER_1 English(EN) · notatechguy ·

    VAORA aligns vision-language model reasoning with physical actions VAORA, a new reward design on arXiv, targets hallucinated reasoning and reasoning-action misa

    VAORA aligns vision-language model reasoning with physical actions VAORA, a new reward design on arXiv, targets hallucinated reasoning and reasoning-action misalignment in vision-language models on physical tasks. https://www. notatechguy.com/vaora-aligns-v ision-language-model-r…