PulseAugur
实时 13:08:26
English(EN) Vision-language models for chest radiography do not always need the image

新研究增强了用于医疗、检索和机器人任务的视觉-语言模型

研究人员正在开发新方法来改进各种领域的视觉-语言模型(VLM)。一篇论文介绍了CoT-Mediate,一个用于评估生成推理如何影响VLM在医疗情境下预测的框架,发现推理的位置比其声明的来源更关键。另一项研究提出了ReToken,一个可学习的单一嵌入,通过选择相关的视觉标记来增强视觉检索,在Visual Haystacks和LVBench等基准测试中显示出显著的提升。对于机器人领域,BioVLN是一个专为生物医学实验室视觉-语言导航设计的新模拟平台,解决了精确仪器交互的需求。此外,还提出了一种名为SeGP-CL的方法用于VLM的持续学习,旨在保持跨模态语义几何并对抗灾难性遗忘。最后,S-GRPO提供了一个用于大型VLM的统一的训练后框架,整合了监督微调和强化学习,以提高适应性并保持通用能力。 AI

影响 VLM推理、检索、机器人导航、持续学习和统一训练后框架方面的这些进展正在推动AI在专业和通用应用中的能力边界。

排序理由 该集群包含多篇关于视觉-语言模型及相关技术进展的研究论文。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 422 个来源。 我们如何撰写摘要 →

新研究增强了用于医疗、检索和机器人任务的视觉-语言模型

报道来源 [422]

  1. arXiv cs.AI TIER_1 English(EN) · Di Wu, Xiaohui Zhu ·

    视觉-语言模型智能体之间潜在通信的事后稀疏编码

    arXiv:2608.10198v1 Announce Type: new Abstract: Latent-space communication allows heterogeneous vision-language model agents to exchange continuous representations without serializing visual and reasoning states into text. Vision Wormhole realizes this approach by translating vis…

  2. arXiv cs.AI TIER_1 English(EN) · Meiwen Ding, Song Xia, Chenqi Kong, Xudong Jiang ·

    针对商用多模态大语言模型的隐蔽视觉提示注入攻击

    arXiv:2603.29418v2 Announce Type: replace-cross Abstract: Although multimodal large language models (MLLMs) are increasingly deployed in real-world applications, their instruction-following behavior leaves them vulnerable to prompt injection attacks. Existing prompt injection met…

  3. arXiv cs.CL TIER_1 English(EN) · Mizanur Rahman, Arshia Azimlu, Shadikur Rahman, Md Tahmid Rahman Laskar, Amran Bhuiyan, Shafiq Joty, Enamul Hoque Prince ·

    VisEditBench:视觉语言模型能否根据多模态反馈编辑可视化代码?

    arXiv:2608.10408v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization authoring is inherently iterative: users frequently revise existi…

  4. arXiv cs.CL TIER_1 English(EN) · Changhao Xiang, Shangyu Xing, Zhen Wu, Jianbing Zhang, Xinyu Dai ·

    多模态语码转换:将视觉对象穿插到语言中以实现显式对象级对齐

    arXiv:2608.11167v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment su…

  5. arXiv cs.LG TIER_1 English(EN) · Binesh Sadanandan, Vahid Behzadan ·

    预测熵作为医学视觉-语言模型错误和释义不稳定性联合筛查方法

    arXiv:2604.08941v2 Announce Type: replace Abstract: Medical Vision-Language Models (VLMs) answering binary presence questions on chest radiographs can fail in two linked ways: they are confidently wrong, and they change answers when a clinically equivalent question is rephrased. …

  6. arXiv cs.AI TIER_1 English(EN) · Yuan Wang, Hualiang Wang, Yixin Chen, Songtao Jiang, Shujian Gao, Jiaming Lin, Siming Fu, Jian Wu, Zuozhu Liu ·

    MedUP:唤醒医疗视觉语言模型中的统一理解与感知

    arXiv:2608.10635v1 Announce Type: cross Abstract: Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging. Existing approaches either verbalize regions as coordinate strings or re…

  7. arXiv cs.AI TIER_1 English(EN) · Li Wenjie, Yash Jangir, Ignacy Stepka, Yash Agarwal, Marion Kipsang, Yonatan Bisk ·

    迷失于重建:在视觉-语言-动作模型中将动作表征与语言对齐

    arXiv:2608.10484v1 Announce Type: cross Abstract: Action verbs describe not only the physical outcomes of actions, but also how those actions are performed. Yet action representations in vision-language-action models (VLAs) are typically optimized for reconstruction under L1/L2 l…

  8. arXiv cs.AI TIER_1 English(EN) · Foundation Model Team, XPeng Inc ·

    XCoT-VLA:面向视觉-语言-动作驾驶的可执行思维链

    arXiv:2608.10976v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models can connect scene understanding, semantic reasoning, and trajectory generation for autonomous driving. However, verbose natural-language Chain-of-Thought (CoT) is poorly suited to real-time contro…

  9. arXiv cs.AI TIER_1 English(EN) · Gesina Schwalbe, Mert Keser, Moritz Bayerkuhnlein, Edgar Heinert, Annika M\"utze, Marvin Keller, Sparsh Tiwari, Georgii Mikriukov, Diedrich Wolter, Jae Hee Lee, Matthias Rottmann ·

    解释、验证和对齐视觉语言模型嵌入中的语义层次结构

    arXiv:2603.26798v2 Announce Type: replace-cross Abstract: Vision-language model (VLM) encoders such as CLIP enable strong retrieval and zero-shot classification in a shared image-text embedding space, yet the semantic organization of this space is rarely inspected. We present a p…

  10. arXiv cs.AI TIER_1 English(EN) · R\'ois\'in Luo ·

    ZetaGPT: 无位置编码状态空间注意力语言模型的参考实现

    arXiv:2608.09432v1 Announce Type: cross Abstract: Transformer-based language models rely on self-attention, whose computation is permutation-equivariant and therefore lacks an intrinsic mechanism for representing token order. Existing architectures address this limitation by expl…

  11. arXiv cs.AI TIER_1 English(EN) · Jieyu Zhang, Ziqi Gao, Luke Zettlemoyer, Ranjay Krishna ·

    视觉-语言基础作为双向概念对应

    arXiv:2608.07886v1 Announce Type: cross Abstract: Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the corresponding…

  12. arXiv cs.LG TIER_1 English(EN) · Ao Zhou, Zhiwei Jiang, Zifeng Cheng, Cong Wang, Shufan Yang, Haoru Chen, Qing Gu ·

    面向动态分布感知的视觉语言表示学习中的不确定性跟踪

    arXiv:2608.09011v1 Announce Type: new Abstract: Uncertainty Quantification (UQ) aims to measure the reliability of model predictions, serving as a critical safeguard for deploying Vision-Language Models (VLMs) in safety-critical scenarios. Post-hoc approaches are widely adopted d…

  13. arXiv cs.CL TIER_1 English(EN) · Zhanna Mukhametsharip (Saarland University, Germany), Vera Demberg (Saarland University, Germany, Max Planck Institute for Informatics, Germany), Varsha Suresh (Max Planck Institute for Informatics, Germany) ·

    PragMatch:在大型视觉语言模型中区分实用性不一致与跨模态不匹配

    arXiv:2608.09772v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason about relationships between images and text or rely on superficial correlations…

  14. Hugging Face Daily Papers TIER_1 English(EN) ·

    VisEditBench:视觉语言模型能否根据多模态反馈编辑可视化代码?

    Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization authoring is inherently iterative: users frequently revise existing visualizations to repair flawed charts or ada…

  15. arXiv cs.AI TIER_1 English(EN) · Bhavika Jalli, Nikhil Korati Prasanna, Jayanta Choudhury ·

    一张图胜过千tokens:视觉语言模型如何在提高准确性的同时降低AI能源成本

    arXiv:2608.07427v1 Announce Type: new Abstract: LLM inference accounts for over 90% of AI operational energy, scaling directly with input token count---a critical inefficiency for telecom network analytics and numerical time-series data analysis (NTSDA), where raw multivariate KP…

  16. arXiv cs.AI TIER_1 English(EN) · Nikos Theodoridis, Reenu Mohandas, Ganesh Sistu, Anthony Scanlan, Ciar\'an Eising, Tim Brophy ·

    探索轻量级视觉语言模型在自动驾驶中的视觉概念

    arXiv:2603.06054v2 Announce Type: replace-cross Abstract: The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios. However,…

  17. arXiv cs.AI TIER_1 English(EN) · Ying Chen, Weizhen Li, Zhe Hu, Zhenjiang Li, Rui Jiang, Zhifeng Gu, Lihuang Fang, Jiangping Liu, Lei Yi, Jie Chen ·

    Capek 0.5:面向具身智能的以执行为中心的视觉语言模型

    arXiv:2608.06756v1 Announce Type: new Abstract: Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reaso…

  18. arXiv cs.LG TIER_1 English(EN) · Minseok Kang, Hyunwoo Kim, Chanyoung Kim, Minwoo Kim, Jaekoo Lee, Dahuin Jung ·

    Prune Once:视觉语言模型的免重训练、任务无关剪枝

    arXiv:2608.06901v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal tasks through large-scale pre-training, yet their rapidly increasing computational and memory requirements pose significant challenges…

  19. arXiv cs.CL TIER_1 English(EN) · Rahul Murali Shankar, Titus von der Malsburg, Sebastian Pad\'o ·

    视觉世界实验中的注视行为可用现成的语言-视觉编码器进行建模

    arXiv:2608.07282v1 Announce Type: new Abstract: The recent advances in neural language models have also spurred much work in computational psycholinguistics, asking whether neural LMs are also promising models of human language processing. However, work has been overwhelmingly fo…

  20. Hugging Face Daily Papers TIER_1 English(EN) ·

    视觉-语言基础作为双向概念对应

    ConCor-1 treats vision-language grounding as bidirectional concept correspondence, jointly predicting text spans, image segments, and cross-modal matches without prespecified phrases.

  21. arXiv cs.LG TIER_1 English(EN) · Rasul Khanbayov, Hasan Kurban ·

    一致性存在可计算盲点:一种用于视觉语言图像阅读的无标签可靠性交换理论

    arXiv:2608.05675v1 Announce Type: new Abstract: Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change. This has a known blind spot, a systematic misreading survives the perturbation and gets certi…

  22. arXiv cs.AI TIER_1 English(EN) · J. de Curt\`o, Dayani Plasencia, Diego S\'anchez, I. de Zarz\`a ·

    Zero-Shot 视觉语言控制中的视觉接地

    arXiv:2608.06154v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that decisions are grounded in visual input: simulator dynamics and conservative action priors can p…

  23. arXiv cs.AI TIER_1 English(EN) · Qinwu Xu ·

    通过分布偏移下的分阶段偏好优化减少视觉语言模型的幻觉

    arXiv:2605.16411v2 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likel…

  24. Hugging Face Daily Papers TIER_1 English(EN) ·

    Capek 0.5:面向具身智能的以执行为中心的视觉语言模型

    Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reasoned about, and verified. Meeting these demands r…

  25. Hugging Face Daily Papers TIER_1 English(EN) ·

    一致性存在可计算盲点:一种无标签视觉语言图文识别可靠性的交换理论

    Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change. This has a known blind spot, a systematic misreading survives the perturbation and gets certified wrong, which we show is computable, not jus…

  26. arXiv cs.CL TIER_1 English(EN) · Sirun Li, Minghao Liu, Ling Dai, Yong Li, Haoxin Lyu, Junting Zhou, Fan Zhang ·

    SIGNPOST-Bench:对多模态大语言模型中文本-视觉冲突解决能力的基准测试

    arXiv:2608.04244v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they confli…

  27. arXiv cs.LG TIER_1 English(EN) · Chen Zhong, Xiao An, Zijie Wang, Jiepan Li, Guangyi Yang, Wei He ·

    DIVE:面向高效视觉语言模型的动态迭代视觉证据构建

    arXiv:2608.04496v1 Announce Type: cross Abstract: Visual inputs in vision-language models (VLMs) are often encoded into substantially longer token sequences than text, making visual tokens a major bottleneck for efficient inference. Abundant recent methods address this bottleneck…

  28. arXiv cs.LG TIER_1 English(EN) · Hongyu Zhang, Cheng Yan, Xiang Xia, Wuyang Zhang ·

    超越全局路由聚合:面向MoE视觉语言模型的相位感知专家合并

    arXiv:2608.04454v1 Announce Type: cross Abstract: Mixture-of-experts vision-language models (MoE-VLMs) increase model capacity with sparse expert activation, yet deployment requires storing the full expert pool. Training-free expert merging reduces this burden, and many routing-b…

  29. arXiv cs.LG TIER_1 English(EN) · Grzegorz Gruszczynski, Pawel Olszowiec, Michal Byra, Grzegorz Stefanski, Alberto Presta ·

    循环视觉Transformer的训练交叉路口:循环、神经ODE和深度监督

    arXiv:2608.04879v1 Announce Type: new Abstract: Vision Transformers (ViTs) achieve strong image-recognition performance, but their parameter count grows linearly with depth when each block is independently parameterized. Single-block recurrent ViTs (bViT) remove this growth by re…

  30. arXiv cs.CL TIER_1 English(EN) · Ali Khoramfar, Mohammad Javad Dousti, Alireza Mohamadian, Heshaam Faili ·

    评估视觉语言模型在视觉和文本扰动下的诊断鲁棒性

    arXiv:2608.04885v1 Announce Type: cross Abstract: Standard accuracy metrics for VLMs often mask significant reliability failures in sensitive domains. In this work, we utilize a histopathology-validated brain MRI dataset to systematically assess the diagnostic robustness of four …

  31. arXiv cs.AI TIER_1 English(EN) · Cheng Yin, Yankai Lin, Wang Xu, Sikyuen Tam, Xiangrui Zeng, Zhiyuan Liu, Zhouping Yin ·

    DeepThinkVLA:增强视觉-语言-动作模型的推理能力

    arXiv:2511.15669v3 Announce Type: replace-cross Abstract: Does Chain-of-Thought (CoT) reasoning genuinely improve Vision Language Action (VLA) models, or does it merely add overhead? Existing CoT-VLA systems report limited and inconsistent gains, yet no prior work has rigorously …

  32. arXiv cs.AI TIER_1 English(EN) · Houze Xu, Jizhong Li, Ziyi Ye ·

    面向长时域规划的显式语言记忆在视觉-语言-动作模型中的应用

    arXiv:2608.04765v1 Announce Type: cross Abstract: Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language understanding, and robotic control. However, existing VLA models still face major challenges in long-horizon tasks: sparse e…

  33. arXiv cs.AI TIER_1 English(EN) · De Jiang, Zhengyang Zhang, Kehong Yuan, Shaohua Ma ·

    CARGO-VL:面向视觉语言模型的风险约束分组优化反事实仲裁

    arXiv:2608.04509v1 Announce Type: new Abstract: Vision-language systems combine images with retrieved text, but these sources can disagree or jointly fail to support an answer. Reliable models must identify the trustworthy source and abstain when neither is adequate. Existing pos…

  34. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越全局路由聚合:面向MoE视觉语言模型的阶段感知专家合并

    Mixture-of-experts vision-language models (MoE-VLMs) increase model capacity with sparse expert activation, yet deployment requires storing the full expert pool. Training-free expert merging reduces this burden, and many routing-based methods aggregate routing statistics across a…

  35. arXiv cs.AI TIER_1 English(EN) · Chunyang Jiang, Pingping Zhang, Yuzhi Zhao, Wenao Ma, Zhijian Hou, Mengyang Wu, Yiyang Cai, Senkang Hu, Sitong Cheng, Chi-Min Chan, Wei Xue, Yike Guo ·

    面向多模态大语言模型自我改进的基于失败信息的图像自增强

    arXiv:2608.03733v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimodal data that are costly to annotate. Self-augmentati…

  36. arXiv cs.AI TIER_1 English(EN) · Mohammad Rostami ·

    视觉语言模型中的上下文崩溃及其缓解方法?

    arXiv:2608.02830v1 Announce Type: cross Abstract: Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely assumed to improve as more demonstrations are supplied. We show the opposite: as d…

  37. arXiv cs.AI TIER_1 English(EN) · Jinquan Zhang, Dongfu Yin, Run Yang, Yufeng Yan, Zhen Tian, F. Richard Yu ·

    结构感知鲁棒微调:防御视觉-语言-动作机器人免受物理注意力劫持

    arXiv:2608.03231v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies promise general robotic manipulation, but their robustness against physical-world attacks remains fragile. In particular, we show that physically realizable adversarial patches can reliably in…

  38. arXiv cs.AI TIER_1 English(EN) · Tianbao Jiang, Weicong Ni, Gerard de Melo, Linlin Wang ·

    测试时对大型视觉语言模型进行对齐:一种轨迹引导的结构化采样方法

    arXiv:2608.03204v1 Announce Type: cross Abstract: Post-training reinforcement learning (RL) algorithms are commonly used to align large vision-language models (LVLMs) with human intent and the requirements of visual reasoning tasks. However, existing RL-based alignment methods ar…

  39. arXiv cs.AI TIER_1 English(EN) · Paribesh Regmi, Qingshuang Chen, Chi Zhang, Heba Aly, Yelin Kim, Hongda Mao ·

    面向视频-语言模型高效推理的自适应两阶段视觉令牌剪枝

    arXiv:2608.03112v1 Announce Type: cross Abstract: Vision-language models excel at image and video understanding but suffer from high inference latency due to the need to process thousands of tokens per image, limiting their deployment on resource-constrained edge devices and in r…

  40. arXiv cs.AI TIER_1 English(EN) · Inkyu Sa, Konstantin Stulov, Rajat Bhageria ·

    ValueFormer:一种具有阶段感知标签的因果Transformer价值函数,用于半自主视觉-语言-动作策略

    arXiv:2608.02958v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies trained by behavior cloning fail silently: from the action stream alone, a collapsing rollout looks much like one making clean progress, because imitation supplies no notion of progress. Reinf…

  41. arXiv cs.CL TIER_1 English(EN) · Haz Sameen Shahgir, Xiaofu Chen, Yu Fu, Erfan Shayegani, Nael Abu-Ghazaleh, Yova Kementchedjhieva, Yue Dong ·

    VLMs 需要语言:视觉语言模型为追求语义锚点而忽略视觉细节

    arXiv:2604.02486v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have achieved impressive performance across a wide range of multimodal tasks. However, they often fail on tasks that require fine-grained visual perception, even when the required information …

  42. arXiv cs.CL TIER_1 English(EN) · Jiashu Yao, Haoyu Wen, Siyuan Gao, Yuhang Guo, Zeming Liu, Heyan Huang ·

    HomeSafeBench:用于自由探索家庭安全检查的具身视觉语言模型的基准

    arXiv:2509.23690v2 Announce Type: replace-cross Abstract: Safety hazards in the home are a leading cause of preventable domestic injuries, motivating an automated inspector that actively explores a home and reports hazards before they cause harm. We introduce HomeSafeBench, the f…

  43. Hugging Face Daily Papers TIER_1 English(EN) ·

    BridgeVLA++:一种数据高效、可泛化且带记忆增强的3D操作视觉-语言-动作框架

    Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution shifts, and …

  44. Hugging Face Daily Papers TIER_1 English(EN) ·

    MuRA:用于高效有效测试时视觉语言泛化的多秩适应

    Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-time adaptation (TTA) via Low-Rank Adaptation offers a parameter-efficient solution, we identify a fundamental bottleneck in curre…

  45. Hugging Face Daily Papers TIER_1 English(EN) ·

    面向多模态大语言模型自改进的故障信息图像自增强

    Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimodal data that are costly to annotate. Self-augmentation offers a promising alternative by enabling mo…

  46. Hugging Face Daily Papers TIER_1 English(EN) ·

    视觉Token数量减少何时能加速多模态推理?一项跨决策位置和硬件的收支平衡研究

    Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead, shared work, and the operators each policy can avoid. A stage-level decomposition reconciles these components with measured end-…

  47. arXiv cs.CL TIER_1 English(EN) · Shalom Kachko, Raz Lapid, Margarita Vald, Almog Dubin, Moshe Sipper ·

    透过 LENS:视觉语言模型表征的局部几何分解

    arXiv:2608.00561v1 Announce Type: cross Abstract: Vision-language models (VLMs) process image patches and text tokens in a shared residual stream, but the local geometry through which the two modalities interact remains poorly understood. Most interpretability methods identify gl…

  48. arXiv cs.CL TIER_1 English(EN) · Jing Wu, Jianhua Wu, Jiayi Guan, Jiahong Chen, Jinghui Lu, Hangjun Ye, Bingzhao Gao, Long Chen ·

    SpatioLM:迈向视觉语言模型通用物理空间智能

    arXiv:2608.01899v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or external spatial encoders, which increase complexity …

  49. arXiv cs.CL TIER_1 English(EN) · Eugene Lee, Ting-Yu Chang, Jui-Huang Tsai, Jiajie Diao, Chen-Yi Lee ·

    Large Language Model 对视觉编码器进行分层预训练

    arXiv:2604.00086v2 Announce Type: replace-cross Abstract: The field of computer vision has experienced significant advancements through scalable vision encoders and multimodal pre-training frameworks. However, existing approaches often treat vision encoders and large language mod…

  50. arXiv cs.LG TIER_1 English(EN) · Mayank Nautiyal, Li Ju, Andreas Hellander, Ekta Vats, Prashant Singh ·

    GeoFlowVLM:几何感知联合不确定性用于冻结视觉语言嵌入

    arXiv:2605.13352v2 Announce Type: replace Abstract: Standard dual-encoder vision-language models that map images and text to deterministic points on a shared unit hypersphere through $\ell_2$ normalization typically expose neither \emph{aleatoric} uncertainty (cross-modal ambigui…

  51. arXiv cs.LG TIER_1 English(EN) · Leyan Xue, Feng Xiong, Mingjun Ma, Changqing Zhang ·

    提炼学生所见:面向视觉语言模型的 Fisher-Projected On-Policy Distillation

    arXiv:2608.01263v1 Announce Type: new Abstract: On-policy distillation (OPD) samples trajectories from the current student policy and minimizes token-level divergence between student and teacher next-token distributions at prefixes along those trajectories. This aligns the distil…

  52. arXiv cs.CL TIER_1 English(EN) · Jin Cui, Chuanchang Su, Jiayi Lu, Xinyue Long, Boran Zhao, Pengju Ren ·

    HAFI-VLM:从频率视角诊断和增强视觉语言模型中的视觉感知

    arXiv:2608.02124v1 Announce Type: cross Abstract: Vision-language models (VLMs) remain unreliable when predictions require fine-grained visual evidence. We identify a previously overlooked cause: spectral response rigidity. Despite substantial frequency variation across images an…

  53. Hugging Face Daily Papers TIER_1 English(EN) ·

    SIGNPOST-Bench:多模态大语言模型文本-视觉冲突解决基准测试

    Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled coun…

  54. Hugging Face Daily Papers TIER_1 English(EN) ·

    SpatialQuery:对视觉语言模型中基于几何的多实例空间推理进行基准测试

    Vision-language models (VLMs) achieve strong semantic understanding but remain unreliable in metric spatial reasoning, particularly when queries require comparing multiple instances of the same object category. We study this problem through the Closest-Instance Distance Query (CI…

  55. arXiv cs.LG TIER_1 English(EN) · Mushir Akhtar, M. Tanveer ·

    缓解临床漂移下医疗视觉语言模型中的类别尾部覆盖不足问题

    arXiv:2607.28696v1 Announce Type: new Abstract: Medical vision-language models (VLMs) can retain high observed marginal coverage after clinical shift while substantially under-covering an individual disease class. The affected class varies with acquisition protocol and backbone g…

  56. arXiv cs.CL TIER_1 English(EN) · Senyu Fei, Xiaopeng Yu, Siyin Wang, Xianzhong Zhao, Jingjing Gong, Xipeng Qiu ·

    WCM:用于视觉-语言-动作强化学习的世界批评模型

    arXiv:2607.29613v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on…

  57. arXiv cs.AI TIER_1 English(EN) · Md Ashikur Rahman, Md Arifur Rahman, Niamul Hassan Samin, Abdullah Ibne Hanif Arean, Juena Ahmed Noshin ·

    步级视觉定位忠实度预测长时域视觉语言模型在分布外泛化能力

    arXiv:2603.06828v2 Announce Type: replace-cross Abstract: We uncover a behavioral law of long-horizon vision-language models: models that maintain temporally grounded beliefs generalize better. Standard benchmarks measure only final-answer accuracy, which obscures how models use …

  58. Hugging Face Daily Papers TIER_1 English(EN) ·

    CRAFT:面向视觉语言模型的视频令牌递归自适应融合压缩

    In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the prefill stage to rise sharply. Such visual sequences are highly redundant along the spatio-temporal dimension, yet a high compressi…

  59. Hugging Face Daily Papers TIER_1 English(EN) ·

    内松外衡:面向视觉语言混合专家模型的几何引导负载均衡

    Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this token mix. We call the standard token-level Switch auxiliary loss Std-Aux. Std-Aux balances only the mixed load, so large image…

  60. arXiv cs.LG TIER_1 English(EN) · Chiyuan He, Zihuan Qiu, Fanman Meng, Runtong Zhang, Linfeng Xu, Qingbo Wu, Hongliang Li ·

    通过语义-几何保持实现视觉语言模型的持续学习

    arXiv:2603.12055v3 Announce Type: replace-cross Abstract: Continual learning of pretrained vision-language models (VLMs) is prone to catastrophic forgetting, yet current approaches adapt to new tasks without explicitly preserving the cross-modal semantic geometry inherited from p…

  61. arXiv cs.AI TIER_1 English(EN) · Zhe Liu, Quan Lu, Zhaohui Du, Zhe Wang, Huanbo Jin, Jiaming Gu, Qi Wang, Ting Xiao, Minting Pan, Dongzhan Zhou ·

    BioVLN:生物医学实验室视觉语言导航的模拟平台

    arXiv:2607.26914v1 Announce Type: cross Abstract: Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms are designed for household environments and treat a target as an object center or an arbit…

  62. arXiv cs.CL TIER_1 English(EN) · Yuming Yan, Kai Tang, Sihong Chen, Ke Xu, Dan Hu, Qun Yu, Pengfei Hu ·

    S-GRPO:大型视觉语言模型的统一训练后处理

    arXiv:2604.16557v2 Announce Type: replace-cross Abstract: Current post-training methodologies for adapting Large Vision-Language Models (LVLMs) generally fall into two paradigms: Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). Despite their prevalence, both approach…

  63. arXiv cs.LG TIER_1 English(EN) · Supratik Bhowal, Subhrajyoti Basu, Aritra Gir Mahanta, Anik Pal Chowdhury ·

    位置而非出处:在医学视觉语言模型中区分推理中介与谄媚

    arXiv:2607.27304v1 Announce Type: new Abstract: Medical vision-language models (VLMs) generate chain-of-thought (CoT) reasoning before answering clinical questions, but whether this reasoning causally influences predictions remains unclear. We present CoT-Mediate, a behavioral fr…

  64. arXiv cs.LG TIER_1 English(EN) · Yao Xiao, Reuben Tan, Zhen Zhu, Yuqun Wu, Jianfeng Gao, Derek Hoiem ·

    ReToken:一个用于改进视觉检索的视觉语言模型

    arXiv:2607.28627v1 Announce Type: cross Abstract: Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present Re…

  65. Hugging Face Daily Papers TIER_1 English(EN) ·

    WCM:用于视觉-语言-动作强化学习的世界批评模型

    Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or single-frame VLM bac…

  66. Hugging Face Daily Papers TIER_1 English(EN) ·

    扩大视觉语言模型规模不足以缓解偏见

    Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly understood at scale. We present the first large-scale empirical study of 194 publicly available VLMs, including 16 model families, c…

  67. Hugging Face Daily Papers TIER_1 English(EN) ·

    ReToken:一个用于改进视觉检索的视觉语言模型

    Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding trained as an …

  68. Hugging Face Daily Papers TIER_1 English(EN) ·

    RL^2-VLA:用于视觉-语言-动作模型的自适应RL潜在成分约束,带测试时缩放

    Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance often degrades on challenging and out-of-domain tasks. Recent test-time steering and scaling methods improve performance without extensive data collection and retraini…

  69. Hugging Face Daily Papers TIER_1 English(EN) ·

    MedARC:3D医学视觉语言模型的无训练自适应冗余视觉标记压缩

    Integrating 3D medical images with vision-language models (VLMs) holds substantial promise for computer-aided diagnosis. However, volumetric images generate prohibitively long visual-token sequences with considerable spatial and inter-slice redundancy. Existing token compression …

  70. arXiv cs.AI TIER_1 English(EN) · Minhyeok Lee, Chiyoung Kim, Chanhoe Gu, Seongrok Kim, Sanghyuk Roy Choi, Donghwan Hwang, Donghun Ryu, Seokhyun Kim ·

    CoTinyVLA:用于亚十亿参数视觉-语言-动作模型的思维链蒸馏

    arXiv:2607.25487v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness benchmark use three- to seven-billion-parameter backbones whose memory demands ca…

  71. arXiv cs.AI TIER_1 English(EN) · Zonghe Liu (University of Hong Kong), Shanyuan Jie (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences), Xiaoquan Sun (Huazhong University of Science and Technology), Chen Cao (University of Hong Kong), Zetian Xu (University of Hong … ·

    SAM3D-引导的以对象为中心的表示对齐,用于视觉-语言-动作模型

    arXiv:2607.25912v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, especially und…

  72. Hugging Face Daily Papers TIER_1 English(EN) ·

    HumanCLAW:视觉语言模型能否通过身体行动?

    Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed t…

  73. Hugging Face Daily Papers TIER_1 English(EN) ·

    TurboVLA:RTX 4090 上实现 32 Hz 实时视觉-语言-动作模型,VRAM 占用 <1 GB

    Vision-language-action (VLA) models commonly adopt an LLM-centric V to L to A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial compu…

  74. arXiv cs.AI TIER_1 English(EN) · Renjie Liang ·

    廉价探针预测3D-CT视觉-语言模型昂贵的训练成本

    arXiv:2607.22771v1 Announce Type: cross Abstract: Picking the frozen image encoder for a 3D~CT vision--language model (VLM), together with the token-compression scheme on top of it, is a search over many candidates. There are several encoders, several ways to compress their token…

  75. arXiv cs.AI TIER_1 English(EN) · Ali Ansari, Yasmin Mohammadi, Farnoush Nili, Parsa Esmaeilkhani, Longin Jan Latecki, Eduard Dragut ·

    ERUnderstand:评估视觉语言模型在结构化ER图上的表现

    arXiv:2607.24707v1 Announce Type: new Abstract: Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiting AI-assisted database engineering. We introduce ER…

  76. arXiv cs.AI TIER_1 English(EN) · Jun Ling, Tao Huang, Junzhuo Liu, Bowen Tang, Peng Wang ·

    GOTS:高分辨率视觉语言模型的高效贪婪正交令牌选择

    arXiv:2607.23913v1 Announce Type: new Abstract: Modern vision-language models (VLMs) increasingly rely on dynamic or high-resolution visual encoding, producing thousands of visual tokens that substantially increase downstream language-model inference cost. Existing token-reductio…

  77. arXiv cs.AI TIER_1 English(EN) · Ridwan Mahbub, Mohammed Saidul Islam, Md Tahmid Rahman Laskar, Mizanur Rahman, Mir Tafseer Nayeem, Enamul Hoque ·

    视觉-语言模型中的图表欺骗:从漏洞到缓解

    arXiv:2607.22600v1 Announce Type: new Abstract: Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or inverted axes, distorted aspect ratios, inappropriate encodings, and misleading color mappin…

  78. arXiv cs.LG TIER_1 English(EN) · Shuai Wang, Daoan Zhang, Zhe Tang, Hao Cheng, Jiaheng Wei ·

    具有噪声学生策略内自蒸馏的自增强视觉语言模型

    arXiv:2607.23125v1 Announce Type: new Abstract: Post-training enables vision-language models (VLMs) to understand human instructions and perform various downstream tasks. Current post-training methods usually rely on human-annotated data, distillation from external models, reinfo…

  79. arXiv cs.CL TIER_1 English(EN) · Khai-Nguyen Nguyen, Zixin Tang, Ankur Mali, Mary ALexandria Kelly ·

    如同双语婴儿:视觉基础化双语语言模型的优势

    arXiv:2210.05487v3 Announce Type: replace Abstract: Unlike most neural language models, humans learn language in a rich, multi-sensory and, often, multi-lingual environment. Current language models typically fail to fully capture the complexities of multilingual language use. We …

  80. arXiv cs.CL TIER_1 English(EN) · M M Asif Ferdous ·

    更大还是更便宜?规模与量化对图像退化下视觉语言模型不确定性信号的影响

    arXiv:2607.24440v1 Announce Type: cross Abstract: Vision-language models (VLMs) deployed on consumer hardware must decide when to answer and when to defer, and that decision depends on having a confidence signal that tracks correctness. A practitioner with a fixed memory budget f…

  81. arXiv cs.CL TIER_1 English(EN) · Sultan Alshehri, Zhantao Yang, Han Zhang, Marios Savvides ·

    相似性并非逻辑:面向双编码器视觉语言模型的因子化推理

    arXiv:2607.23052v1 Announce Type: cross Abstract: Dual-encoder vision-language models (VLMs) expose a similarity interface that enables zero-shot retrieval but fails compositional constraints: queries like "umbrella and no person" retrieve images containing both, even when concep…

  82. arXiv cs.AI TIER_1 English(EN) · Alexander G. Ororbia, Ankur Mali, Mary Alexandria Kelly, David Reitter ·

    如同婴儿:视觉情境下的神经语言习得

    arXiv:1805.11546v3 Announce Type: replace-cross Abstract: We examine the benefits of visual context in training neural language models to perform next-word prediction. A multi-modal neural architecture is introduced that outperform its equivalent trained on language alone with a …

  83. arXiv cs.AI TIER_1 English(EN) · Khang Nhat Hoang Vo ·

    从像素到提示词:视觉语言模型

    arXiv:2605.07544v3 Announce Type: replace Abstract: When you read a paper about a new Vision-Language Model today, it can be easy to forget how strange this idea would have sounded not so long ago. Teaching machines to see was already hard. Teaching them to read and generate lang…

  84. arXiv cs.AI TIER_1 English(EN) · Qing Yang, Xun Wang, Ziguan Wang, Zhenjiang Li, Hongqiang Wang, Dongdong Weng ·

    面向视觉-语言-动作操控的Real2Sim2Real:基于AMD ROCm的流水线

    arXiv:2607.22997v1 Announce Type: cross Abstract: Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the real world -- has emerged as the next major frontier for AI, echoed by industry leaders such as Jensen Huang (``the n…

  85. Hugging Face Daily Papers TIER_1 English(EN) ·

    看见还是认知?多模态大语言模型中的视觉上下文敏感性

    Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. However, they often fail on vision-centric tasks, especially when visual evidence conflicts with pretrained knowledge. We explore t…

  86. arXiv cs.CL TIER_1 English(EN) · Gwang Gook Lee, Kenan Emir Ak, Jay Mohta, Yan Xu, Dimitrios Dimitriadis ·

    视觉语言模型是阅读还是改写?关于视觉语言模型转录忠实度的探讨

    arXiv:2607.21617v1 Announce Type: cross Abstract: Vision Language Models (VLMs) are increasingly used in place of traditional OCR pipelines for document understanding. In this paper, we show they do not always act as faithful transcribers: when text is imperfect, they often tend …

  87. Hugging Face Daily Papers TIER_1 English(EN) ·

    PerceptionBench:评估多模态大语言模型中的原子视觉感知

    We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures i…

  88. Hugging Face Daily Papers TIER_1 English(EN) ·

    N_0-VTLA:使用潜在触觉令牌扩展视觉-触觉-语言-动作模型

    We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current vision-bas…

  89. Hugging Face Daily Papers TIER_1 English(EN) ·

    UltraViT:面向大型视觉语言模型、针对延迟优化的设备端视觉编码器

    Large Vision-Language Models (LVLMs) remain bottlenecked by massive computational footprints, precluding their deployment on resource-constrained edge devices. While efforts to compress LVLMs focus heavily on vision token reduction or smaller language models, the vision encoder i…

  90. arXiv cs.AI TIER_1 English(EN) · Dongbin Na ·

    何时基于推理的防护栏效率不高?ResponseGuard:一种用于实时审核的快速视觉语言防护

    arXiv:2607.21401v1 Announce Type: cross Abstract: A vision-language AI assistant returns its answer as a stream of generated tokens. Therefore, a safety guard that watches that answer has to keep up with the stream and stop a harmful reply before a user reads it. Recent vision-la…

  91. arXiv cs.CL TIER_1 English(EN) · Aditi Gupta, Yossi Gandelsman ·

    面向视觉语言模型模态顺序一致性的测试时训练

    arXiv:2607.20351v1 Announce Type: cross Abstract: We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently …

  92. arXiv cs.AI TIER_1 English(EN) · Karan Goyal, Afreen Hossain, Debojyoti Das, Vishal Bhutani ·

    ENTRAP-VL:视觉-语言模型中双重上下文纠缠的分类探测器

    arXiv:2607.20092v1 Announce Type: cross Abstract: Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mec…

  93. Hugging Face Daily Papers TIER_1 English(EN) ·

    面向视觉语言模型模态顺序一致性的测试时训练

    We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently outperforms question-first prompting, revealing a …

  94. arXiv cs.AI TIER_1 English(EN) · Gautam Rajendrakumar Gare, Jia Shi, Zhiqiu Lin, Deepak Pathak, John Galeotti, Deva Ramanan ·

    属性应来自图像,而非类别名称:面向视觉语言模型的分布条件属性选择

    arXiv:2607.18695v1 Announce Type: cross Abstract: A popular route to interpretable zero-shot classification asks a large language model (LLM) to describe each class name and prompts CLIP with the resulting descriptors. We show that these descriptors carry little visual evidence o…

  95. arXiv cs.AI TIER_1 English(EN) · Sibo Wang, Jie Zhang, Shiguang Shan, Xilin Chen, Wen Gao ·

    大型视觉语言模型鲁棒性增强的双对抗微调

    arXiv:2607.18958v1 Announce Type: cross Abstract: While Large Vision-Language Models (LVLMs), represented by LLaVA and GPT-4V, have demonstrated remarkable capabilities, their visual inputs remain vulnerable to adversarial attacks, posing significant security risks. Existing defe…

  96. Hugging Face Daily Papers TIER_1 English(EN) ·

    ENTRAP-VL:用于视觉语言模型中双重上下文纠缠的分类探测器

    Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whet…

  97. arXiv cs.CL TIER_1 English(EN) · Sudharshan Balaji, Yili Ren, Guangjing Wang, Yimin Chen, Ning Wang ·

    一种模态遗忘所有:增强视觉语言模型中的跨模态遗忘

    arXiv:2607.16442v1 Announce Type: cross Abstract: Machine unlearning is widely used to remove hazardous knowledge from large language models. Modern Vision-Language Models (VLMs), however, process both text and visual inputs, raising a fundamental security question: does unlearni…

  98. arXiv cs.AI TIER_1 English(EN) · Jianing Guo, Zhenhong Wu, Chang Tu, Yiyao Ma, Xiangqi Kong, Zhiqian Liu, Jiaming Ji, Shuning Zhang, Yuanpei Chen, Kai Chen, Qi Dou, Yaodong Yang, Xianglong Liu, Huijie Zhao, Weifeng Lv, Simin Li ·

    RobustVLA:关于视觉-语言-动作模型在多模态扰动下的鲁棒性研究

    arXiv:2510.00037v5 Announce Type: replace-cross Abstract: In Vision-Language-Actionf(VLA) models, robustness to real-world perturbations is critical for deployment. Existing methods target simple visual disturbances, overlooking the broader multi-modal perturbations that arise in…

  99. arXiv cs.AI TIER_1 English(EN) · Tuan Duong Trinh, Naveed Akhtar, Basim Azam ·

    推理的双刃剑:视觉-语言-动作模型中的架构与跨阶段鲁棒性

    arXiv:2607.17786v1 Announce Type: cross Abstract: Does adding a reasoning step make a Vision-Language-Action (VLA) model more robust to perturbation? Intuitively, a policy that reasons before acting should absorb a perturbed input better than one that maps observations directly t…

  100. arXiv cs.CL TIER_1 English(EN) · Tomer Ullman ·

    幻觉-幻觉:视觉语言模型看到不存在的幻觉

    arXiv:2412.18613v2 Announce Type: replace-cross Abstract: Illusions are entertaining, but they are also a useful diagnostic tool in cognitive science, philosophy, and neuroscience. A typical illusion shows a gap between how something `really is' and how something `appears to be',…

  101. arXiv cs.CL TIER_1 English(EN) · Wei Chen, Liangmin Wu, Yunhai Hu, Zhiyuan Li, Zhiyuan Cheng, Yicheng Qian, Lingyue Zhu, Zhipeng Hu, Luoyi Liang, Qiang Tang, Zhen Liu, Han Yang ·

    AutoNeural:为NPU推理共同设计视觉语言模型

    arXiv:2512.02924v3 Announce Type: replace Abstract: While Neural Processing Units (NPUs) offer high theoretical efficiency for edge AI, state-of-the-art Vision--Language Models (VLMs) tailored for GPUs often falter on these substrates. We attribute this hardware-model mismatch to…

  102. arXiv cs.AI TIER_1 English(EN) · Yijiang Li, Huiqi Zou, Bingyang Wang, Ziang Xiao ·

    通过动态、多轮交互对视觉语言模型进行情境化评估

    arXiv:2607.14499v1 Announce Type: new Abstract: Multi-modal Large Language Models (MLLMs) have made substantial advances on benchmarks, yet their real-world effectiveness remains uncertain. This gap stems from the fundamental misalignment between benchmarks in controlled, static …

  103. arXiv cs.AI TIER_1 English(EN) · Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman ·

    SceneBind:跨越视觉、音频和语言的“何处”与“何物”绑定

    arXiv:2607.15265v1 Announce Type: cross Abstract: We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what …

  104. arXiv cs.LG TIER_1 English(EN) · Pegah Khayatan, Sara Meziane, Jayneel Parekh, Matthieu Cord ·

    DiMaS:用于引导视觉-语言-动作模型的分布匹配

    arXiv:2607.14280v1 Announce Type: cross Abstract: Flow-matching-based vision-language-action (VLA) models have emerged as powerful policies for robotic manipulation, yet a critical capability remains underexplored: fine-grained behavioral control, the ability to govern how a robo…

  105. arXiv cs.AI TIER_1 English(EN) · Wei Li, Peijin Jia, Yuan Ma, Xuefeng Jiang, Titong Jiang, Sheng Sun, Yujian Li, Xin Wen, Han Hong, Zhikang Liu, Bailin Li, Kun Zhan ·

    FoMoVLA:为视觉-语言-动作模型搭建视觉预测与运动指导的桥梁

    arXiv:2607.14739v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have achieved impressive results in visuomotor policy learning, yet remain fundamentally reactive, mapping current observations and language to actions without explicit forward prediction of wor…

  106. arXiv cs.AI TIER_1 English(EN) · Eli Shlizerman ·

    SceneBind:跨越视觉、音频和语言的“何处”与“何物”绑定

    We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial struc…

  107. arXiv cs.AI TIER_1 English(EN) · Kun Zhan ·

    FoMoVLA:为视觉-语言-动作模型搭建视觉预测与运动引导的桥梁

    Vision-Language-Action (VLA) models have achieved impressive results in visuomotor policy learning, yet remain fundamentally reactive, mapping current observations and language to actions without explicit forward prediction of world dynamics. Existing visual foresight methods pre…

  108. arXiv cs.LG TIER_1 English(EN) · Charlotte Morissette, Amin Abyaneh, Wei-Di Chang, Anas Houssaini, David Meger, Hsiu-Chin Lin, Jonathan Tremblay, Gregory Dudek ·

    面向视觉-语言-动作模型的触觉模态融合

    arXiv:2603.14604v2 Announce Type: replace-cross Abstract: We propose TacFiLM, a lightweight modality-fusion approach that integrates visual-tactile signals into vision-language-action (VLA) models. While advances in VLAs have introduced robot policies that are both generalizable …

  109. arXiv cs.AI TIER_1 English(EN) · Jose Mart\'inez-Fajardo, Pablo Pueyo, Fernando Caballero, Luis Merino ·

    从语言到导航目标:一种使用RGB-D感知的移动机器人语义导航的视觉-语言方法

    arXiv:2607.13624v1 Announce Type: cross Abstract: Natural language interaction provides an intuitive way for non-expert users to communicate with robotic platforms. However, transforming user requests into executable navigation actions remains a challenging task, requiring the in…

  110. arXiv cs.AI TIER_1 English(EN) · Luyuan Jia, Yinfeng Yu ·

    HRO:用于大型语言模型零样本物体目标导航的层级房间到物体框架

    arXiv:2607.13072v1 Announce Type: cross Abstract: Zero-shot object-goal navigation aims to enable an intelligent agent to explore and navigate to objects of unknown categories in an unfamiliar environment without specific target training. In zero-shot navigation tasks, pre-traine…

  111. Hugging Face Daily Papers TIER_1 English(EN) ·

    从语言到导航目标:一种使用RGB-D感知的移动机器人语义导航的视觉-语言方法

    Natural language interaction provides an intuitive way for non-expert users to communicate with robotic platforms. However, transforming user requests into executable navigation actions remains a challenging task, requiring the integration of language understanding, environment p…

  112. arXiv cs.AI TIER_1 English(EN) · Luis Merino ·

    从语言到导航目标:一种使用RGB-D感知的移动机器人语义导航的视觉-语言方法

    Natural language interaction provides an intuitive way for non-expert users to communicate with robotic platforms. However, transforming user requests into executable navigation actions remains a challenging task, requiring the integration of language understanding, environment p…

  113. arXiv cs.AI TIER_1 English(EN) · Jiawei Liang, Jianjie Huang, Ruoyu Chen, Xianghao Jiao, Siyuan Liang, Shiming Liu, Xiaochun Cao ·

    多模态大语言模型视觉归因的证据重组与预测上下文残差化

    arXiv:2509.22415v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains difficult to inspect. Recent logit-lens attribution methods project each visual-token…

  114. arXiv cs.LG TIER_1 English(EN) · Yuanjie Lu, Beichen Wang, Zhengqi Wu, Yang Li, Xiaomin Lin, Chengzhi Mao, Xuesu Xiao ·

    APPLV: 从视觉-语言-动作模型中学习自适应规划器参数

    arXiv:2603.08862v2 Announce Type: replace-cross Abstract: Autonomous navigation in highly constrained environments remains challenging for mobile robots. Classical navigation approaches offer safety assurances but require environment-specific parameter tuning; end-to-end learning…

  115. arXiv cs.AI TIER_1 English(EN) · Hiroto Osaka, Shohei Taniguchi, Gouki Minegishi, Kai Yamashita, Masahiro Suzuki, Yutaka Matsuo ·

    视觉语言模型推理中的视觉访问边界

    arXiv:2607.12815v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires conti…

  116. arXiv cs.AI TIER_1 English(EN) · Peiyu Li, Xiaobao Huang, Ting Hua, Nitesh V. Chawla ·

    CrochetBench:视觉语言模型能否从描述转向克罗歇特领域的实践?

    arXiv:2511.09483v3 Announce Type: replace Abstract: While multimodal large language models can describe visual content, their ability to generate executable procedures remains underexplored. CrochetBench presented in this paper evaluates this shift from describing to doing throug…

  117. arXiv cs.AI TIER_1 English(EN) · Maijunxian Wang, Yijiang Li, Bingyang Wang, Tianwei Zhao, Ran Ji, Qingying Gao, Emmy Liu, Hokin Deng, Dezhi Luo ·

    Egocentric Bias in Vision-Language Models

    arXiv:2602.15892v2 Announce Type: replace-cross Abstract: Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce FlipSet, a diagnostic benchmark for Level-2 visual perspective taking (L2 VPT) in visio…

  118. Hugging Face Daily Papers TIER_1 English(EN) ·

    视觉语言模型推理中的视觉访问边界

    Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires continued access to image tokens, or whether it mainl…

  119. arXiv cs.AI TIER_1 English(EN) · Yutaka Matsuo ·

    视觉语言模型推理中的视觉访问边界

    Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires continued access to image tokens, or whether it mainl…

  120. Hugging Face Daily Papers TIER_1 English(EN) ·

    CoRe:视觉语言模型跨图像比较推理的综合框架

    Cross-image comparative reasoning remains challenging for vision-language models (VLMs), especially when correct prediction requires fine-grained attribute grounding and globally consistent reasoning. We present CoRe, a unified framework for this problem. CoRe includes: (i) CoRe-…

  121. arXiv cs.AI TIER_1 English(EN) · Mingjie Xie, Guangjun He, Dongli Xu, Youtian Lin, Hongjue Li, Pengming Feng, Jian Guan, Yue Deng ·

    SynCLIP:同义词一致的语言-图像预训练,用于鲁棒的开放词汇密集感知

    arXiv:2607.11008v1 Announce Type: cross Abstract: Open-vocabulary dense perception (OVDP) aims to localize objects unseen during training by leveraging textual knowledge. Despite the remarkable progress of recent CLIP-based approaches, we identify a critical limitation: synonym-i…

  122. arXiv cs.AI TIER_1 English(EN) · Liuyi Wang, Kai Sheng, Zongtao He, Jinlong Li, Yongrui Qin, Haojie Dai, Xiangyi Wang, Jingwei Yang, Qingqing Yan, Chengju Liu, Qijun Chen ·

    具身视觉语言导航的全面调查与系统性现实世界评估

    arXiv:2607.09792v1 Announce Type: cross Abstract: Navigation is a fundamental capability of autonomous systems, yet most existing approaches rely on highly structured models and strong prior assumptions, limiting their robustness in open and uncertain real-world environments. Vis…

  123. arXiv cs.AI TIER_1 English(EN) · Byungkun Lee, Dongyoon Hwang, Dongjin Kim, Hojoon Lee, Minho Park, Jaegul Choo ·

    像机器人一样看:面向视觉-语言-动作模型的机器人中心点图

    arXiv:2607.11498v1 Announce Type: cross Abstract: Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, c…

  124. arXiv cs.AI TIER_1 English(EN) · Murad Farzulla ·

    Vision-Language模型中的上下文相关可供性计算

    arXiv:2603.04419v2 Announce Type: replace-cross Abstract: We characterize the phenomenon of context-dependent affordance computation in vision-language models (VLMs). Our primary study uses Qwen3-VL-30B-A3B ($n = 3{,}213$ scene-context pairs from COCO-2017: 479 images under 7 age…

  125. arXiv cs.AI TIER_1 English(EN) · Hayeon Kim, Ji Ha Jang, Junghun James Kim, Se Young Chun ·

    超乎寻常的视觉语言模型中的基于不确定性的组合对齐与部分到整体的语义代表性

    arXiv:2603.22042v3 Announce Type: replace-cross Abstract: While Vision-Language Models (VLMs) have achieved remarkable performance, their Euclidean embeddings remain limited in capturing hierarchical relationships such as part-to-whole or parent-child structures, and often face c…

  126. arXiv cs.AI TIER_1 English(EN) · Shengzhuo Yang, Ronghao Yu, Chuanjie Lv, Linpeng Peng, Hang Yu, Jie Ren, Jiajun Lv, Yong Liu ·

    TS-Mask VLA:用于具有有效桥接的视觉-语言-动作模型的二维时空掩码

    arXiv:2607.09818v1 Announce Type: cross Abstract: Vision-language-action (VLA) models aim to understand natural-language instructions and visual observations, and to generate and execute corresponding actions as embodied agents. Recently, autoregressive token-based action generat…

  127. arXiv cs.AI TIER_1 English(EN) · Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, … ·

    ABot-N1:迈向通用视觉语言导航基础模型

    arXiv:2607.10383v1 Announce Type: cross Abstract: Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic polici…

  128. arXiv cs.LG TIER_1 English(EN) · Finn Ferchau, Daniel Pommer, Cristian Axenie ·

    面向工业机器人操作的视觉-语言-动作模型LoRA微调效率研究

    arXiv:2607.10172v1 Announce Type: cross Abstract: Deploying billion-parameter Vision-Language-Action (VLA) models on industrial hardware requires fine-tuning to bridge the embodiment gap. Full Fine-Tuning (FFT) provides maximal plasticity but requires data centre-grade GPUs. We p…

  129. arXiv cs.AI TIER_1 English(EN) · Jaegul Choo ·

    像机器人一样看:面向视觉-语言-动作模型的机器人中心点图

    Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene i…

  130. arXiv cs.AI TIER_1 English(EN) · Shravan Murlidaran, Miguel P. Eckstein ·

    十年间视觉-语言AI模型准确性与视觉认知错误的演变

    arXiv:2607.09654v1 Announce Type: cross Abstract: Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactions or behaviors, only a handfu…

  131. arXiv cs.LG TIER_1 English(EN) · Xingyu Zhu, Huanshen Wu, Shuo Wang, Beier Zhu, Jiannan Ge, Jiaheng Zhang, Long Chen ·

    通过测试时提示适应来增强视觉语言模型

    arXiv:2607.09450v1 Announce Type: cross Abstract: Pre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations. Existing test-time adaptation methods typically rely on sample…

  132. arXiv cs.AI TIER_1 English(EN) · Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen ·

    Scalable Visual Pretraining for Language Intelligence

    arXiv:2607.09657v1 Announce Type: cross Abstract: The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations…

  133. Hugging Face Daily Papers TIER_1 English(EN) ·

    ABot-N1:迈向通用视觉语言导航基础模型

    Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map observations directly to actions, yet …

  134. arXiv cs.AI TIER_1 English(EN) · Kai Chen ·

    Scalable Visual Pretraining for Language Intelligence

    The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that can…

  135. arXiv cs.AI TIER_1 English(EN) · Miguel P. Eckstein ·

    十年间视觉-语言AI模型准确性与视觉认知错误的演变

    Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactions or behaviors, only a handful of non-curated human descriptions as a benchmark…

  136. arXiv cs.LG TIER_1 English(EN) · Long Chen ·

    通过测试时提示自适应增强视觉语言模型

    Pre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations. Existing test-time adaptation methods typically rely on sample-level confidence heuristics, overlooking the intr…

  137. arXiv cs.AI TIER_1 English(EN) · Qi Lyu, Baicheng Liu, Xudong Wang, Jiahua Dong, Lianqing Liu, Zhi Han ·

    LEEVLA:在视觉-语言-动作的潜在环境演化中识别关键要素

    arXiv:2607.08182v1 Announce Type: cross Abstract: Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all visual tokens uniformly and reasoning with human-…

  138. arXiv cs.AI TIER_1 English(EN) · Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu, Vaidehi Patil ·

    跨视觉、语言、视频和音频的多模态遗忘:方法、数据集和基准的调查

    arXiv:2607.07907v1 Announce Type: cross Abstract: With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associations that originate from their training data. Retrai…

  139. arXiv cs.AI TIER_1 English(EN) · Emily Jin, Joy Hsu, Yiqing Xu, Weiyu Liu, Nick Haber, Jiajun Wu ·

    APIVOT:具有交错视觉-语言思考的自适应规划

    arXiv:2607.08024v1 Announce Type: cross Abstract: Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility. To successfully execute a task, a robot must decompose goals, select task-relevant objects, and sequence actions, while…

  140. Hugging Face Daily Papers TIER_1 English(EN) ·

    Scalable Visual Pretraining for Language Intelligence

    The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that can…

  141. Hugging Face Daily Papers TIER_1 English(EN) ·

    LEEVLA:在视觉-语言-动作的潜在环境演化中识别关键要素

    Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all visual tokens uniformly and reasoning with human-selected factors, which lack mechanisms to emphasi…

  142. arXiv cs.AI TIER_1 English(EN) · Zhi Han ·

    LEEVLA:在视觉-语言-动作的潜在环境演化中识别关键要素

    Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all visual tokens uniformly and reasoning with human-selected factors, which lack mechanisms to emphasi…

  143. arXiv cs.AI TIER_1 English(EN) · Juyi Lin, Amir Taherin, Arash Akbari, Arman Akbari, Lei Lu, Guangyu Chen, Taskin Padir, Xiaomeng Yang, Weiwei Chen, Yiqian Li, Xue Lin, David Kaeli, Pu Zhao, Yanzhi Wang ·

    VOTE:轨迹集成投票的视觉-语言-动作优化

    arXiv:2507.05116v5 Announce Type: replace-cross Abstract: Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic manipulation tasks guided by natural language. However, current VLA models suffer from two drawbacks: (i) generation of mass…

  144. arXiv cs.AI TIER_1 English(EN) · Chethan Krishnamurthy Ramanaik, Tobias Callies, Michael Hecht, Eirini Ntoutsi ·

    从中间谱子空间视角看视觉-语言模型的对抗脆弱性

    arXiv:2607.07375v1 Announce Type: cross Abstract: Adversarial vulnerability in deep neural networks (DNNs) has been studied from the perspectives of decision-boundary geometry, feature robustness, input-output Jacobians, and the instability of inverse problems. Here, we focus on …

  145. arXiv cs.CL TIER_1 English(EN) · Kaito Watanabe, Taisei Yamamoto, Tomoki Doi, Hitomi Yanaka ·

    视觉语言模型使用空间指示性表达的多语言能力评估

    arXiv:2607.07251v1 Announce Type: new Abstract: One of the expected abilities of vision-language models (VLMs) is spatial reasoning ability based on a given text and image. To evaluate the spatial reasoning abilities of VLMs, we focus on the use of spatial deictic expressions, wh…

  146. arXiv cs.AI TIER_1 English(EN) · Inkyu Sa, Chanoh Park, Hea-Min Lee, Donghee Noh, Ho Seok Ahn ·

    用于无人机机器人和双臂操作的视觉语言动作(VLA)模型:一篇综述

    arXiv:2607.06706v1 Announce Type: cross Abstract: Vision Language Action (VLA) models unify visual perception, natural-language understanding, and action generation within a single foundation model, allowing a robot to follow instructions such as fold the towel or fly to the red …

  147. arXiv cs.AI TIER_1 English(EN) · Peter Bohm, Saimunur Rahman, Abdelwahed Khamis, Sagun Man Singh Shrestha, Chris McCool, Peyman Moghadam ·

    GemNav:使用多模态大语言模型的离散标记视觉机器人导航

    arXiv:2607.06882v1 Announce Type: cross Abstract: Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke action head, and training on thousands of hours of cross-embodiment datasets. We ask whether t…

  148. arXiv cs.AI TIER_1 English(EN) · Zhiying Du, Bei Liu, Yaobo Liang, Yichao Shen, Haidong Cao, Xiangyu Zheng, Zhiyuan Feng, Zuxuan Wu, Jiaolong Yang, Yu-Gang Jiang ·

    HiMoE-VLA:用于通用视觉-语言-动作策略的层次化混合专家模型

    arXiv:2512.05693v2 Announce Type: replace-cross Abstract: Generalist vision--language--action (VLA) policies are typically trained on heterogeneous mixtures of robot demonstrations spanning diverse embodiments, action spaces, and observation configurations. Modeling such heteroge…

  149. arXiv cs.CL TIER_1 English(EN) · Vaidehi Patil ·

    跨越视觉、语言、视频和音频的多模态遗忘:方法、数据集和基准的调查

    With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associations that originate from their training data. Retraining after deletion requests or policy updates is …

  150. arXiv cs.AI TIER_1 English(EN) · Eirini Ntoutsi ·

    从中间谱子空间视角看视觉-语言模型的对抗脆弱性

    Adversarial vulnerability in deep neural networks (DNNs) has been studied from the perspectives of decision-boundary geometry, feature robustness, input-output Jacobians, and the instability of inverse problems. Here, we focus on the spectral structure of intermediate linear tran…

  151. arXiv cs.CL TIER_1 English(EN) · Hitomi Yanaka ·

    视觉语言模型使用空间指示语的多语能力评估

    One of the expected abilities of vision-language models (VLMs) is spatial reasoning ability based on a given text and image. To evaluate the spatial reasoning abilities of VLMs, we focus on the use of spatial deictic expressions, which are defined as spatial expressions whose ref…

  152. arXiv cs.LG TIER_1 English(EN) · Haitao Lin, Hanyang Yu, Jingshun Huang, He Zhang, Yonggen Ling, Ping Tan, Xiangyang Xue, Yanwei Fu ·

    PoseVLA: 通用姿态预训练,用于可泛化的视觉-语言-动作策略

    arXiv:2602.19710v3 Announce Type: replace-cross Abstract: Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision. Since these model…

  153. arXiv cs.AI TIER_1 English(EN) · Sohwi Lim, Lee Hyoseok, Jungjoon Park, Tae-Hyun Oh ·

    CLAY: 视觉-语言嵌入空间中的条件视觉相似度调制

    arXiv:2604.11539v2 Announce Type: replace-cross Abstract: Human perception of visual similarity is inherently adaptive and subjective, depending on the users' interests and focus. However, most image retrieval systems fail to reflect this flexibility, relying on a fixed, monolith…

  154. arXiv cs.CL TIER_1 English(EN) · Liang Chen, Weichu Xie, Yiyan Liang, Hongfeng He, Hans Zhao, Zhibo Yang, Zhiqi Huang, Haoning Wu, Haoyu Lu, Y. charles, Yiping Bao, Yuantao Fan, Guopeng Li, Haiyang Shen, Xuanzhong Chen, Wendong Xu, Shuzheng Si, Zefan Cai, Wenhao Chai, Ziqi Huang, Fangfu… ·

    BabyVision: 超越语言的视觉推理

    arXiv:2601.06521v2 Announce Type: replace-cross Abstract: While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a cruc…

  155. arXiv cs.LG TIER_1 English(EN) · Ryuji Oi, Hikari Otsuka, Kosuke Matsushima, Yuki Ichikawa, Masato Motomura, Tatsuya Kaneko, Daichi Fujiki ·

    面向视觉-语言-动作模型的无训练加速:动作缓存与精炼

    arXiv:2607.06370v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow matching-based VLA models have shown remarkable success due to their capability to generate prec…

  156. Hugging Face Daily Papers TIER_1 English(EN) ·

    GemNav:使用多模态大语言模型的离散标记视觉机器人导航

    Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke action head, and training on thousands of hours of cross-embodiment datasets. We ask whether this recipe is necessary. In this paper, we introdu…

  157. Hugging Face Daily Papers TIER_1 English(EN) ·

    面向机器人操作的双重潜在记忆视觉-语言-动作模型

    LaMem-VLA introduces a latent-memory-native framework that integrates historical experience into vision-language-action reasoning through coordinated memory components operating in the same latent space.

  158. arXiv cs.LG TIER_1 English(EN) · Daichi Fujiki ·

    面向视觉-语言-动作模型的无训练加速:动作缓存与精炼

    Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow matching-based VLA models have shown remarkable success due to their capability to generate precise and smooth action sequences and capture multim…

  159. Hugging Face Daily Papers TIER_1 English(EN) ·

    面向视觉-语言-动作模型的无训练加速:动作缓存与精炼

    Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow matching-based VLA models have shown remarkable success due to their capability to generate precise and smooth action sequences and capture multim…

  160. arXiv cs.AI TIER_1 English(EN) · Suhyeok Jang, Dongyoung Kim, Changyeon Kim, Youngsuk Kim, Jinwoo Shin ·

    面向视觉-语言-动作模型的无验证器测试时采样

    arXiv:2510.05681v2 Announce Type: replace-cross Abstract: Vision-Language-Action models (VLAs) have demonstrated remarkable performance in robot control. However, they remain fundamentally limited in tasks that require high precision due to their single-inference paradigm. While …

  161. arXiv cs.AI TIER_1 English(EN) · Qingqian Yang, Hao Wang, Sai Qian Zhang, Jian Li, Yang Hua, Miao Pan, Tao Song, Zhengwei Qi, Haibing Guan ·

    pFedNavi:面向具身AI的结构感知个性化联邦视觉-语言导航

    arXiv:2602.14401v2 Announce Type: replace-cross Abstract: Vision-Language Navigation VLN requires large-scale trajectory instruction data from private indoor environments, raising significant privacy concerns. Federated Learning FL mitigates this by keeping data on-device, but va…

  162. arXiv cs.CL TIER_1 English(EN) · Israfel Salazar, Stella Frank, Dan Oneata, Desmond Elliott, Constanza Fierro ·

    视觉-语言模型中的视觉信息流路径

    arXiv:2607.03358v1 Announce Type: cross Abstract: We study how visual information is routed in vision-language models (VLMs). Using causal patching on controlled synthetic and natural datasets, we find that models rely on two distinct pathways to solve visual tasks: A direct path…

  163. arXiv cs.CL TIER_1 English(EN) · Khang Nhat Hoang Vo, Artem Vazhentsev, Artem Shelmanov, Timothy Baldwin, Yova Kementchedjhieva ·

    是“看不见”还是“不知道”?探究视觉语言模型中的错误归因

    arXiv:2607.04683v1 Announce Type: cross Abstract: Vision-language models (VLMs) perform well on visual question answering with high-quality images but struggle when questions require knowledge beyond what is clearly and directly visible. In such settings, uncertainty quantificati…

  164. arXiv cs.LG TIER_1 English(EN) · Harsh Goel, S P Sharan, Sahil Shah, Minkyu Choi, Joungbin An, Kristen Grauman, Sandeep P. Chinchali ·

    激励视觉语言模型搜索长视频问答

    arXiv:2607.02959v1 Announce Type: cross Abstract: We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process. VSeek utilizes a natural language-driven search to iden…

  165. arXiv cs.LG TIER_1 English(EN) · Zelin Zhao, Min Shi, Bo Yuan, Haotian Xue, Jialuo Li, Lama Moukheiber, Humphrey Shi, Yongxin Chen ·

    WorldBagel:揭示统一多模态模型在视觉-语言-动作-世界建模中的强大能力

    arXiv:2607.03461v1 Announce Type: cross Abstract: World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in Vision-Language-Action-World (VLAW) modeling. Meanwhile, unified vision-langu…

  166. arXiv cs.LG TIER_1 English(EN) · Jaeyoung Kim, Eunseok Kim, Dongsuk Jang ·

    几何在起作用吗?双曲视觉语言模型中层级结构的运行点审计

    arXiv:2607.05268v1 Announce Type: cross Abstract: Whether a hyperbolic representation model uses its geometry cannot be read off its curvature parameter: what matters is the dimensionless operating point $\sqrt{c}\rho$ and whether the radial and cone machinery is active there. We…

  167. arXiv cs.AI TIER_1 English(EN) · Suhyeong Park, Junha Jung, Jungwoo Park, Jaewoo Kang ·

    所有视觉标记都同等重要吗?面向视觉语言检索的保留对象证据的标记合并

    arXiv:2607.04605v1 Announce Type: cross Abstract: Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce t…

  168. arXiv cs.AI TIER_1 English(EN) · Riccardo Renzulli, Gabriele Spadaro, Shruthi Gowda, Alaa Eddine Mazouz, Van-Tam Nguyen ·

    TORINO:通过视觉语言模型中可解释的概念重叠实现令牌缩减

    arXiv:2607.04593v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated impressive capabilities across different tasks, but their computational cost is dominated by the large number of visual tokens fed to the language model. Existing token reduction met…

  169. arXiv cs.AI TIER_1 English(EN) · Xinchuan Qiu, Yi Yu ·

    用于视觉-语言-动作学习的简单到复杂结构化演示

    arXiv:2607.04591v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation by integrating visual perception, language understanding, and robot action generation. Existing research has primarily focused on im…

  170. arXiv cs.AI TIER_1 English(EN) · Kai Tang, Jinhao You, Bohua Zhang, Yichen Guo, Yiding Sun, Dongxu Zhang, Chenxi Li, Xiande Huang, Shanghang Zhang ·

    SeeMe:通过有效的视觉令牌工程减轻大型视觉语言模型的幻觉

    arXiv:2607.04163v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual understanding tasks such as image captioning and visual question answering. However, they remain susceptible to hallucinations, generating content th…

  171. arXiv cs.AI TIER_1 English(EN) · Seung Il Lee, Qinqian Lei, Daguang Xu, Dong Yang, Robby T. Tan, Yixin Chen, Bo Wang ·

    基于Token的意图识别与大型视觉语言模型

    arXiv:2607.03595v1 Announce Type: cross Abstract: Affordance grounding aims to localize image regions that support a specific action, serving as a core capability for physical intelligence and embodied perception. Previous studies have primarily relied on weakly supervised learni…

  172. arXiv cs.AI TIER_1 English(EN) · Li Ji, Siyin Wang, Pengfang Qian, Xiaopeng Yu, Yihai Tian, Zhaoye Fei, Jingjing Gong, Xipeng Qiu ·

    HiMe:用于长时域视觉-语言-动作控制的分层具身记忆

    arXiv:2607.03449v1 Announce Type: cross Abstract: Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian tasks requiring long-term memory and reasoning due to their reliance on immediate observations. Existing solutions fac…

  173. arXiv cs.AI TIER_1 English(EN) · Qi Liu, Yabei Li, Hongsong Wang, Heng Zhang, Lei He ·

    AnchorVLA:连接离散决策与连续轨迹,实现视觉-语言-动作规划

    arXiv:2607.03182v1 Announce Type: cross Abstract: Autonomous driving planning requires translating navigation intent, traffic rules, dynamic interactions, and language instructions into executable continuous trajectories. Vision-Language-Action models have been introduced into dr…

  174. arXiv cs.AI TIER_1 English(EN) · Chengzhen Yu, Canran Xiao, Siyuan Ma, Yang Liu ·

    文本作为部分约束:核心-残差对齐实现鲁棒的视觉语言学习

    arXiv:2607.03143v1 Announce Type: cross Abstract: Vision-language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified, making similarity brittle and overly confident under paraphrase and omitted details. We ai…

  175. arXiv cs.AI TIER_1 English(EN) · Junchi Liao, Jiawen Deng, Fuji Ren ·

    VISTA:审计视觉-语言模型中的语义分歧

    arXiv:2607.02995v1 Announce Type: cross Abstract: Vision-language models can exhibit visual concept-conditioned divergence: given images containing demographic features, corporate logos, or ideological symbols, some models produce unusually uniform responses that differ from what…

  176. arXiv cs.AI TIER_1 English(EN) · Zikai Zhang, Rui Hu, Olivera Kotevska, Jiahao Xu ·

    大型视觉语言模型云边推理中的视觉令牌操纵攻击

    arXiv:2607.02819v1 Announce Type: cross Abstract: Cloud-edge Large Vision-Language Model (LVLM) inference enables efficient deployment by splitting computation between edge devices and cloud servers. In this process, intermediate vision tokens are transmitted from the edge to the…

  177. arXiv cs.AI TIER_1 English(EN) · Kaiyun Yang, Ruilin Yang, Zhimin Yao, J. Wang, Wei Ge ·

    Criterion-Conditional In-Context Learning: Evaluating Criterion-Shift Adaptation in Vision-Language Models

    arXiv:2607.02575v1 Announce Type: cross Abstract: Vision-language models can perform new tasks without parameter updates through in-context learning (ICL), whose core mechanism is utilizing the support set for task induction. In the standard ICL setting, once the task is induced,…

  178. arXiv cs.AI TIER_1 English(EN) · Guli Zhu, Chenwei Wu, Liyue Shen ·

    评估与理解用于医学视觉语言模型的模型编辑技术

    arXiv:2607.05310v1 Announce Type: new Abstract: Model editing promises a fast, targeted way to correct post-deployment mistakes in medical vision-language models (VLMs) without costly retraining. However, existing multimodal model editing benchmarks focus on general-purpose tasks…

  179. arXiv cs.AI TIER_1 English(EN) · Mansi Phute, Ravikumar Balakrishnan ·

    VISOR:基于视觉输入的输出重定向转向技术,用于视觉语言模型

    arXiv:2508.08521v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) are increasingly being used in a broad range of applications, bringing their security and behavioral control to the forefront. While existing approaches for behavioral control or output redire…

  180. arXiv cs.AI TIER_1 English(EN) · Ravikumar Balakrishnan, Mansi Phute ·

    VISOR++:基于通用视觉输入的大型视觉语言模型转向

    arXiv:2509.25533v2 Announce Type: replace-cross Abstract: As Vision Language Models (VLMs) are deployed across safety-critical applications, understanding and controlling their behavioral patterns has become increasingly important. Existing behavioral control methods face signifi…

  181. arXiv cs.AI TIER_1 English(EN) · George-Andrei Dima, R\u{a}zvan-Alexandru Sm\u{a}du, Dumitru-Clementin Cercel ·

    面向罗马尼亚视觉语言模型的参数高效多模态指令调优

    arXiv:2512.14926v2 Announce Type: replace-cross Abstract: Focusing on low-resource languages is an essential step toward democratizing generative AI. In this work, we contribute to reducing the multimodal NLP resource gap for Romanian. We translate the widely known Flickr30K data…

  182. arXiv cs.AI TIER_1 English(EN) · Liyue Shen ·

    评估与理解用于医学视觉语言模型的模型编辑技术

    Model editing promises a fast, targeted way to correct post-deployment mistakes in medical vision-language models (VLMs) without costly retraining. However, existing multimodal model editing benchmarks focus on general-purpose tasks and do not reflect realistic clinical domain re…

  183. arXiv cs.LG TIER_1 English(EN) · Dongsuk Jang ·

    几何在起作用吗?双曲视觉语言模型中层级结构的运行点审计

    Whether a hyperbolic representation model uses its geometry cannot be read off its curvature parameter: what matters is the dimensionless operating point $\sqrt{c}ρ$ and whether the radial and cone machinery is active there. We develop a battery of necessary-condition diagnostics…

  184. Hugging Face Daily Papers TIER_1 English(EN) ·

    SteelBench:在真实工业环境中评估视觉语言模型

    Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test vision-language models under the visual and procedural conditions of real industrial CCTV, where workers appear as distant figur…

  185. arXiv cs.CL TIER_1 English(EN) · Yova Kementchedjhieva ·

    是看不见还是不知道?探究视觉-语言模型中的错误归因

    Vision-language models (VLMs) perform well on visual question answering with high-quality images but struggle when questions require knowledge beyond what is clearly and directly visible. In such settings, uncertainty quantification should not only indicate whether the model is l…

  186. Hugging Face Daily Papers TIER_1 English(EN) ·

    所有视觉标记都同等重要吗?面向视觉语言检索的对象-证据保留标记合并

    Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce this cost, yet they can remove or collapse object- …

  187. arXiv cs.CL TIER_1 English(EN) · Jaewoo Kang ·

    所有视觉标记都同等重要吗?用于视觉-语言检索的对象-证据保留标记合并

    Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce this cost, yet they can remove or collapse object- …

  188. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Jaewoo Kang ·

    所有视觉标记都同等重要吗?用于视觉语言检索的对象-证据保留标记合并

    Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce this cost, yet they can remove or collapse object- …

  189. Hugging Face Daily Papers TIER_1 English(EN) ·

    用于视觉-语言-动作学习的简单到复杂结构化演示

    Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation by integrating visual perception, language understanding, and robot action generation. Existing research has primarily focused on improving model architectures, training strategies, …

  190. Hugging Face Daily Papers TIER_1 English(EN) ·

    所有视觉标记都同等重要吗?用于视觉语言检索的对象-证据保留标记合并

    Object-aware token merging framework SaMer compresses image-side tokens while preserving query-selectable visual evidence, achieving significant storage reduction and improved retrieval performance.

  191. arXiv cs.AI TIER_1 English(EN) · Sung June Kim, Sangpil Kim, Honglak Lee ·

    用于视觉语言导航中语义探索的路径级事后指导

    arXiv:2607.01754v1 Announce Type: new Abstract: On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader state distribution. However, such exploration inevitably leads to trajectories that deviate fr…

  192. arXiv cs.AI TIER_1 English(EN) · Ninghao Zhang, Bin Zhu, Shijie Zhou, Jingjing Chen ·

    通过无训练注意力重校准恢复VLA模型中的语言基础

    arXiv:2603.06001v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models enable robots to perform manipulation tasks directly from natural language instructions and are increasingly viewed as a foundation for generalist robotic policies. However, their reliab…

  193. arXiv cs.AI TIER_1 English(EN) · Phillip Howard, Xin Su, Kathleen C. Fraser ·

    大型视觉语言模型中的跨文化价值归因

    arXiv:2604.09945v2 Announce Type: replace-cross Abstract: The rapid adoption of large vision-language models (LVLMs) in recent years has been accompanied by growing fairness concerns due to their propensity to reinforce harmful societal stereotypes. While significant attention ha…

  194. arXiv cs.AI TIER_1 English(EN) · Qi Lyu, Jiahua Dong, Baichen Liu, Xudong Wang, Mingfei Han, Yulun Zhang, Fahad Shahbaz Khan, Salman Khan, Lianqing Liu, Zhi Han ·

    SAB-LVLM:面向大型视觉语言模型的显著性感知二值化

    arXiv:2607.01876v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal understanding, yet their enormous parameter scale and cross-modal computation incur substantial memory and latency overhead, severely limiting re…

  195. arXiv cs.AI TIER_1 English(EN) · Tien-Huy Nguyen, Minh-Nhat Nguyen, Nguyen Nhat Huy, Hung Viet Nguyen, Huy Nguyen Minh Nhat, Thanh-Huy Nguyen, Cuong Tuan Nguyen, Hoang M. Le, Dat Nguyen, Phat Kim Huynh, Min Xu, Ulas Bagci ·

    ESC:面向可靠视觉语言模型的自主情感纠错

    arXiv:2607.02089v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved strong performance across diverse multimodal tasks, yet they remain vulnerable to unreliable reasoning. Existing self-correction methods mitigate these issues but typically rely on post-…

  196. arXiv cs.AI TIER_1 English(EN) · Rintaro Otsubo, Ryo Fujii, Reina Ishikawa, Taiki Kanaya, Kanta Sawafuji, Hiroki Kajita, Shigeki Sakai, Hideo Saito, Ryo Hachiuma ·

    AnyGroundBench:面向视觉语言模型视频定位的专业领域基准测试

    arXiv:2607.02269v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation protocols are largely confined to zero-shot assessments on general, daily-life benchmarks. This…

  197. arXiv cs.AI TIER_1 English(EN) · Ryo Hachiuma ·

    AnyGroundBench:面向视觉语言模型视频定位的专业领域基准测试

    Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation protocols are largely confined to zero-shot assessments on general, daily-life benchmarks. This creates a critical disconnect from real-world app…

  198. Hugging Face Daily Papers TIER_1 English(EN) ·

    SAB-LVLM:面向大型视觉语言模型的显著性感知二值化

    Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal understanding, yet their enormous parameter scale and cross-modal computation incur substantial memory and latency overhead, severely limiting real-world deployment on resource-constrained device…

  199. Hugging Face Daily Papers TIER_1 English(EN) ·

    用于视觉语言导航中语义探索的路径级事后指导

    On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader state distribution. However, such exploration inevitably leads to trajectories that deviate from expert demonstrations, resulting in a semanti…

  200. arXiv cs.AI TIER_1 English(EN) · Lukas Kuhn, Giuseppe Serra, Randall Balestriero, Florian Buettner ·

    LeVLJEPA:无需负样本的端到端视觉语言预训练

    arXiv:2607.00784v1 Announce Type: cross Abstract: Vision-language pretraining remains dominated by contrastive objectives, whereas vision-only self-supervised learning has largely adopted non-contrastive methods. At the same time, the role of vision-language encoders has shifted:…

  201. arXiv cs.LG TIER_1 English(EN) · Seokhee Jin, Changhwan Sung, Sunung Mun, Hoyoung Kim, Jungseul Ok ·

    AdaBoosting Text Prompts for Vision-Language Models

    arXiv:2607.00684v1 Announce Type: new Abstract: The classification accuracy of pretrained Vision-Language Models (VLMs) relies on the quality of the text prompts. Handcrafted templates and Large Language Model (LLM)-generated descriptions not only make predictions more interpreta…

  202. arXiv cs.LG TIER_1 English(EN) · Kai Hu, Akash Bharadwaj, Weichen Yu, Matt Fredrikson ·

    窃取补丁大小:对抗性地操纵视觉语言模型

    arXiv:2607.00174v1 Announce Type: cross Abstract: We present a black-box model-stealing attack that recovers private vision-tokenizer configurations of deployed vision-language models (VLMs), including the visual patch size and input preprocessing pipeline. The key idea is a task…

  203. arXiv cs.LG TIER_1 English(EN) · Sajjad Ghiasvand, Haniyeh Ehsani Oskouie, Mahnoosh Alizadeh, Ramtin Pedarsani ·

    MMLoP:多模态低秩提示用于高效的视觉语言适应

    arXiv:2602.21397v2 Announce Type: replace-cross Abstract: Prompt learning has become a dominant paradigm for adapting vision-language models (VLMs) such as CLIP to downstream tasks without modifying pretrained weights. While extending prompts to both vision and text encoders acro…

  204. arXiv cs.AI TIER_1 English(EN) · Shaoheng Zhang, Zhichen Li, Jie Mei ·

    DART-VLN:离散视觉语言导航的测试时记忆衰减和反循环正则化

    arXiv:2607.01043v1 Announce Type: cross Abstract: Memory-based discrete vision-language navigation (VLN) agents must act under partial observability, yet even strong frozen backbones remain vulnerable at test time. Two common failure modes are stale historical evidence at memory …

  205. arXiv cs.AI TIER_1 English(EN) · Arpita Nema, Hanwei Zhu, Xi Zhang, Weisi Lin ·

    LongVQUBench:视觉语言模型长期视频质量理解的基准测试

    arXiv:2607.01086v1 Announce Type: cross Abstract: The evaluation of long-term video quality understanding remains an open challenge for large vision-language models (LVLMs). Existing video quality benchmarks predominantly focus on short clips and isolated distortions, overlooking…

  206. Hugging Face Daily Papers TIER_1 English(EN) ·

    AnyGroundBench:面向视觉语言模型视频定位的专业领域基准

    Vision-Language Models struggle with domain adaptation in specialized spatio-temporal video grounding tasks, highlighting limitations in zero-shot generalization and in-context learning capabilities.

  207. arXiv cs.AI TIER_1 English(EN) · Weisi Lin ·

    LongVQUBench:视觉语言模型长期视频质量理解的基准测试

    The evaluation of long-term video quality understanding remains an open challenge for large vision-language models (LVLMs). Existing video quality benchmarks predominantly focus on short clips and isolated distortions, overlooking the temporal continuity, cumulative degradation, …

  208. Hugging Face Daily Papers TIER_1 English(EN) ·

    DART-VLN:离散视觉语言导航的测试时记忆衰退和反循环正则化

    Memory-based discrete vision-language navigation (VLN) agents must act under partial observability, yet even strong frozen backbones remain vulnerable at test time. Two common failure modes are stale historical evidence at memory readout and inefficient local backtracking during …

  209. arXiv cs.AI TIER_1 English(EN) · Jie Mei ·

    DART-VLN:离散视觉-语言导航的测试时记忆衰减与反循环正则化

    Memory-based discrete vision-language navigation (VLN) agents must act under partial observability, yet even strong frozen backbones remain vulnerable at test time. Two common failure modes are stale historical evidence at memory readout and inefficient local backtracking during …

  210. arXiv cs.AI TIER_1 English(EN) · Florian Buettner ·

    LeVLJEPA:无需负例的端到端视觉语言预训练

    Vision-language pretraining remains dominated by contrastive objectives, whereas vision-only self-supervised learning has largely adopted non-contrastive methods. At the same time, the role of vision-language encoders has shifted: they are increasingly deployed not as zero-shot c…

  211. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Xiangxiang Chu ·

    M2Note:通过错误笔记学习持续演进的视觉语言模型

    Vision Language Models (VLMs) have demonstrated remarkable capabilities in multimodal reasoning tasks, yet they still suffer from recurring failures, such as skipping key visual checks, misapplying domain rules, and hallucinating unsupported concepts. Most existing solutions rely…

  212. arXiv cs.LG TIER_1 English(EN) · Jungseul Ok ·

    AdaBoosting Text Prompts for Vision-Language Models

    The classification accuracy of pretrained Vision-Language Models (VLMs) relies on the quality of the text prompts. Handcrafted templates and Large Language Model (LLM)-generated descriptions not only make predictions more interpretable, but also enable reuse of the same prompts a…

  213. arXiv cs.AI TIER_1 English(EN) · Nan Li, Albert Gatt, Massimo Poesio ·

    看见不等于分享:一些视觉语言模型在不对称对话中高估共同点

    arXiv:2606.31719v1 Announce Type: cross Abstract: In collaborative dialogue, shared perception does not guarantee shared interpretation. Mutual understanding must be established through interaction. We investigate whether vision-language models (VLMs) can distinguish what could b…

  214. arXiv cs.AI TIER_1 English(EN) · Ta Duc Huy, Trang Nguyen, Townim Chowdhury, Ankit Yadav, Minh-Son To, Zhibin Liao, Johan W. Verjans, Vu Minh Hieu Phan ·

    视觉语义熵:视觉语言模型能识别视觉歧义吗?

    arXiv:2606.31407v1 Announce Type: cross Abstract: Vision-language models can produce confident answers on visually ambiguous inputs, resulting in biased predictions. Common entropy-based methods, such as Semantic Entropy (SE), rely on output diversity. Yet our analysis shows that…

  215. arXiv cs.CL TIER_1 English(EN) · Kaier Liang, Hengde Dai, Cristian-Ioan Vasile ·

    ViTL:通过视觉语言模型进行基于时间逻辑的零样本自然语言导航

    arXiv:2606.30696v1 Announce Type: cross Abstract: Enabling robots to follow natural language commands to complete zero-shot long-horizon tasks remains challenging. It requires extracting implicit temporal and logical constraints from natural language commands and executing multip…

  216. arXiv cs.LG TIER_1 English(EN) · Cl\'ement Fuchs, Tim Bary, Beno\^it Macq ·

    面向视觉-语言模型的图像分类的局部一致性预测

    arXiv:2606.31577v1 Announce Type: cross Abstract: Conformal predictions have attracted significant attention in the field of uncertainty quantification, mainly because of their strong marginal coverage guarantees. Full conditional guarantee is not an attainable goal, a well known…

  217. arXiv cs.LG TIER_1 English(EN) · Aditi Naiknaware, Salimeh Sekeh ·

    T-QPM:赋能开放世界视觉语言模型的时间域外检测与域泛化

    arXiv:2603.18481v2 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection remains a critical challenge in open-world learning, where models must adapt to evolving data distributions. While recent vision-language models (VLMS) like CLIP enable multimodal OOD de…

  218. arXiv cs.AI TIER_1 English(EN) · Massimo Poesio ·

    看见不等于分享:一些视觉语言模型在不对称对话中高估共同点

    In collaborative dialogue, shared perception does not guarantee shared interpretation. Mutual understanding must be established through interaction. We investigate whether vision-language models (VLMs) can distinguish what could be shared from what has been shared between dialogu…

  219. arXiv cs.CL TIER_1 English(EN) · Vu Minh Hieu Phan ·

    视觉语义熵:视觉语言模型能识别视觉歧义吗?

    Vision-language models can produce confident answers on visually ambiguous inputs, resulting in biased predictions. Common entropy-based methods, such as Semantic Entropy (SE), rely on output diversity. Yet our analysis shows that overconfident visual embeddings suppress output d…

  220. arXiv cs.LG TIER_1 English(EN) · Haitao Wu, Qirui Zhang, Zhouheng Yao, Shangquan Sun, Qihao Zheng, Mianxin Liu, Chi Zhang, Wanli Ouyang, Chunfeng Song, Changqing Zhang, Jiamin Wu ·

    BrainJanus:一个用于理解和生成大脑、视觉和语言的统一模型

    arXiv:2606.30319v1 Announce Type: cross Abstract: Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged as a critical frontier in neuroscience. However, existing approaches predominantly treat brain encoding and decodi…

  221. arXiv cs.CL TIER_1 English(EN) · Hee-Seon Kim, Minbeom Kim, Seokil Ham, Changick Kim ·

    利用视觉编码器漏洞实现大型视觉语言模型的通用对抗性扰动

    arXiv:2412.08108v3 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable performance on multimodal tasks but remain highly vulnerable to small adversarial perturbations in input images. Existing attacks typically target the vision en…

  222. arXiv cs.CL TIER_1 English(EN) · Matteo Farina, Vishaal Udandarao, Thao Nguyen, Selim Kuzucu, Maximilian B\"other, Andreas Hochlehnert, Adhiraj Ghosh, Marianna Nezhurina, Karsten Roth, Joschka Struber, Yuhui Zhang, Sebastian Dziadzio, Elaine Sui, Soumya Jahagirdar, Dhruba Ghosh, Hasan H… ·

    DataComp-VLM:改进的视觉语言模型开放数据集

    arXiv:2606.28551v1 Announce Type: cross Abstract: Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curation strategies. We introduce DataComp for VLMs (DC…

  223. arXiv cs.AI TIER_1 English(EN) · Simon Roschmann, Paul Krzakala, Sonia Mazelet, Quentin Bouniot, Zeynep Akata ·

    SOTAlign:通过最优传输实现单模态视觉和语言模型的半监督对齐

    arXiv:2602.23353v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical model of the world. Recent work exploits this convergence by aligning frozen pretraine…

  224. arXiv cs.AI TIER_1 English(EN) · Chen Yang, Yuhao Wei, Ze Xu, Ziheng Zou, Shuang Liang, Delin Ouyang, Lingfeng Qi, Jie Li, Guofa Li ·

    LWDrive:逐层世界模型引导的视觉语言模型规划用于自动驾驶

    arXiv:2606.29879v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning. However, trajectories directly generated by VLMs often encode only coarse driving…

  225. arXiv cs.AI TIER_1 English(EN) · Rahul Chowdhury, Timothy A Rupprecht, Xuan Shen, Pu Zhao, Yanzhi Wang ·

    ScAle:注意力头缩放作为视觉语言模型空间推理的最小适配器

    arXiv:2606.29579v1 Announce Type: cross Abstract: Spatial reasoning remains a persistent challenge for many vision language models (VLMs), and improving it typically requires fine-tuning with substantial additional parameters. Our preliminary analysis reveals that rescaling activ…

  226. arXiv cs.AI TIER_1 English(EN) · Jongoh Jeong, Sun-Kyung Lee, Kuk-Jin Yoon ·

    面向视觉语言数据集蒸馏的秩感知双曲对齐

    arXiv:2606.29464v1 Announce Type: cross Abstract: Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can efficiently train contrastive vision-language models under strict data and compute budgets. Most…

  227. arXiv cs.AI TIER_1 English(EN) · Xiao Wang, Liye Jin, Dan Xu, Yuehang Li, Lan Chen, Yaowei Wang, Yonghong Tian, Jin Tang ·

    使用VLMs动态解析和更新自然语言规范以实现鲁棒的视觉-语言跟踪

    arXiv:2606.29357v1 Announce Type: cross Abstract: Vision-language tracking guided by natural language specifications leverages high-level semantic cues of target objects to substantially boost tracking accuracy and robustness. Existing studies have verified that adaptively optimi…

  228. arXiv cs.AI TIER_1 English(EN) · Minh-Loi Nguyen, Nghiem Tuong Diep, Hung Khang Nguyen, Minh Le, Doanh Le Thien, Hoang H. Tran, Dung D. Le, Vu N. Duong, Daniel Sonntag, An Thai Le, Duy Minh Ho Nguyen, Vien Anh Ngo, Tran Van Nhiem ·

    RoboGaze:通过结构化视觉语言分析评估机器人世界模型

    arXiv:2606.28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning. However, evaluating these videos is challenging: visually realistic outputs often violate physical laws, temporal consis…

  229. arXiv cs.AI TIER_1 English(EN) · Guanglong Sun, Shuang Cui, Bo Lei, Liyuan Wang, Zihan Zhai, Hongwei Yan, Hang Su, Jun Zhu, Yi Zhong ·

    ComMem:用于视觉语言模型测试时自适应的互补记忆系统

    arXiv:2606.28719v1 Announce Type: new Abstract: Test-time adaptation (TTA) of vision-language models (VLMs) is essential for their robust deployment in dynamic, real-world environments. However, existing TTA methods often adapt locally without accumulating knowledge over time, or…

  230. arXiv cs.AI TIER_1 English(EN) · Yichen Guo, Kai Tang, Fenglai Lin, Yiding Sun, Dongshuo Zhang, Wenya Wang, Lin William Cong, Shanghang Zhang ·

    FADE:通过减少大型视觉语言模型中语言先验的主导地位来减轻幻觉

    arXiv:2606.29431v1 Announce Type: new Abstract: Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucination, generating content inconsistent with the input image. Recent studies attribute this to the dominance of language …

  231. arXiv cs.LG TIER_1 English(EN) · Kwun Ho Ngan, Saman Sadeghi Afgeh, Joe Townsend, Artur d'Avila Garcez ·

    对比式视觉语言学习结合释义与否定

    arXiv:2511.16527v2 Announce Type: replace-cross Abstract: Contrastive vision-language models continue to be the dominant approach for image-text retrieval. Contrastive Language-Image Pre-training (CLIP) trains two neural networks to align their image and text embeddings in a shar…

  232. Hugging Face Daily Papers TIER_1 English(EN) ·

    3D HAMSTER:通过3D轨迹引导,在分层视觉语言动作模型中实现规划与控制的桥梁

    3D HAMSTER framework enhances robot manipulation by integrating a vision-language model with depth encoding to generate metrically accurate 3D trajectories for point cloud-based control policies.

  233. Hugging Face Daily Papers TIER_1 English(EN) ·

    眼见不等于分享:某些视觉语言模型在不对称对话中高估共同点

    Vision-language models struggle to distinguish between shared and interpreted visual information in dialogue, relying on static map cues rather than dynamic grounding processes.

  234. arXiv cs.LG TIER_1 English(EN) · Jiamin Wu ·

    BrainJanus:一个用于理解和生成大脑、视觉和语言的统一模型

    Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged as a critical frontier in neuroscience. However, existing approaches predominantly treat brain encoding and decoding as isolated tasks, relying heavily on unimodal …

  235. Hugging Face Daily Papers TIER_1 English(EN) ·

    Dynamo:面向视觉语言智能体的动态技能-工具演化

    Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dynamo, a training-free framework that adapts a frozen VLM without any weight updates. On a small labeled training subset, the agent inspects i…

  236. arXiv cs.AI TIER_1 Deutsch(DE) · Dong-Jae Lee, Sunghyun Baek, Junmo Kim ·

    IWP:大型视觉语言模型中的隐式权重剪枝作为Token剪枝

    arXiv:2604.00757v2 Announce Type: replace-cross Abstract: Large Vision Language Models show impressive performance across image and video understanding tasks, yet their computational cost grows rapidly with the number of visual tokens. Existing token pruning methods mitigate this…

  237. arXiv cs.CL TIER_1 English(EN) · Jaume Guasch-Mart\'i, Enrique Lopez-Cuena, Mart\'in Su\'arez-Fern\'andez, Jordi Bayarri-Planas, Anna Arias-Duart, Dario Garcia-Gasulla ·

    Aloe-Vision:面向医疗保健的鲁棒视觉语言模型

    arXiv:2606.27500v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) specialized in healthcare are emerging as a promising research direction due to their potential impact in clinical and biomedical applications. However, progress is constrained by the scarcity …

  238. arXiv cs.CL TIER_1 English(EN) · Niclas Lietzow, Danielle Bitterman, Carsten Eickhoff, William Rudman, Michal Golovanevsky ·

    Vision-Default, Prior-Override: 视觉-语言模型中感知-知识冲突的因果机制

    arXiv:2606.28273v1 Announce Type: new Abstract: Vision-language models must reconcile visual evidence with memorized world knowledge when the two conflict. How they resolve this conflict shapes the reliability of multimodal systems, yet prior work characterizes it behaviorally wi…

  239. Hugging Face Daily Papers TIER_1 English(EN) ·

    BrainJanus:一个用于理解和生成大脑、视觉和语言的统一模型

    BrainJanus represents the first unified brain model integrating brain, vision, and language through a shared Omni space, enabling bidirectional mapping between neural activity and sensory stimuli via a tokenized representation and autoregressive architecture.

  240. Hugging Face Daily Papers TIER_1 English(EN) ·

    面向视觉语言数据集蒸馏的感知层级双曲对齐

    Vision-language dataset distillation method using rank-aware hyperbolic alignment to optimize synthetic image-text pairs for efficient contrastive model training while preserving modality-specific diversity.

  241. arXiv cs.CL TIER_1 English(EN) · Michal Golovanevsky ·

    Vision-Default, Prior-Override: 视觉-语言模型中感知-知识冲突的因果机制

    Vision-language models must reconcile visual evidence with memorized world knowledge when the two conflict. How they resolve this conflict shapes the reliability of multimodal systems, yet prior work characterizes it behaviorally without a component-level causal account. We combi…

  242. arXiv cs.AI TIER_1 English(EN) · Chenyang Zhang, Anqi Dong, Guangming Zhu, Nuoye Xiong, Siyuan Wang, Lin Mei, Liang Zhang ·

    通过最优传输语义流连接视觉与语言概念

    arXiv:2606.26891v1 Announce Type: cross Abstract: Concept Bottleneck Models (CBMs) promise transparent reasoning by predicting through human-interpretable concepts, yet their effectiveness fundamentally depends on how well visual and textual representations are aligned or matched…

  243. arXiv cs.AI TIER_1 English(EN) · Haoxiang Sun, Tao Wang, Li Yuan, Jian Zhao, Jiancheng Lv ·

    从结构到协同:多模态大语言模型中视觉-语言感知范式演进的调查研究

    arXiv:2606.26196v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, especially following the introduction of models such as OpenAI's O-series and DeepSeek's R-se…

  244. arXiv cs.AI TIER_1 English(EN) · Kai Tang, Jinhao You, Yichen Guo, Yiding Sun, Dongxu Zhang, Wenya Wang, Hanze Li, Tao Luo, Renyuan Li, Xiande Huang ·

    通过层间一致性聚合缓解大型视觉语言模型中的幻觉

    arXiv:2505.12343v2 Announce Type: replace-cross Abstract: Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucinations, where generated content is inconsistent with the input image. Existing training-free hallucination mit…

  245. arXiv cs.CL TIER_1 English(EN) · Byung-Kwan Lee, Ryo Hachiuma, Yong Man Ro, Yu-Chiang Frank Wang, Yueh-Hua Wu ·

    GenRecal:从大型到小型视觉语言模型的校准后生成

    arXiv:2506.15681v4 Announce Type: replace Abstract: Recent advancements in vision-language models (VLMs) have leveraged large language models (LLMs) to achieve performance on par with closed-source systems like GPT-4V. However, deploying these models in real-world scenarios, part…

  246. Hugging Face Daily Papers TIER_1 English(EN) ·

    DataComp-VLM:改进的视觉语言模型开放数据集

    DataComp for VLMs (DCVLM) establishes a comprehensive benchmark for evaluating data curation strategies in vision-language models, demonstrating that data mixing rather than filtering significantly improves model performance at scale.

  247. arXiv cs.CL TIER_1 English(EN) · Dario Garcia-Gasulla ·

    Aloe-Vision:面向医疗保健的鲁棒视觉语言模型

    Large Vision-Language Models (LVLMs) specialized in healthcare are emerging as a promising research direction due to their potential impact in clinical and biomedical applications. However, progress is constrained by the scarcity of high-quality medical multimodal data, concerns …

  248. arXiv cs.AI TIER_1 English(EN) · Hui Li ·

    使用联合稀疏自编码器引导视觉语言模型

    Sparse Autoencoders (SAEs) have shown promise for analyzing language models, but applying them to vision-language models (VLMs) often yields representations that are difficult to use as controllable cross-modal steering directions. We introduce the Joint Sparse Autoencoder (JSAE)…

  249. arXiv cs.AI TIER_1 English(EN) · Ahmad Algadhi, Ahmed Alzuhair, Omar Alkhulaif, Muzammil Behzad ·

    教授:面向视觉语言模型的教师多重无监督提示蒸馏

    arXiv:2606.23897v1 Announce Type: cross Abstract: Prompt distillation compresses large vision-language models (VLMs) such as CLIP into lightweight student models by matching teacher predictions on unlabeled domain images. PromptKD (CVPR 2024) established this paradigm with a sing…

  250. Hugging Face Daily Papers TIER_1 English(EN) ·

    4DVLT:以世界线为中心的视觉语言跟踪实现动态场景理解

    4D dynamic scene understanding requires grounding language to a persistent worldline that binds identity, metric 3D motion, and synchronized multi-view 2D projections. Existing paradigms capture only part of this structure: large multimodal models reason over rich visual evidence…

  251. arXiv cs.CL TIER_1 English(EN) · Yusuf Salcan (Computer Vision Group, University of Freiburg, Germany, CRIION-AI Lab, Freiburg, Germany), Simon Ging (Computer Vision Group, University of Freiburg, Germany, Adaptive & Agentic AI), Robin Schirrmeister (Department of Radiology, Medical Cen… ·

    面向放射学的可扩展空间定位2D视觉语言模型训练

    arXiv:2606.20477v1 Announce Type: cross Abstract: We study how to train visually grounded vision-language models (VLMs) for radiology without manual spatial annotations. We introduce RefRad2D, a large-scale bilingual (German/English) dataset of 1.2M CT and MR image-text pairs der…

  252. arXiv cs.AI TIER_1 English(EN) · Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh ·

    用于胸部放射学的视觉语言模型并非总是需要图像

    arXiv:2606.17710v1 Announce Type: cross Abstract: Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image. That inference is unsafe: a model exploiting finding-name priors scores like one that reads…

  253. arXiv cs.CL TIER_1 English(EN) · Soroosh Tayebi Arasteh ·

    用于胸部放射学的视觉语言模型并非总是需要图像

    Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image. That inference is unsafe: a model exploiting finding-name priors scores like one that reads the scan, and no standard benchmark separates the…

  254. arXiv cs.CV TIER_1 English(EN) · Parsa Esmaeilkhani, Longin Jan Latecki ·

    用于视觉语言模型中补丁级别解释的 Logit Lens 监督

    arXiv:2602.01530v2 Announce Type: replace Abstract: Modern autoregressive Vision-Language Models (VLMs) can generate fluent answers while their visual-token representations become weakly tied to the image regions from which they originate. This limits patch-level explainability: …

  255. arXiv cs.CV TIER_1 English(EN) · Hunter Schofield, Mohammed Elmahgiubi, Mohammad Mahdavian, Richard Shi, Jinjun Shan, Amir Rasouli, Dongfeng Bai ·

    Chain of Spatial Thoughts:面向视觉语言模型的模态无关空间定位

    arXiv:2608.10278v1 Announce Type: new Abstract: Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomous driving. Although recent vision-language models (VLMs) have achieved impressi…

  256. arXiv cs.CV TIER_1 English(EN) · Naren Kumar S, Tirth Bhatt, Mayank Singh ·

    何处可寻?:视觉语言模型中视觉编码器的因果追踪

    arXiv:2608.10758v1 Announce Type: new Abstract: Vision-language models can describe an image with remarkable accuracy, yet a more fundamental question remains unanswered: what visual information actually drives their answers? In this work, we investigate this question through cau…

  257. arXiv cs.CV TIER_1 English(EN) · Kiet T. Nguyen, Hanbo Shim, Jinwoo Kim, Seunghoon Hong ·

    面向视觉语言模型的空间推理的多视图关系蒸馏

    arXiv:2608.10864v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile, leading to failures in spatial reasoning needed for embodied AI, robotics, and …

  258. arXiv cs.CV TIER_1 English(EN) · Yufei Zhang, Chenlu Zhan, Hongwei Wang ·

    当视觉信号误导时:视觉语言模型属性幻觉的机制研究

    arXiv:2608.11024v1 Announce Type: new Abstract: Attribute hallucination---where vision-language models (VLMs) correctly identify an object but mischaracterize its properties---is prevalent yet mechanistically poorly understood. The dominant explanation, language-prior dominance, …

  259. arXiv cs.CV TIER_1 English(EN) · Zhijie Wu, Kento Kawaharazuka, Kei Okada ·

    面向视觉-语言-动作模型的自适应 KV 缓存重用的神经内省门控

    arXiv:2608.10824v1 Announce Type: cross Abstract: Vision-Language-Action(VLA) models map camera images and language instructions directly to motor commands through a single autoregressive transformer. In real-time control, they still spend substantial compute recomputing key-valu…

  260. arXiv cs.CV TIER_1 English(EN) · Issa Sugiura, Keito Sasagawa, Keisuke Nakao, Koki Maeda, Ziqi Yin, Zhishen Yang, Shuhei Kurita, Yusuke Oda, Ryoko Tokuhisa, Daisuke Kawahara, Naoaki Okazaki ·

    Jagle:为视觉语言模型构建大规模日语多模态后训练数据集

    arXiv:2604.02048v2 Announce Type: replace Abstract: Developing vision-language models (VLMs) that generalize across diverse tasks requires large-scale training datasets with diverse content. In English, such datasets are typically constructed by aggregating and curating numerous …

  261. arXiv cs.CV TIER_1 English(EN) · Ningxin Pan, Hanyu Li, Yehui Tang ·

    ComplexityWorld:对可验证视觉决策能力进行基准测试的视觉语言模型

    arXiv:2608.07584v1 Announce Type: new Abstract: Vision-language models (VLMs) have made rapid progress in visual perception and increasingly support real-world tasks that depend on images. Many such tasks, however, require more than rec- ognizing what an image contains: a model m…

  262. arXiv cs.CV TIER_1 English(EN) · Sajjad Ghiasvand, Yifan Yang, Mahnoosh Alizadeh, Ramtin Pedarsani ·

    ZOMP:面向视觉语言模型的零阶多模态提示调优

    arXiv:2608.08060v1 Announce Type: new Abstract: Fine-tuning vision-language models such as CLIP typically requires backpropagation (BP) through the full model, which is infeasible when only forward-pass access is available, as is common for memory-constrained edge devices and pro…

  263. arXiv cs.CV TIER_1 English(EN) · Xuan Yao, Yuze Zhu, Junyu Gao, Zongmeng Wang, Changsheng Xu ·

    SC$^{2}$-WM:一种具有闭环反馈的自纠正世界模型,用于连续环境中的视觉与语言导航

    arXiv:2608.07548v1 Announce Type: cross Abstract: Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to make fine-grained navigation decisions under partial observability. However, most existing methods rely on open-loop execution, lacking mechanis…

  264. arXiv cs.CV TIER_1 English(EN) · Zhewei Zhang, Puyue Wang, Guanren Qiao, Yijie Weng, Jiawei Hu, Guo Li, Lujia Wang, Junyan Wang, Tao Gu, Hongliang Lu, Guiliang Liu, Hong Jia, Xinhu Zheng ·

    LIRA:用于视觉-语言-动作解码的本地跨层信息路由

    arXiv:2608.07596v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models transform representations from pretrained vision-language models (VLMs) into robot actions, yet the interface that routes intermediate VLM features into action decoders remains underexplored. Ex…

  265. arXiv cs.CV TIER_1 English(EN) · Joseph Bingham ·

    用于在视觉感知数据中进行人类指代表达基础的动态语义框架

    arXiv:2608.08663v1 Announce Type: cross Abstract: Humans converge on shared names for novel, hard-to-describe objects through repeated interaction, a process psycholinguists call lexical entrainment. Leading vision-language models fail at this: recent empirical work documents tha…

  266. arXiv cs.CV TIER_1 English(EN) · Hongjin Ji, Guoyang Xia, Luoyang Sun, Fangxiang Feng, Lei Ren ·

    VANE:通过未来视觉表征预测实现可靠的视觉-语言-动作模型测试时训练

    arXiv:2608.09448v1 Announce Type: cross Abstract: Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains difficult to use reliably in closed-loop manipulation. A shared adaptation space …

  267. arXiv cs.CV TIER_1 English(EN) · Kexin Ma, Jing Xiao, Chaofeng Chen, Geyong Min, Guibo Zhu, Jinqiao Wang, Liang Liao ·

    大型视觉语言模型中用于任务感知令牌修剪的解耦相似性

    arXiv:2604.11240v2 Announce Type: replace Abstract: Token pruning has emerged as an effective approach to reduce the substantial computational overhead of Large Vision-Language Models (LVLMs) by discarding less informative visual tokens while preserving performance. However, exis…

  268. arXiv cs.CV TIER_1 English(EN) · Guiyu Zhao, Longteng Guo, Yanghong Mei, Zilin Zhu, Yu Zhang, Bin Cao, Mingming Yu, Xingjian He, Jie Jiang, Jing Liu ·

    AtlasVLA:面向视觉-语言-动作模型的持续世界-自我状态建模

    arXiv:2608.06729v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camer…

  269. arXiv cs.CV TIER_1 English(EN) · Ziheng Liu, Quantao Yang ·

    TEMPO: 视觉-语言-动作模型语义-动作解耦强化学习后训练

    arXiv:2608.07314v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are commonly adapted to downstream manipulation tasks via supervised fine-tuning (SFT) or online reinforcement learning (RL) post-training. SFT is prone to distribution mismatch, and existing RL…

  270. arXiv cs.CV TIER_1 English(EN) · Harisankar Babu, Benjamin Coors, Christopher Lang, Hendrik Berkemeyer, Tamim Asfour, Simon Foell ·

    驾驶视觉-语言-动作模型中规划令牌的深度探测与剪枝

    arXiv:2608.07361v1 Announce Type: cross Abstract: Vision-language-action (VLA) models route driving decisions through a deep language model, but it is unclear how much of that depth the action itself requires. We study a representative driving VLA whose entire plan is carried by …

  271. arXiv cs.CV TIER_1 English(EN) · Thong Nguyen, Vinh-Hien Do, Quynh Vo, Cong-Duy Nguyen, See-Kiong Ng ·

    DynaPix:视觉语言模型能识别确切的未来吗?

    arXiv:2608.05505v1 Announce Type: new Abstract: Acting in a physical scene requires knowing its real later state, not a plausible one. Current evaluations often accept words or a realistic-looking image, so the predicted state is never checked against the true one. We introduce D…

  272. arXiv cs.CV TIER_1 English(EN) · Jihoon Oh, Kento Kawaharazuka, Kei Okada ·

    VLAff:用于统一可操作性的视觉-语言-可操作性模型

    arXiv:2608.05215v1 Announce Type: cross Abstract: Learning manipulation skills from human videos is promising for scalable robot learning. However, the embodiment mismatch between humans and robots makes this challenging. One promising solution is to learn object-centric actionab…

  273. arXiv cs.CV TIER_1 English(EN) · Mahyar Ghazanfari, Amin Tabrizian, Arsyi Aziz, Binshuai Wang, Peng Wei ·

    一个段落胜过千言万语:重新思考用于视觉语言检索的文本监督

    arXiv:2608.05260v1 Announce Type: new Abstract: Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from detailed textual descriptions. While methods such as Long-CLIP extend the token l…

  274. arXiv cs.CV TIER_1 English(EN) · Jingyan Jiang, Yaru Sun, Xiao Chen, Jiazhen Huang, Caiting Li, Zhijian He, Yin Chen, Pingting Hao ·

    尊重你的零样本不确定性:测试时间自适应视觉语言模型的保守校准

    arXiv:2608.05945v1 Announce Type: new Abstract: Test-time adaptation (TTA) can improve the recognition accuracy of vision-language models under distribution shift, but often degrades calibration, making predictive confidence unreliable for downstream decision-making. Many existin…

  275. arXiv cs.CV TIER_1 English(EN) · Xi Xiao, Xingjian Li, Cheng Han, Tianyang Wang, Lin Zhao, Yunbei Zhang, Guosheng Hu, Runmin Jiang, Xi Li, Xiao Wang, Min Xu ·

    使用级联语义适配视觉基础模型

    arXiv:2608.05393v1 Announce Type: new Abstract: Prompt tuning, a leading parameter-efficient adaptation paradigm in NLP, has recently been extended to computer vision. Visual prompt tuning (VPT) adapts pre-trained vision transformers (ViTs) by updating a small set of additional p…

  276. arXiv cs.CV TIER_1 English(EN) · Behnam Raoufi, Hossein Sharify, Mohamad Mahdee Ramezanee, Khosrow Hajsadeghi, Saeed Bagheri Shouraki ·

    CLIP-Joint-Detect:端到端联合训练目标检测器与对比式视觉-语言监督

    arXiv:2512.22969v2 Announce Type: replace Abstract: Conventional object detectors rely on cross-entropy classification, which can be vulnerable to class imbalance and label noise. We propose CLIP-Joint-Detect, a simple and detector-agnostic framework that integrates CLIP-style co…

  277. arXiv cs.CV TIER_1 English(EN) · Weihan Cai, Hao Tan, Zichang Tan, Jun Wan, Xinping Gao ·

    释放视觉语言模型在可泛化 AI 生成图像检测中的潜力

    arXiv:2608.04935v1 Announce Type: new Abstract: Recent work has shown that a simple linear probe on frozen representations from modern vision foundation models (VFMs) can achieve state-of-the-art AIGI detection performance, substantially outperforming specialized detectors in cha…

  278. arXiv cs.CV TIER_1 English(EN) · Yan Zhang, Yinan Wu, Haoran Duan, Jungong Han ·

    CofactVLA:通过反事实干预消除视觉-语言-动作模型中的混淆

    arXiv:2608.04396v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have driven significant progress in robotic manipulation, yet they fundamentally struggle with the vision-override phenomenon. Driven by the severe modality imbalance between dense visual streams …

  279. arXiv cs.CV TIER_1 English(EN) · Gengyuan Liu, Nanzhou Wang, Chang Liu, Qinwen Wu, Zhenhao Wang, Jiacong Wang, Bokui Chen, Xiangyang Ji ·

    MuRA:用于高效有效测试时视觉语言泛化的多秩自适应

    arXiv:2608.03885v1 Announce Type: new Abstract: Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-time adaptation (TTA) via Low-Rank Adaptation offers a parameter-efficient solutio…

  280. arXiv cs.CV TIER_1 English(EN) · Hao Dou, Ruiwen Tian ·

    视觉Token数量减少何时能加速多模态推理?一项跨决策位置和硬件的收支平衡研究

    arXiv:2608.03649v1 Announce Type: new Abstract: Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead, shared work, and the operators each policy can avoid. A stage-level decompositio…

  281. arXiv cs.CV TIER_1 English(EN) · Amir Sabbaghziarani, Mohammadsajad Abavisani, Sergey Plis ·

    自信但不可靠:对视觉语言模型在脑部MRI上的行为安全审计

    arXiv:2608.02790v1 Announce Type: new Abstract: Vision-language models (VLMs), including medical specialists, are increasingly proposed for medical imaging, yet their stated confidence is rarely evaluated separately from correctness. We use brain MRI as a controlled, high-stakes …

  282. arXiv cs.CV TIER_1 English(EN) · Lucy Lin, Ayush Jain, Yifan Liu, Katerina Fragkiadaki ·

    Qwen-3D:一个用于空间理解的通用3D视觉语言模型

    arXiv:2608.02980v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to frame-centric tokenization and limited context windows. 3D geometry provides a na…

  283. arXiv cs.CV TIER_1 English(EN) · Zihan Wang, Tong Liu, Zhiwei Wang, Tao Huang, Wentao Jiang, Sihan Ma, Shanshan Ye, Xiaohui Yang, Jing Zhang ·

    LocAnyMed:用于多模态医学图像的视觉-语言基础

    arXiv:2608.03322v1 Announce Type: new Abstract: Medical visual grounding connects free-form clinical queries to spatial evidence in medical images and is an important component of interpretable medical artificial intelligence. However, general-purpose grounding models are predomi…

  284. arXiv cs.CV TIER_1 English(EN) · Xiangyun Huang, Xiangchen Wang, Runfeng Lin, Yihao Xu, Kangyu Huang, Jiang Hengchen, Xiwang Dong, Lin Jiarong ·

    从路线到步骤:在视觉与语言导航中分离语义进展与局部执行

    arXiv:2608.03143v1 Announce Type: new Abstract: Vision-and-Language Navigation (VLN) requires an agent to follow a route-level instruction by executing its constituent steps from egocentric visual observations. Existing VLM-based navigators typically supervise both capabilities t…

  285. arXiv cs.CV TIER_1 English(EN) · Yaozhi Wen, Jialong Guo, Zhenliang Ni, Han Shu, Xinghao Chen ·

    SlimVLM:用于高效视觉语言模型的敏感感知动态结构化剪枝与自适应视觉标记选择

    arXiv:2608.03580v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have demonstrated remarkable performance in processing and understanding both text and images, their large parameter sizes lead to significant computational overhead, limiting their deployment on …

  286. arXiv cs.CV TIER_1 English(EN) · Landi He, Mingde Yao, Shawn Young, Lijian Xu ·

    DiffPrune:用于视觉语言模型中 token 修剪的可微分信息节流

    arXiv:2608.01985v1 Announce Type: new Abstract: Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. The key is to learn a score that measures whether a token is useful. Existing methods typically rely on Gumbel…

  287. arXiv cs.CV TIER_1 English(EN) · Yanbin Hu, Jin Cui, Jun Ye, Jiepeng Zhou, Jiangcheng Song, Boran Zhao, Pengju Ren ·

    提炼RGB可恢复内容:仅RGB视觉语言模型的特权3D证据

    arXiv:2608.00110v1 Announce Type: new Abstract: 3D scene understanding requires reasoning about entity existence, spatial layout, and object relations, yet RGB images alone often provide insufficient 3D cues. Existing 3D-VLMs commonly rely on depth or 3D-position-aware inputs at …

  288. arXiv cs.CV TIER_1 English(EN) · Meibo Hu, Guohao Sun, Annemarie D. Ross, Sheng Li, Zhiqiang Tao ·

    用于手语翻译的注意力引导视觉语言模型

    arXiv:2608.00235v1 Announce Type: new Abstract: Vision-language models (VLMs) have emerged as a powerful framework for multimodal video understanding. However, they remain limited in the sign language translation task, where we identify a key failure mode of existing VLMbased tra…

  289. arXiv cs.CV TIER_1 English(EN) · Guoliang You, Haifan Gong, Xiaomeng Chu ·

    CT视觉语言学习的语义校准证据组合

    arXiv:2608.00239v1 Announce Type: new Abstract: Learning transferable representations from CT-report pairs requires combining whole-volume context with anatomy-specific evidence. Existing methods typically emphasize either global CT-report alignment or fine-grained anatomy-level …

  290. arXiv cs.CV TIER_1 English(EN) · Ziang Wu, Peng Jin, Qishen Yin, Munan Ning, Hao Li, Peizhen Zhang, Li Yuan ·

    内松外衡:面向视觉语言混合专家模型的几何引导负载均衡

    arXiv:2608.00574v1 Announce Type: new Abstract: Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this token mix. We call the standard token-level Switch auxiliary loss Std-Aux. Std-…

  291. arXiv cs.CV TIER_1 English(EN) · Hashmat Shadab Malik, Toluwani Aremu, Samuele Poppi, Muzammal Naseer, Salman Khan ·

    ReACT-CLIP:视觉-语言模型的响应感知测试时防御

    arXiv:2608.01067v1 Announce Type: new Abstract: Training-free test-time defenses offer a practical way to improve the adversarial robustness of CLIP-style vision--language models without modifying the pretrained model. However, their correction strength is typically fixed for a n…

  292. arXiv cs.CV TIER_1 English(EN) · Puzhuo Zheng, Hasan Kurban ·

    是解码格式,而非扰动:审计基于一致性的选择以实现视觉-语言测试时缩放

    arXiv:2608.01207v1 Announce Type: new Abstract: Test-time scaling lifts large language model reasoning by sampling many candidate solutions and selecting among them, yet the same recipe transfers poorly to vision-language models (VLMs): recent work shows that simple majority voti…

  293. arXiv cs.CV TIER_1 English(EN) · Omid Nejati Manzari, Guillaume Lajoie, Hassan Rivaz ·

    用于通用符号推理的递归视觉语言模型

    arXiv:2608.01534v1 Announce Type: new Abstract: Hard symbolic-reasoning tasks such as Sudoku, maze pathfinding, and ARC remain challenging for LLMs due to their fixed-depth autoregressive reasoning, which limits systematic search, refinement, and backtracking. While recursive mod…

  294. arXiv cs.CV TIER_1 English(EN) · Ashfak Yeafi, Mehedi Hasan, Md Khairul Islam ·

    线性多时间尺度记忆作为一种内存高效的视觉语言桥梁

    arXiv:2608.01614v1 Announce Type: new Abstract: Vision-Language Models (VLMs) face a critical computational bottleneck when processing high-resolution imagery due to the $O(N^2)$ memory complexity of Softmax Multi-Head Attention (MHA). While substituting MHA with independent Mult…

  295. arXiv cs.CV TIER_1 English(EN) · Yu Chen, Xiaohong Li, Xiaole Wang, Jianjin Zhang, Jun Sun, Yafeng Deng ·

    CRAFT:视频令牌的递归自适应融合压缩用于视觉语言模型

    arXiv:2608.01644v1 Announce Type: new Abstract: In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the prefill stage to rise sharply. Such visual sequences are highly redundant along the…

  296. arXiv cs.CV TIER_1 English(EN) · Hai Nguyen, Tung Vu, Cong Tran ·

    SpatialQuery:对视觉语言模型中基于几何的多实例空间推理进行基准测试

    arXiv:2608.01709v1 Announce Type: new Abstract: Vision-language models (VLMs) achieve strong semantic understanding but remain unreliable in metric spatial reasoning, particularly when queries require comparing multiple instances of the same object category. We study this problem…

  297. arXiv cs.CV TIER_1 English(EN) · Hongjie Zhou, Shiqin Wang, Haoyang Chen, Haonan Guo, Di Wang, Juhua Liu, Fu Lin, Yong Luo ·

    RSVideo:您的视觉语言模型准备好处理遥感视频了吗?

    arXiv:2608.02039v1 Announce Type: new Abstract: Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, actions, interactions, and scene changes that cannot be captured by isolated images.…

  298. arXiv cs.CV TIER_1 English(EN) · Xuanhui Lin, Junhao Dong, Mingrong Gong, Yucheng Chen, Xinghua Qu, Yew-Soon Ong ·

    一枚硬币的两面:视觉语言模型跨任务攻击的协同演化搜索

    arXiv:2608.02137v1 Announce Type: new Abstract: Vision-language models (VLMs) exhibit strong generalization across multimodal tasks but remain vulnerable to adversarial perturbations. Existing attacks typically follow single-trajectory gradient optimization or task-specific objec…

  299. arXiv cs.CV TIER_1 English(EN) · Yan Huang, Guowei Wang, Xu Wang, Kangjun Liu, Xin Lin ·

    面向视觉语言模型的测试时自适应的局部边缘恢复

    arXiv:2608.02216v1 Announce Type: new Abstract: Vision-language models (VLMs) such as CLIP exhibit remarkable zero-shot capabilities, yet their performance frequently degrades sharply under unexpected test-time distribution shifts. While Test-Time Adaptation (TTA) offers a promis…

  300. arXiv cs.CV TIER_1 English(EN) · Brian Song, Michael A. Lepori, Ellie Pavlick ·

    语言上下文重塑视觉语言模型中的视觉表征

    arXiv:2608.00035v1 Announce Type: cross Abstract: Goal-directed visual processing is a hallmark of human visual intelligence, resulting in representations that support downstream tasks such as categorization or search. Though vision-language models (VLMs) are often faced with the…

  301. arXiv cs.CV TIER_1 English(EN) · Qi Lv, Jianming Xing, Zhao Yang, Mingyuan Yao, Yinan Shi, Yawei Jueluo, Mike Zheng Shou, Xiang Deng ·

    Hermite曲线作为视觉-语言-动作模型的轨迹先验

    arXiv:2608.01265v1 Announce Type: cross Abstract: Despite recent progress in Vision-Language-Action (VLA) models for robotic manipulation, the action chunk remains a weakly structured interface. Existing work typically flatten each chunk into per-timestep controls, relying on imp…

  302. arXiv cs.CV TIER_1 English(EN) · Li Lin, Wujun Xu, Weiwei Meng, Kaiwen Xia, Kang Hao Cheong, Shuai Wang ·

    FibVLA:一种采用斐波那契采样的高效时序视觉-语言-动作模型

    arXiv:2607.29596v1 Announce Type: cross Abstract: Vision-language-action models (VLAs), which leverage the cognition of multimodal information to infer physical-world actions, provide a generalized solution for embodied AI applications. Conventional VLAs usually concentrate on cu…

  303. arXiv cs.CV TIER_1 English(EN) · Jiasheng Li, Zhong Ji, Yan Zhang, Huihui Li ·

    校准优先于推理:强大的视觉令牌缩减以对抗视觉语言模型中的语义漂移

    arXiv:2607.27700v1 Announce Type: new Abstract: Large Vision-Language Models (VLMs) suffer from prohibitive inference overhead due to long sequences of visual tokens. However, existing visual token reduction methods mainly improve efficiency by pruning or compressing redundant to…

  304. arXiv cs.CV TIER_1 English(EN) · Nguyen Duc Thai, Junhao Dong, Sua Qi Rong, Hua Yu, Yew-Soon Ong ·

    统一视觉语言模型中的对抗性鲁棒模型专家

    arXiv:2607.27897v1 Announce Type: new Abstract: Vision-language models (VLMs), such as CLIP, are vulnerable to adversarial attacks, posing a serious problem for real-life applications and deployment. Adversarial fine-tuning emerges as a prominent defense method; however, differen…

  305. arXiv cs.CV TIER_1 English(EN) · Ioannis Sarridis, Ioannis Kompatsiaris, Symeon Papadopoulos ·

    扩大视觉-语言模型规模不足以缓解偏见

    arXiv:2607.28211v1 Announce Type: new Abstract: Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly understood at scale. We present the first large-scale empirical study of 194 publicl…

  306. arXiv cs.CV TIER_1 English(EN) · Jialuo He, Huangxun Chen ·

    面向高效视觉语言模型的能量驱动自适应视觉令牌剪枝

    arXiv:2603.05950v2 Announce Type: replace Abstract: Visual token reduction is critical for accelerating Vision-Language Models (VLMs), since visual inputs are represented as token sequences that introduce substantial computational overhead in the LLM backbone. However, most pruni…

  307. arXiv cs.CV TIER_1 English(EN) · Enyi Shi, Fei Shen, Chuancheng Shi, Shuyi Miao, Linxia Zhu, Pengyang Shao, Jinhui Tang, Tat-Seng Chua ·

    面向多语言视觉-语言大模型的定向可解释安全神经元增强

    arXiv:2604.08881v2 Announce Type: replace Abstract: With the widespread deployment of vision-language large models (VLLMs), their safety alignment faces dual challenges across languages and modalities. Existing methods model multilingual and multimodal safety separately, overlook…

  308. arXiv cs.CV TIER_1 English(EN) · Darsha Udayanga, Pin-Yu Chen, Payel Das, Qiang Ji ·

    视觉语言模型能否推理图像中的 AI 编辑?

    arXiv:2607.28464v1 Announce Type: new Abstract: Detection and localization of AI-tampered images are critical for trustworthy AI, yet modern generative models have made such manipulations increasingly difficult to identify. While traditional binary classifiers can detect image ta…

  309. arXiv cs.CV TIER_1 English(EN) · Yitao Zhu, Mengjun Liu, Yingji Fu, Haowen Pang, Anqi Qiu ·

    MedARC:3D医学视觉-语言模型的无训练自适应冗余视觉标记压缩

    arXiv:2607.26554v1 Announce Type: new Abstract: Integrating 3D medical images with vision-language models (VLMs) holds substantial promise for computer-aided diagnosis. However, volumetric images generate prohibitively long visual-token sequences with considerable spatial and int…

  310. arXiv cs.CV TIER_1 English(EN) · Jinkun Zhao, Kui Zhang, Wenjun Wu ·

    TPCD:视觉语言模型中的音调-压力对比解码和无标签门控瓶颈

    arXiv:2607.26536v1 Announce Type: new Abstract: High-pressure prompts can push vision-language models (VLMs) into unsupported commitments, such as reading illegible text, reporting indeterminate times, or affirming absent objects. This paper asks whether the pressure-induced dist…

  311. arXiv cs.CV TIER_1 English(EN) · Zhongbin Guo, Jiahe Liu, Wenyu Gao, Yushan Li, Xiaomin He, Chengzhi Li, Ping Jian ·

    LISA-3D:通过多视图一致性将语言-图像分割提升至三维

    arXiv:2512.01008v2 Announce Type: replace Abstract: Text-driven 3D reconstruction requires masks that understand free-form instructions and remain stable under viewpoint changes. We present LISA-3D, a two-stage framework that adapts the instruction-following segmenter LISA with g…

  312. arXiv cs.CV TIER_1 English(EN) · Hengyi Xie, Chenfei Yao, Xianjin Wu, Xuanyang Xi, Yiping Tang, Di Xu, Yingying Zhu, Dingkang Liang, Xiang Bai, Han Ding ·

    TurboVLA:RTX 4090 上实现 32 Hz 实时视觉-语言-动作模型,VRAM 占用 <1 GB

    arXiv:2607.27205v1 Announce Type: new Abstract: Vision-language-action (VLA) models commonly adopt an LLM-centric $V \to L \to A$ pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Alth…

  313. arXiv cs.CV TIER_1 English(EN) · Siyao Li, Jiawei Gu, Shuai Liu, Kairui Hu, Zekun Li, Linjie Li, Chengcheng Tang, Po-Chen Wu, Ivan Shugurov, Lingni Ma, Michael Zollhoefer, Sizhe An, Abhay Mittal, Amy Zhao, Ranjay Krishna, Manling Li, Ziwei Liu, Chuan Guo ·

    HumanCLAW:视觉语言模型能否通过身体行动?

    arXiv:2607.27180v1 Announce Type: new Abstract: Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a ba…

  314. arXiv cs.CV TIER_1 English(EN) · Yunzhan Fu, Enyu Bao, Xiangyu Shen, Yihao Wu, Chunbo Jiang, Fangli Guan, Liqi Yan ·

    SCALPEL:通过 LLM 驱动的编码器学习实现医学视觉语言表示的语义跨模态对齐

    arXiv:2607.26885v1 Announce Type: new Abstract: Vision-language pre-training (VLP) serves as a cornerstone for medical multimodal representation learning. However, existing medical VLP frameworks are often constrained by the limited context windows and shallow representational ca…

  315. arXiv cs.CV TIER_1 English(EN) · Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu, Serge Belongie, V\'esteinn Sn{\ae}bjarnarson ·

    看见还是知道?多模态大语言模型中的视觉上下文敏感性

    arXiv:2607.26326v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. However, they often fail on vision-centric tasks, especially when visual evidence c…

  316. arXiv cs.CV TIER_1 English(EN) · Mingkuan Feng, Zhengqi Wen, Jianhua Tao ·

    解耦视觉处理:通过特定模态的Transformer替换实现高效多模态适应

    arXiv:2607.26596v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual understanding within a unified transformer architecture. However, fine-tuning all parameters of these models for vi…

  317. arXiv cs.CV TIER_1 English(EN) · Zichao Lin, Yifeng Xie, Bowen Qu, Haiming Wang, Jia Li, Haoning Wu, Yuhao Dong, Zuhao Yang, Jinguo Zhu, Haoyu Lu, Zijia Zhao, Tongtian Yue, Zhangyang Qi, Junwei Yang, Mengfan Dong, Peizhou Cao, Chenzhuang Du, Zaida Zhou, Haotian Yao, Hao Yang, Hongcheng … ·

    PerceptionBench:评估多模态大语言模型中的原子视觉感知

    arXiv:2607.24957v1 Announce Type: new Abstract: We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evalua…

  318. arXiv cs.CV TIER_1 English(EN) · Riccardo Andrea Izzo, Gianluca Bardaro, Matteo Matteucci ·

    行动、思考或弃权:面向视觉-语言-动作模型的面向复杂性的自适应推理

    arXiv:2603.05147v2 Announce Type: replace Abstract: Current research on Vision-Language-Action (VLA) models predominantly focuses on enhancing generalization through reasoning techniques. While effective, these improvements increase computational complexity and inference latency.…

  319. arXiv cs.CV TIER_1 English(EN) · Daojie Peng, Fulong Ma, Jun Ma ·

    用于高效和可泛化视觉-语言导航的结构化观察语言

    arXiv:2603.27577v2 Announce Type: replace Abstract: Vision-Language Navigation (VLN) requires an embodied agent to navigate complex environments by following natural language instructions, which typically demands tight fusion of visual and language modalities. Existing VLN method…

  320. arXiv cs.CV TIER_1 English(EN) · Shaofei Lei ·

    MAViE:一种用于细粒度视觉感知和高效多模态推理的多尺度自适应视觉编码器

    arXiv:2607.24424v1 Announce Type: new Abstract: Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features can discard text, local attributes, and spatial relationships, while high-resoluti…

  321. arXiv cs.CV TIER_1 English(EN) · Yuqi Liu, Shengju Qian, Tianyuan Qu, Mingxian Lin, Zixuan Wang, Xin Wang, Bei Yu, Jiaya Jia ·

    MemVLN:用于视觉与语言导航的片段式和程序式记忆

    arXiv:2607.23504v1 Announce Type: new Abstract: Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to maintain long-horizon visual history for trajectory consistency while executing actions with low latency. Existing video-based VLN approaches typi…

  322. arXiv cs.CV TIER_1 English(EN) · Ioannis Maniadis Metaxas, Adrian Bulat, Alberto Baldrati, Anestis Zaganidis, Yassine Ouali, Hyeonuk Kim, Georgios Tzimiropoulos ·

    UltraViT:面向大型视觉语言模型、针对设备端延迟优化的视觉编码器

    arXiv:2607.23373v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) remain bottlenecked by massive computational footprints, precluding their deployment on resource-constrained edge devices. While efforts to compress LVLMs focus heavily on vision token reduction …

  323. arXiv cs.CV TIER_1 English(EN) · Moshiur Farazi, Sameera Ramasinghe, Bekir Sait Ciftler, Mahbub Ahmed Turza, Shafin Rahman ·

    大门常闭:关于向冻结的视觉语言模型注入辅助信号

    arXiv:2607.23335v1 Announce Type: new Abstract: Auxiliary signal pathways in VLMs are routinely fitted with learnable gates so the optimiser can decide how much of the signal to admit. We find that the optimiser almost always decides on zero: across five injection designs, every …

  324. arXiv cs.CV TIER_1 English(EN) · Yihao Wu, Chenyi Xu, Liqi Yan, Chenhuan Cai, Geyong Min, Bin Lin, Fangli Guan, Jianhui Zhang, Pan Li ·

    迈向双脑最小充分表征用于视觉语言导航

    arXiv:2607.23181v1 Announce Type: new Abstract: Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multimodal large models and world-model-based methods have…

  325. arXiv cs.CV TIER_1 English(EN) · Futa Waseda, Saku Sugawara, Isao Echizen ·

    优质文本、强大视觉:语言在增强视觉语言模型视觉鲁棒性中的作用

    arXiv:2507.16257v2 Announce Type: replace Abstract: Defending pre-trained vision-language models (VLMs), such as CLIP, against adversarial attacks is crucial, as these models are widely used in diverse zero-shot tasks, including image classification. However, existing adversarial…

  326. arXiv cs.CV TIER_1 English(EN) · Pau de Jorge, C\'esar Roberto de Souza, Bj\"orn Michele, Mert B\"ulent Sar{\i}y{\i}ld{\i}z, Philippe Weinzaepfel, Florent Perronnin, Diane Larlus, Yannis Kalantidis ·

    任务对齐:跨多样化视觉任务实用模型合并的简单代理

    arXiv:2604.12935v2 Announce Type: replace Abstract: Efficiently merging several models fine-tuned for different tasks, but stemming from the same pretrained base model, is of great practical interest. Despite extensive prior work, most evaluations of model merging in computer vis…

  327. arXiv cs.CV TIER_1 English(EN) · Futa Waseda, Shojiro Yamabe, Daiki Shiono, Kento Sasaki, Tsubasa Takahashi ·

    读还是不读?视觉语言模型中字体攻击鲁棒性和文本识别的统一基准

    arXiv:2512.11899v2 Announce Type: replace Abstract: Large vision-language models (LVLMs) are vulnerable to typographic attacks, where misleading text inserted into an image can override visual understanding. However, existing evaluation protocols and defenses are largely focused …

  328. arXiv cs.CV TIER_1 English(EN) · Wenqi Marshall Guo, Qingyun Qian, Shiyu Zhou, Guoping Luo, Shan Du ·

    MissingBench-Verified:探究视觉语言模型检测缺失物体部件的能力

    arXiv:2607.18673v1 Announce Type: new Abstract: Vision Language Models (VLMs) are well known for hallucinating non-existent objects in images. Objects with missing parts present a unique challenge for VLMs, stemming from both real-world knowledge bias and the scarcity of such ima…

  329. arXiv cs.CV TIER_1 English(EN) · Xiping Li, Jianghong Ma ·

    AIM-CoT:面向视觉语言推理的主动信息驱动多模态思维链

    arXiv:2509.25699v4 Announce Type: replace Abstract: Interleaved-Modal Chain-of-Thought (I-MCoT) advances vision-language reasoning, such as Visual Question Answering (VQA). This paradigm integrates specially selected visual evidence from the input image into the context of Vision…

  330. arXiv cs.CV TIER_1 English(EN) · Wei Chen, Zhiyuan Li, Shuo Xin ·

    OmniVLM:一种令牌压缩、参数量低于十亿的视觉语言模型,用于高效的设备端推理

    arXiv:2412.11475v3 Announce Type: replace Abstract: We present OmniVLM, a sub-billion-parameter vision-language model for efficient on-device inference. OmniVLM introduces a token compression mechanism that reduces visual token sequence length from 729 to 81 tokens, significantly…

  331. arXiv cs.CV TIER_1 English(EN) · Tarun Tomar ·

    搜索任务特定视觉路径:跨视觉语言模型的进化块剪枝

    arXiv:2607.17052v1 Announce Type: new Abstract: Vision-language models normally execute the same complete vision encoder for every question, even when OCR, counting, object, attribute, and spatial queries may not require identical computation. We study whether fixed-budget combin…

  332. arXiv cs.CV TIER_1 English(EN) · Zhifang Zhang, Qiqi Tao, Jiaqi Lv, Na Zhao, Lei Feng, Joey Tianyi Zhou ·

    TokenSwap:大型视觉语言模型组合理解中的后门攻击

    arXiv:2509.24566v2 Announce Type: replace Abstract: Large vision-language models (LVLMs) have achieved impressive performance across a wide range of vision-language tasks, while they remain vulnerable to backdoor attacks. Existing backdoor attacks on LVLMs aim to force the victim…

  333. arXiv cs.CV TIER_1 English(EN) · Biao Chen, Yunqian Yu, Xiangxu Zhao, Zhongshu Chen, Mengmeng Jing, Lin Zuo ·

    面向视觉-语言模型的U型多粒度学习

    arXiv:2607.14966v1 Announce Type: new Abstract: The prompt learning paradigm for vision-language models is effective yet faces a granularity dilemma: global prompts lack fine-grained semantic awareness, while local prompts ignore contextual associations, limiting cross-task gener…

  334. arXiv cs.CV TIER_1 English(EN) · Lin Zuo ·

    面向视觉-语言模型的U型多粒度学习

    The prompt learning paradigm for vision-language models is effective yet faces a granularity dilemma: global prompts lack fine-grained semantic awareness, while local prompts ignore contextual associations, limiting cross-task generalization. This dilemma exists in dense predicti…

  335. arXiv cs.CV TIER_1 English(EN) · Yu Fang, Yuchun Feng, Dong Jing, Jiaqi Liu, Yue Yang, Zhenyu Wei, Daniel Szafir, Mingyu Ding ·

    当视觉超越语言:评估和缓解VLA中的反事实失败

    arXiv:2602.17659v2 Announce Type: replace Abstract: Vision-Language-Action models (VLAs) promise to ground language instructions in robot control, yet in practice often fail to faithfully follow language. When presented with instructions that lack strong scene-specific supervisio…

  336. arXiv cs.CV TIER_1 English(EN) · Xuanyi Hao, Zuoyuan Zhang, Zhibo Wang, Xiaoyi Pang, Jiahui Hu, Jiacheng Du, Shuguo Zhuo ·

    面向高效视觉语言模型的高效无注意力轻量级Token缩减

    arXiv:2607.13500v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have achieved strong performance in multimodal understanding, yet remain challenging to deploy on resource-constrained edge devices due to the substantial computational overhead of processing numerous v…

  337. arXiv cs.CV TIER_1 English(EN) · Shuguo Zhuo ·

    面向高效视觉语言模型的无注意力轻量级Token缩减

    Vision-Language Models (VLMs) have achieved strong performance in multimodal understanding, yet remain challenging to deploy on resource-constrained edge devices due to the substantial computational overhead of processing numerous visual tokens. Token reduction is a promising dir…

  338. arXiv cs.CV TIER_1 English(EN) · Mao Chen, Xiangkai Zhang, Zhiyong Liu, Chuankai Liu, Xu Yang ·

    超越所处位置:从跨视图定位学习语义、结构和几何

    arXiv:2607.12429v1 Announce Type: new Abstract: Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same scene across extreme viewpoint gaps. Cross-view localization naturally provides a …

  339. arXiv cs.CV TIER_1 English(EN) · Lin Peng, Cong Wan, Zeyu Guo, SongLin Dong, Yihong Gong ·

    CoRe:视觉语言模型跨图像比较推理的综合框架

    arXiv:2607.12786v1 Announce Type: new Abstract: Cross-image comparative reasoning remains challenging for vision-language models (VLMs), especially when correct prediction requires fine-grained attribute grounding and globally consistent reasoning. We present CoRe, a unified fram…

  340. arXiv cs.CV TIER_1 English(EN) · Xinyue Xu, Zheng Zhang, Kunyang Ma, Ge Zhu, Lianshuai Cao, Lei Wang, Zixuan Li, Yi Cheng ·

    DM-KG:一种提升视觉语言模型在街景图像中空间认知能力的新方法

    arXiv:2607.12319v1 Announce Type: new Abstract: As vision-language models (VLMs) are increasingly deployed in geospatial question answering and visual scene understanding, improving their spatial cognition capability on street view imagery for complex logical reasoning has emerge…

  341. arXiv cs.CV TIER_1 English(EN) · Jiahang Wang, Yirong Yang, Yanqing Zhu, Minghua Luo, Shichao Xie, Fei Liu, Mu Xu ·

    ReflectVLN: 使用反思性推理训练视觉语言导航代理

    arXiv:2607.12680v1 Announce Type: new Abstract: Existing vision-language navigation methods often couple a VLM with waypoint decoders to produce multi-step action plans, but they typically lack an explicit closed-loop mechanism for tracking semantic progress, diagnosing execution…

  342. arXiv cs.CV TIER_1 English(EN) · Sania Waheed, Michael Milford, Sarvapali D. Ramchurn, Shoaib Ehsan ·

    似曾相识的突破:通过视觉语言推理对视觉地点识别进行独立审计

    arXiv:2607.12818v1 Announce Type: new Abstract: Visual place recognition (VPR) is a key enabler of accurate localization and long-term autonomous navigation in robotics applications, such as loop closure detection for simultaneous localisation and mapping (SLAM). However, real-wo…

  343. arXiv cs.CV TIER_1 English(EN) · Jiho Hong, Eunae Kang, Sanghyun Kim, Young-Sik Shin ·

    面向视觉语言导航的实例增强语义地图

    arXiv:2607.12630v1 Announce Type: cross Abstract: Visual Language Navigation (VLN) aims to enable an embodied agent to navigate complex environments by following natural language instructions. Recent approaches build semantic spatial maps and leverage Large Language Models (LLMs)…

  344. arXiv cs.CV TIER_1 English(EN) · Shoaib Ehsan ·

    突发似曾相识:通过视觉语言推理对视觉地点识别进行独立审计

    Visual place recognition (VPR) is a key enabler of accurate localization and long-term autonomous navigation in robotics applications, such as loop closure detection for simultaneous localisation and mapping (SLAM). However, real-world VPR deployment relies on selecting an image …

  345. arXiv cs.CV TIER_1 English(EN) · Yihong Gong ·

    CoRe:视觉语言模型跨图像比较推理的综合框架

    Cross-image comparative reasoning remains challenging for vision-language models (VLMs), especially when correct prediction requires fine-grained attribute grounding and globally consistent reasoning. We present CoRe, a unified framework for this problem. CoRe includes: (i) CoRe-…

  346. arXiv cs.CV TIER_1 English(EN) · Mu Xu ·

    ReflectVLN: 使用反思性推理训练视觉语言导航代理

    Existing vision-language navigation methods often couple a VLM with waypoint decoders to produce multi-step action plans, but they typically lack an explicit closed-loop mechanism for tracking semantic progress, diagnosing execution failures, and recovering from error accumulatio…

  347. arXiv cs.CV TIER_1 English(EN) · Young-Sik Shin ·

    面向视觉语言导航的实例增强语义地图

    Visual Language Navigation (VLN) aims to enable an embodied agent to navigate complex environments by following natural language instructions. Recent approaches build semantic spatial maps and leverage Large Language Models (LLMs) for reasoning and decision making. Despite these …

  348. arXiv cs.CV TIER_1 English(EN) · Xu Yang ·

    超越所处位置:从跨视图定位学习语义、结构和几何

    Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same scene across extreme viewpoint gaps. Cross-view localization naturally provides a promising pathway toward this ability, as it req…

  349. arXiv cs.CV TIER_1 English(EN) · Hao Zheng, Jinyi Huang, Tiantian Zheng, Xun Xu, Tuka Alhanai ·

    用于视频复杂装配动作理解的组合上下文微调视觉语言模型

    arXiv:2607.10797v1 Announce Type: new Abstract: Assembly action understanding is a key enabler for effective human-robot collaborative assembly, yet it remains challenging due to subtle motions and fine-grained hand-object interactions. We adapt vision-language models (VLMs) to t…

  350. arXiv cs.CV TIER_1 English(EN) · Zhaoyang Li, Yanjun Li, Wangkai Li, Yujia Chen, Tianzhu Zhang ·

    用于视觉语言模型保守令牌凝结的光谱热流

    arXiv:2607.10640v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are costly at inference time because they must process long sequences of visual tokens. Existing token pruning methods often degrade under high compression by blindly discarding information, breaking sp…

  351. arXiv cs.CV TIER_1 English(EN) · Quynh Vo, Phuc Dao, Cong-Duy Nguyen, Thong Nguyen ·

    深度比展示更能说明问题:用于视觉语言空间推理的深度序数提示

    arXiv:2607.11173v1 Announce Type: new Abstract: Vision-language models (VLMs) are expected to reason about physical space -- which object is closer, what lies behind what, and how objects are arranged in 3D -- yet they still struggle with such spatial judgments. A natural remedy …

  352. arXiv cs.CV TIER_1 English(EN) · Robert Wijaya, Ngai-Man Cheung ·

    大型视觉语言模型中的认知专家混合体

    arXiv:2607.10796v1 Announce Type: new Abstract: Large Vision Language Models (LVLMs) require strong reasoning over both visual and textual input. Recent work suggests that cognitive elements, especially diverse representations and metacognition, correlate with better performance.…

  353. arXiv cs.CV TIER_1 English(EN) · Yi Cheng ·

    DM-KG:一种提升街景图像中视觉语言模型空间认知能力的新方法

    As vision-language models (VLMs) are increasingly deployed in geospatial question answering and visual scene understanding, improving their spatial cognition capability on street view imagery for complex logical reasoning has emerged as a key research priority. However, existing …

  354. arXiv cs.CV TIER_1 English(EN) · Thong Nguyen ·

    深度比展示更能说明问题:用于视觉语言空间推理的深度序数提示

    Vision-language models (VLMs) are expected to reason about physical space -- which object is closer, what lies behind what, and how objects are arranged in 3D -- yet they still struggle with such spatial judgments. A natural remedy is to show the model a depth map, but we find th…

  355. arXiv cs.CV TIER_1 English(EN) · Yuncheng Yang, Feiyang Ye, Shixian Luo, Yinna Zhu, Lianlei Shan, Wangcai Zhao, Kuo Zhang, Yan Chen, Yong Wu, Yan Xie ·

    MOSAIC:高效异构视觉-语言模型的自适应层间组合

    arXiv:2607.09029v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have achieved success using homogeneous Transformers to process multimedia data. Recent studies show that heterogeneous structures interleaving efficient mechanisms, like linear attention, improve both …

  356. arXiv cs.CV TIER_1 English(EN) · Francis Fernandez, Arash Jahangiri, Salimeh Sekeh ·

    C-GAP:类感知和在线提示改进了具有类别不平衡的视觉语言模型

    arXiv:2607.09008v1 Announce Type: new Abstract: Safety-critical perception systems must reliably detect rare object classes within small label spaces, a setting that long-tailed detection methods, designed for hundreds of classes with dense annotation, fundamentally do not addres…

  357. arXiv cs.CV TIER_1 English(EN) · Meng Wei, Chenyang Wan, Xiqian Yu, Tai Wang, Yuqiang Yang, Xiaohan Mao, Chenming Zhu, Wenzhe Cai, Hanqing Wang, Yilun Chen, Xihui Liu, Jiangmiao Pang ·

    StreamVLN: 通过慢速-快速上下文建模实现流式视觉语言导航

    arXiv:2507.05240v2 Announce Type: replace-cross Abstract: Vision-and-Language Navigation (VLN) in real-world settings requires agents to process continuous visual streams and generate actions with low latency grounded in language instructions. While Video-based Large Language Mod…

  358. arXiv cs.CV TIER_1 English(EN) · Zuntao Liu, Yi Du, Taimeng Fu, Shaoshu Su, Cherie Ho, Chen Wang ·

    用于空间推理的视觉语言记忆

    arXiv:2511.20644v2 Announce Type: replace Abstract: Spatial reasoning is a critical capability for intelligent robots, yet current vision-language models (VLMs) still fall short of human-level performance in video-based spatial reasoning. This gap mainly stems from two challenges…

  359. arXiv cs.CV TIER_1 English(EN) · Mingjia Shi, Shuo Wang, Xiaobo Wang, Sifan Zhou, Kai Wang, Tianyu Fu, Chenxu Zhao, Anyang Su, Ping Jiang, Minghui Wu ·

    深入探讨低秩视觉语言对齐的隐式偏差

    arXiv:2607.08194v1 Announce Type: new Abstract: Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requiring full-parameter updates. We challenge this view and investigate what happens…

  360. arXiv cs.CV TIER_1 English(EN) · Yan Xie ·

    MOSAIC:异构视觉语言模型的高效自适应层间组合

    Vision-Language Models (VLMs) have achieved success using homogeneous Transformers to process multimedia data. Recent studies show that heterogeneous structures interleaving efficient mechanisms, like linear attention, improve both performance and inference latency over homogeneo…

  361. arXiv cs.CV TIER_1 English(EN) · Salimeh Sekeh ·

    C-GAP:类别感知和在线提示可改善不平衡类别的视觉语言模型

    Safety-critical perception systems must reliably detect rare object classes within small label spaces, a setting that long-tailed detection methods, designed for hundreds of classes with dense annotation, fundamentally do not address. Open-vocabulary detectors offer a promising a…

  362. arXiv cs.CV TIER_1 English(EN) · Minghui Wu ·

    深入探讨低秩视觉语言对齐的隐式偏差

    Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requiring full-parameter updates. We challenge this view and investigate what happens when low-rank adaptation is applied to the LLM …

  363. arXiv cs.CV TIER_1 English(EN) · Hongyu Qu, Jianzhe Gao, Xiaobin Hu, Shaohuan Yang, Xinlei Yu, Rui Yan, Wenguan Wang, Xiangbo Shu, Shuicheng Yan ·

    面向机器人操作的双重潜在记忆视觉-语言-动作模型

    arXiv:2607.07608v1 Announce Type: cross Abstract: Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Existing memory-augmented VLAs eith…

  364. arXiv cs.CV TIER_1 English(EN) · Jiajun Wu ·

    APIVOT:具有交错视觉-语言思考的自适应规划

    Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility. To successfully execute a task, a robot must decompose goals, select task-relevant objects, and sequence actions, while ensuring that plans satisfy spatial constraints s…

  365. arXiv cs.CV TIER_1 English(EN) · Shuicheng Yan ·

    用于机器人操作的双重潜在记忆视觉-语言-动作模型

    Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Existing memory-augmented VLAs either expand the observation window or retrieve histo…

  366. arXiv cs.CV TIER_1 English(EN) · Jiho Choi, Jaemin Kim, Sanghwan Kim, Seunghoon Hong, Jin-Hwi Park ·

    沉没的帮助或伤害:大型视觉语言模型中注意力沉没的统一框架

    arXiv:2604.03316v3 Announce Type: replace Abstract: Attention sinks are defined as tokens that attract disproportionate attention. While these have been studied in single modality transformers, their cross-modal impact in Large Vision-Language Models (LVLM) remains largely unexpl…

  367. arXiv cs.CV TIER_1 English(EN) · Dominick Reilly, Manish Kumar Govind, Le Xue, Srijan Das ·

    VisCoP:视觉探针用于视觉语言模型的视频域自适应

    arXiv:2510.13808v2 Announce Type: replace Abstract: Large Vision Language Models (VLMs) excel at general visual reasoning but experience significant performance degradation when deployed in novel domains that exhibit substantial distribution shifts from their pretraining data. Ex…

  368. arXiv cs.CV TIER_1 English(EN) · Xinda Liu, Qinyu Zhang, Weiqing Min, Guohua Geng, Shuqiang Jiang ·

    面向细粒度图像识别的视觉-语言模型结构化-压缩提示调优

    arXiv:2607.06185v1 Announce Type: new Abstract: Fine-grained image recognition poses a significant challenge due to the substantial expertise and effort required for manual annotation. Vision-language models (VLMs) like CLIP provide a compelling zero-shot alternative, reducing re…

  369. arXiv cs.CV TIER_1 English(EN) · Zhiwei Yang, Yuanchen Wu, Nan Zhang, Yucong Meng, Ke Yan, Shouhong Ding ·

    场景图思维:增强多模态大语言模型的结构化视觉推理能力

    arXiv:2607.05716v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong perception and reasoning capabilities. However, most existing models focus on isolated objects and neglect structured relationships for efficient target navigation, l…

  370. arXiv cs.CV TIER_1 English(EN) · Shuqiang Jiang ·

    面向细粒度图像识别的视觉-语言模型结构化-压缩提示调优

    Fine-grained image recognition poses a significant challenge due to the substantial expertise and effort required for manual annotation. Vision-language models (VLMs) like CLIP provide a compelling zero-shot alternative, reducing reliance on extensive labeled data. However, their…

  371. arXiv cs.CV TIER_1 English(EN) · Zaiyu Cheng, Khai-Nguyen Nguyen, Antonio Mastropaolo ·

    视觉语言模型在UML图解释中的先验偏差

    arXiv:2607.02853v1 Announce Type: new Abstract: Vision Language Models (VLMs) are increasingly applied to software engineering artifacts, especially UML class diagrams whose meaning depends on visual notation. Yet, it is unclear whether VLMs actually read such diagrams or instead…

  372. arXiv cs.CV TIER_1 English(EN) · Jaehyun Kwak, Nam Cao, Boryeong Cho, Segyu Lee, Sumyeong Ahn, Se-Young Yun ·

    面向大型视觉语言模型的对抗性攻击的阶段性注意力引导区域排序

    arXiv:2602.04356v2 Announce Type: replace Abstract: Targeted adversarial attacks on Large Vision-Language Models (LVLMs) test whether small image perturbations can steer model responses toward attacker-specified content. Under the standard L-infinity constraint, targeted attacks …

  373. arXiv cs.CV TIER_1 English(EN) · Zhaoxu Li, Chenqi Kong, Yi Yu, Qiangqiang Wu, Xinghao Jiang, Ngai-Man Cheung, Bihan Wen, Alex Kot, Xudong Jiang ·

    SAVER:通过风格感知视觉早期修正减轻大型视觉语言模型的幻觉

    arXiv:2508.03177v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their real-world applicability. Although previous mitiga…

  374. arXiv cs.CV TIER_1 English(EN) · Xiangyu Shi, Ruoxi Yang, Wei Tao, Jiwen Zhang, Yanyuan Qiao, Qi Wu ·

    从区域抵达至实例级接地:视觉与语言导航

    arXiv:2607.03792v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) agents may satisfy conventional success criteria while still failing to establish reliable object-level grounding, because current evaluation protocols mainly reward stopping within a 3-meter r…

  375. arXiv cs.CV TIER_1 English(EN) · Suryanarayana Reddy Yarrabothula, Manisha Chawla, Kunal Sinha, Gagan Raj Gupta, Sashank Lekkala, Ashirvadhan Dosapati, Saikamal Nannuri, Katragadda Ajay RamaSwamy Chowdary Gowtham ·

    SteelBench:在真实工业环境中评估视觉语言模型

    arXiv:2607.05264v1 Announce Type: new Abstract: Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test vision-language models under the visual and procedural conditions of real indust…

  376. arXiv cs.CV TIER_1 English(EN) · Pin Tang, Guoqing Wang, Xiangxuan Ren, Zhongdao Wang, Guodongfang Zhao, Bailan, Chao Ma ·

    PixelPilot:可扩展的视觉-语言-动作模型,用于端到端自动驾驶

    arXiv:2607.04637v1 Announce Type: new Abstract: Vision-Language-Action Models (VLAs), which leverage the advanced reasoning capabilities of Vision-Language Models (VLMs), show promising generalization in complex autonomous driving scenarios. Existing VLAs typically predict and op…

  377. arXiv cs.CV TIER_1 English(EN) · Amanda Adkins, Tarunvidyut Ravisankar, Joydeep Biswas ·

    PreSIST:开放世界场景中视觉-语言信息驱动的物体持久性预测

    arXiv:2607.04057v1 Announce Type: new Abstract: Robots deployed over long periods must reason about environments that change over time. Existing long-term perception systems often address object change reactively, updating their maps only after revisiting a scene and observing th…

  378. arXiv cs.CV TIER_1 English(EN) · Anas Zafar, Leema Krishna Murali, Siddhant Bharadwaj, Ashish Vashist, Jia Wu ·

    医疗视觉语言模型(Medical Vision Language Models)真的能“看见”吗?一个反事实接地框架和硬负例对比训练用于视觉依赖的医疗VLM

    arXiv:2607.03647v1 Announce Type: new Abstract: Large vision language models (VLMs) report strong accuracy on medical question-answering, yet it remains unclear whether they reason from visual evidence or exploit textual shortcuts. We introduce a counterfactual evaluation framewo…

  379. arXiv cs.CV TIER_1 English(EN) · Geng Li, Yuxin Peng ·

    BVS:基于多模态大语言模型的贝叶斯视觉搜索,用于细粒度感知

    arXiv:2607.03184v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) demonstrate impressive general capabilities, they struggle with fine-grained perception in ultra-high-resolution (UHR) images, particularly for tiny objects in cluttered scenes. Existin…

  380. arXiv cs.CV TIER_1 English(EN) · Haoyu Zhang, Yangyang Guo, Mohan Kankanhalli ·

    大型视觉语言模型过载以进行越狱

    arXiv:2607.02961v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) exhibit remarkable vision-language capabilities and are increasingly deployed in real-world applications such as personal assistants, document analysis systems, and embodied agents. However, thei…

  381. arXiv cs.CV TIER_1 English(EN) · Marwane Hariat, David Filliat, Antoine Manzanera ·

    VLRC:视觉-语言重投影一致性作为可扩展信号,用于更好的前馈三维预训练

    arXiv:2607.02707v1 Announce Type: new Abstract: Feed-forward 3D models are commonly trained using either expensive geometric supervision or self-supervised photometric objectives, both of which provide incomplete learning signals. We introduce Vision-Language Reprojection Consist…

  382. arXiv cs.CV TIER_1 English(EN) · Katragadda Ajay RamaSwamy Chowdary Gowtham ·

    SteelBench:在真实工业环境中评估视觉语言模型

    Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test vision-language models under the visual and procedural conditions of real industrial CCTV, where workers appear as distant figur…

  383. arXiv cs.CV TIER_1 English(EN) · Cheng Chen, Yuyu Guo, Pengpeng Zeng, Jingkuan Song, Peng Di, Hang Yu, Lianli Gao ·

    从一对一到多对多:深度视觉语言融合的动态跨层注入

    arXiv:2601.10710v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) create a severe visual feature bottleneck by using a crude, asymmetric connection that links only the output of the vision encoder to the input of the large language model (LLM). This static archite…

  384. arXiv cs.CV TIER_1 English(EN) · Yiqian Liu, Iuliia Kotseruba, John K. Tsotsos ·

    通过深度排序任务解耦视觉语言模型中的图像线索理解与语言偏见

    arXiv:2607.01503v1 Announce Type: new Abstract: In this paper, we study depth perception of vision-language models (VLMs) to isolate the effects of pictorial depth cues and disentangle vision and language influences on model performance. To this end, we combine depth-ordering and…

  385. arXiv cs.CV TIER_1 English(EN) · Lev Sorokin, Chen Yang, Ken E. Friedl, Andrea Stocco ·

    面向车内场景理解的基于搜索的视觉语言模型测试

    arXiv:2607.02300v1 Announce Type: new Abstract: In the automotive domain, in-car scene understanding (ISU) enables the detection of safety-critical events, such as driver distraction, and supports drivers or passengers by analyzing the in-car scene and adapting the environment (e…

  386. arXiv cs.CV TIER_1 English(EN) · Zelin Peng, Yichen Zhao, Yu Huang, Piao Yang, Feilong Tang, Zhengqin Xu, Xiaokang Yang, Wei Shen ·

    NEARL:用于医学视觉语言理解的具有正交正则化的交互式查询自适应

    arXiv:2508.04101v2 Announce Type: replace Abstract: Computer-aided medical image analysis is crucial for disease diagnosis and treatment planning. While vision-language models (VLMs) such as CLIP exhibit strong generalization ability, their direct application to medical imaging r…

  387. arXiv cs.CV TIER_1 English(EN) · Andrea Stocco ·

    面向车内场景理解的基于搜索的视觉语言模型测试

    In the automotive domain, in-car scene understanding (ISU) enables the detection of safety-critical events, such as driver distraction, and supports drivers or passengers by analyzing the in-car scene and adapting the environment (e.g., ambient lighting). The industry is increasi…

  388. arXiv cs.CV TIER_1 English(EN) · Sangyun Chung, Youngjoon Yu, Se Yeon Kim, Youngchae Chee, Yong Man Ro ·

    面向多样化传感器理解的增强视觉-语言模型:成本效益优化与基准测试

    arXiv:2412.20750v3 Announce Type: replace Abstract: Large-scale Vision-Language Models (VLMs) have achieved notable progress in aligning visual inputs with text. However, their ability to deeply understand the unique physical properties of non-RGB vision sensor images remains lim…

  389. arXiv cs.CV TIER_1 English(EN) · Diogo Gl\'oria-Silva, Jo\~ao Cardeira, Manuel Letras da Luz, Afonso Simpl\'icio, Gon\c{c}alo Vinagre, Diogo Tavares, Rafael Ferreira, In\^es Calvo, In\^es Vieira, David Semedo, Jo\~ao Magalh\~aes ·

    AMALIA-VL:原生欧洲葡萄牙语开源视觉与语言模型

    arXiv:2606.19100v3 Announce Type: replace Abstract: Large Vision and Language Models (LVLMs) have advanced rapidly, yet European Portuguese (pt-PT) remains systematically underserved by existing open-source multimodal models, which either conflate it with Brazilian Portuguese or …

  390. arXiv cs.CV TIER_1 English(EN) · Ke Wu, Yanan Zhang, Yingjie Gao, Wenhao Li, Chenyu Zhou, XinZhu Ma, Jiaxin Chen, Di Huang ·

    DroneFINE:无人机图像的视觉语言检测器的域感知参数高效微调

    arXiv:2607.00338v1 Announce Type: new Abstract: Object detection for Unmanned Aerial Vehicles (UAVs) working in open and dynamic environments is a highly challenging task. While Vision-Language Models (VLMs) have offered a powerful solution for universal object detection, adaptin…

  391. arXiv cs.CV TIER_1 English(EN) · Kensuke Nakamura, Byung-Woo Hong ·

    通过视觉语言模型的上下文推理实现个性化物体识别与定位

    arXiv:2607.00357v1 Announce Type: new Abstract: Personalized object localization (POL) localizes an object instance in a query image based on a few reference images with bounding-box annotations and a target object label. The pioneering method, IPLoc, solves this task through in-…

  392. arXiv cs.CV TIER_1 English(EN) · Kyan Mahajan, Mohammad Saqlain ·

    SpiralFovea:输入自适应注视式分词作为资源自适应推理的第三个杠杆

    arXiv:2607.00780v1 Announce Type: new Abstract: Most adaptive-inference techniques for foundation models change what the model does - early exit, MoE routing, KV-cache compression, dynamic attention sparsity. The input that hits the backbone, however, remains a fixed-grid tokenis…

  393. arXiv cs.CV TIER_1 English(EN) · Mohammad Saqlain ·

    SpiralFovea:输入自适应注视点标记化作为资源自适应推理的第三个杠杆

    Most adaptive-inference techniques for foundation models change what the model does - early exit, MoE routing, KV-cache compression, dynamic attention sparsity. The input that hits the backbone, however, remains a fixed-grid tokenisation indifferent to image content. We argue tha…

  394. arXiv cs.CV TIER_1 English(EN) · Benoît Macq ·

    面向视觉-语言模型的图像分类的局部一致性预测

    Conformal predictions have attracted significant attention in the field of uncertainty quantification, mainly because of their strong marginal coverage guarantees. Full conditional guarantee is not an attainable goal, a well known fact in conformal predictions literature. As a re…

  395. arXiv cs.CV TIER_1 English(EN) · Dieter Schmalstieg ·

    边绘制边思考:用于增量式三维场景图的异步视觉语言代理

    Open-vocabulary 3D scene graph methods typically operate in two stages: first reconstruct, then enrich with vision-language models, leaving the graph unqueryable during exploration. We argue that this sequential coupling is unnecessary and propose an asynchronous architecture in …

  396. arXiv cs.CV TIER_1 English(EN) · Fawaz Sammani, Tzoulio Chamiti, Nikos Deligiannis ·

    关于视觉语言模型测试时缩放的研究

    arXiv:2606.28864v1 Announce Type: new Abstract: Test-time scaling is a paradigm where large models use additional compute at inference to achieve better performance, without changing model weights. While it has been widely studied for Large Language Models (LLMs), its applicabili…

  397. arXiv cs.CV TIER_1 English(EN) · Enshen Zhou, Yibo Li, Jingkun An, Jiayuan Zhang, Shanyu Rong, Mengzhen Liu, Yi Han, Yuheng Ji, Huajie Tan, Jiawei He, Pengwei Wang, Zhongyuan Wang, Cheng Chi, Lu Sheng, Shanghang Zhang ·

    面向机器人视觉-语言模型的空间轨迹推理

    arXiv:2512.13660v3 Announce Type: replace-cross Abstract: Spatial tracing, as a fundamental embodied interaction ability for robots, is inherently challenging as it requires multi-step metric-grounded reasoning compounded with complex spatial referring and real-world metric measu…

  398. arXiv cs.CV TIER_1 English(EN) · Wenhui Liao, Hongliang Li, Pengyu Xie, Xinyu Cai, Yufan Shen, Yi Xin, Qi Qin, Shenglong Ye, Tianbin Li, Ming Hu, Junjun He, Yihao Liu, Wenhai Wang, Min Dou, Bin Fu, Botian Shi, Yu Qiao, Lianwen Jin ·

    HSD:用于文档解析视觉语言模型的无训练加速,具有分层推测解码

    arXiv:2602.12957v3 Announce Type: replace Abstract: Document parsing is a fundamental task in multimodal understanding, supporting a wide range of downstream applications such as information extraction and intelligent document analysis. Benefiting from strong semantic modeling an…

  399. arXiv cs.CV TIER_1 English(EN) · Duo Zheng, Shijia Huang, Yanyang Li, Liwei Wang ·

    Efficient-VLN:一个简单而强大的高效视觉语言导航基线

    arXiv:2512.10310v2 Announce Type: replace Abstract: While Multimodal Large Language Models (MLLMs) have demonstrated significant promise in Vision-Language Navigation (VLN), existing agents remain heavily constrained by systemic bottlenecks across inference, training, and data co…

  400. arXiv cs.CV TIER_1 English(EN) · Atif Belal, Heitor R. Medeiros, Marco Pedersoli, Eric Granger ·

    VLOD-TTA:视觉-语言目标检测器的测试时自适应

    arXiv:2510.00458v3 Announce Type: replace Abstract: Vision-language object detectors (VLODs) such as YOLO-World and Grounding DINO exhibit strong zero-shot generalization, but their performance degrades under distribution shift. Test-time adaptation (TTA) offers a practical way t…

  401. arXiv cs.CV TIER_1 English(EN) · Kai Jiang, Ruishu Zhu, Siqi Huang, Hongyuan Zhang, Xuelong Li ·

    用于减少多模态大语言模型视觉冗余的潜在噪声掩码

    arXiv:2606.30168v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) often fail in fine-grained visual reasoning, as question-relevant visual cues are diluted by dense and redundant image tokens. Recent multimodal reasoning methods usually extend chain-of-thou…

  402. arXiv cs.CV TIER_1 English(EN) · Hong-Han Wang, Yuntao Wang, Hu Ding ·

    MIRROR:通过Gromov-Wasserstein实现语言到图像的语义关系对齐

    arXiv:2606.29462v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) inherit rich relational priors from their language backbones, yet often fail when asked to apply these relationships in visual contexts. We trace this failure to a structural blind spot: proj…

  403. arXiv cs.CV TIER_1 English(EN) · Xuelong Li ·

    用于减少多模态大语言模型视觉冗余的潜在噪声掩码

    Multimodal large language models (MLLMs) often fail in fine-grained visual reasoning, as question-relevant visual cues are diluted by dense and redundant image tokens. Recent multimodal reasoning methods usually extend chain-of-thought from language models into visual or latent s…

  404. arXiv cs.CV TIER_1 English(EN) · Jiajun Cheng, Xianwu Zhao, Sainan Liu, Xiaofan Yu, Ravi Prakash, Patrick J. Codd, Jonathan Elliott Katz, Shan Lin ·

    SurgXBench:用于手术的可解释视觉-语言模型基准测试

    arXiv:2505.10764v4 Announce Type: replace Abstract: Innovations in digital intelligence are transforming robotic surgery with more informed decision-making. Real-time awareness of surgical instrument presence and actions (e.g., cutting tissue) is essential for such systems. Yet, …

  405. arXiv cs.CV TIER_1 English(EN) · Chanik Kang, Rapha\"el Pestourie, Haejun Chung ·

    面向冻结视觉语言模型的VLM感知超表面前端设计

    arXiv:2606.27646v1 Announce Type: new Abstract: Conventional machine-vision pipelines typically rely on high-quality optics that produce clean, human-interpretable images, and optical design has therefore been driven by image-level criteria such as resolution, aberration correcti…

  406. arXiv cs.CV TIER_1 English(EN) · Nan Yang, Zhanwen Liu, Linfeng Zhang, Shangyu Xie, Yang Wang, Wenzhuo Zhou, Xiangmo Zhao ·

    MVPruner:用于加速自动驾驶多视图视觉语言模型的动态令牌修剪

    arXiv:2606.27660v1 Announce Type: new Abstract: Vision-Language Models (VLMs) improve generalization and interpretability in autonomous driving but suffer from efficiency issues due to long visual token sequences, particularly in standard multi-view settings. Existing token pruni…

  407. arXiv cs.CV TIER_1 English(EN) · Chaoxiang Cai, Minghe Weng, Jie Li, Yibo Jiang, Longrong Yang, Zequn Qin, Xi Li ·

    GeMoE:仅需门控熵即可实现基于MoE的大型视觉语言模型中的不确定性感知自适应路由

    arXiv:2606.26287v1 Announce Type: new Abstract: With the increase in model parameters and training data, the instruction following and generalization capabilities of Large VisionLanguage Models (LVLMs) have been significantly improved. Based on the Mixture of Experts (MoE) archit…

  408. arXiv cs.CV TIER_1 English(EN) · Sourabh Sharma, Sonam Gupta, Sadbhawna ·

    通过感知验证的自训练改进视觉语言模型的推理能力

    arXiv:2606.22158v2 Announce Type: replace Abstract: Achieving human-like reasoning in Vision-Language Models (VLMs) remains a long-standing challenge. Recent approaches leverage Chain-of-Thought (CoT) rationales generated by human annotators or proprietary models to improve reaso…

  409. arXiv cs.CV TIER_1 English(EN) · Xiangmo Zhao ·

    MVPruner:加速自动驾驶多视图视觉语言模型的动态令牌剪枝

    Vision-Language Models (VLMs) improve generalization and interpretability in autonomous driving but suffer from efficiency issues due to long visual token sequences, particularly in standard multi-view settings. Existing token pruning methods employ fixed pruning rate allocation …

  410. arXiv cs.CV TIER_1 English(EN) · Haejun Chung ·

    面向冻结视觉语言模型的VLM感知超透镜前端设计

    Conventional machine-vision pipelines typically rely on high-quality optics that produce clean, human-interpretable images, and optical design has therefore been driven by image-level criteria such as resolution, aberration correction, and pixel fidelity. However, such optics are…

  411. arXiv cs.CV TIER_1 English(EN) · Liang Zhang ·

    通过最优传输语义流连接视觉与语言概念

    Concept Bottleneck Models (CBMs) promise transparent reasoning by predicting through human-interpretable concepts, yet their effectiveness fundamentally depends on how well visual and textual representations are aligned or matched. Existing vision-language CBMs often rely on pre-…

  412. arXiv cs.CV TIER_1 English(EN) · Huizhen Shu, Xuying Li, Hongxu Lin, Wenjie Sun, Hui Li ·

    使用联合稀疏自编码器引导视觉语言模型

    arXiv:2606.25657v1 Announce Type: new Abstract: Sparse Autoencoders (SAEs) have shown promise for analyzing language models, but applying them to vision-language models (VLMs) often yields representations that are difficult to use as controllable cross-modal steering directions. …

  413. arXiv cs.CV TIER_1 English(EN) · Qitong Wang, Fan Du, Pranav Maneriker, Jihui Jin, Christopher Rasmussen ·

    迈向低延迟的视觉语言模型,在以自我为中心的视觉理解中实现双重正确预测

    arXiv:2606.25160v1 Announce Type: cross Abstract: The rapid rise of Vision-Language Models (VLMs) in egocentric visual understanding has made low-latency inference in human-robot collaborative (HRC) tasks increasingly critical. Weight pruning techniques developed for VLMs to shri…

  414. arXiv cs.CV TIER_1 English(EN) · Weiran Huang ·

    面向视觉语言模型的黑盒持续学习

    The rapid deployment of Vision-Language Models (VLMs) in dynamic environments necessitates the ability to learn continuously without forgetting. However, traditional continual learning (CL) settings often rely on white-box paradigms, which is increasingly invalidated by the shift…

  415. arXiv cs.CV TIER_1 English(EN) · Thomas Brox ·

    面向放射学的可扩展空间定位二维视觉语言模型训练

    We study how to train visually grounded vision-language models (VLMs) for radiology without manual spatial annotations. We introduce RefRad2D, a large-scale bilingual (German/English) dataset of 1.2M CT and MR image-text pairs derived from clinical practice, with task-specific VQ…

  416. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    ICML 2026 | 视觉-语言模型的语义鲁棒性认证

    <h1 style="font-size: 15px; line-height: 1.85; margin: 24px 0 14px; font-weight: 700; color: #111827;"><span><br /></span></h1><h1 style="font-size: 15px; line-height: 1.85; margin: 24px 0 14px; font-weight: 700; color: #111827;">原文作者:公众号“专知”</h1><p>原文链接:<a href="https://mp.weixi…

  417. Medium — fine-tuning tag TIER_1 English(EN) · Arpit Neewaliya ·

    我如何微调了一个视觉语言模型以理解无人机图像

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@arpitneewaliya/how-i-fine-tuned-a-vision-language-model-for-drone-image-understanding-ab3329d6c210?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1536/1*kzeM1R7oQ…

  418. Towards AI TIER_1 English(EN) · Sanjana Dubey ·

    视觉语言基础:AI如何将“狗”与像素关联,以及其局限性

    <h4><em>A research grounded deep dive</em></h4><h3>1. The Hook</h3><p>Picture an ordinary photo like this one. A vision language model will confidently caption it “a dog in a park,” and it will be right, almost every time [1].</p><figure><img alt="A golden retriever sits on green…

  419. dev.to — LLM tag TIER_1 English(EN) · pixelbank dev ·

    视觉-语言模型 — 深度解析 + 问题:多项式回归误差

    <p><em>A daily deep dive into llm topics, coding problems, and platform features from <a href="https://pixelbank.dev" rel="noopener noreferrer">PixelBank</a>.</em></p> <h2> Topic Deep Dive: Vision-Language Models </h2> <p><em>From the Multimodal LLMs chapter</em></p> <h2> Introdu…

  420. dev.to — LLM tag TIER_1 English(EN) · Vahid Aghajani ·

    视觉语言模型——当AI学会看和说(第三部分,共三部分)

    <blockquote> <p>Originally published on <a href="https://software-engineer-blog.com/content/vision-language-models-when-ai-learns-to-see-and-talk-part-3-of-3?id=51" rel="noopener noreferrer">my blog</a>. Cross-posted here with a canonical link.</p> </blockquote> <p> </p> <p>This …

  421. Mastodon — mastodon.social TIER_1 English(EN) · appassionato ·

    2/ 生成式AI:通过奇瑞与大型语言和视觉模型的集成提供支持,使机器人能够理解自然语言,处理视觉环境

    2/ Generative AI: Powered by Chery's integration with large language and vision models, allowing the robot to understand natural speech, process visual environments, and carry on human-like conversations. Combined with Ai (Artificial Intelligence), the name essentially translates…

  422. Mastodon — mastodon.social TIER_1 English(EN) · notatechguy ·

    VAORA 使视觉-语言模型推理与物理动作对齐 VAORA 是 arXiv 上的一项新奖励设计,旨在解决幻觉推理和推理-动作不匹配问题

    VAORA aligns vision-language model reasoning with physical actions VAORA, a new reward design on arXiv, targets hallucinated reasoning and reasoning-action misalignment in vision-language models on physical tasks. https://www. notatechguy.com/vaora-aligns-v ision-language-model-r…