PulseAugur
中
实时 11:06:58

新方法提升视觉语言模型效率与准确性

研究人员开发了两种新方法来提高视觉语言模型(VLM)的效率和准确性。一种方法SCOPD使用on-policy自蒸馏来帮助VLM适应剪枝的视觉上下文,即使在视觉信息减少的情况下也能显著提高性能。另一种方法VQS通过使用解析图像为结构化记录的程序验证的QA生成系统,解决了训练自演化VLM的挑战,从而获得更准确的答案和更好的模型性能。 AI

影响 这些进展提供了更有效、更准确的训练和利用视觉语言模型的方法,有望加速其在各个领域的应用。

排序理由 两篇不同的研究论文介绍了用于视觉语言模型的新方法。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 18 个来源。 我们如何撰写摘要 →

新方法提升视觉语言模型效率与准确性

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
两篇不同的研究论文介绍了用于视觉语言模型的新方法。
Source corroboration
18 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
3 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+16 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [18]

  1. arXiv cs.AI TIER_1 English(EN) · Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang ·

    面向视觉-语言-动作模型的基于失败库的运行时反馈自我进化学习

    arXiv:2609.39820v1 Announce Type: cross Abstract: Vision-language-action (VLA) models generalize broadly across robotic manipulation tasks, but complex environments require balancing task success with unintended contact. Runtime shields can correct individual actions, but they le…

  2. arXiv cs.CL TIER_1 English(EN) · Yanshu Li, Jiaqian Li, Canran Xiao, Xi Xiao, Tianyang Wang, Yongtai Liu ·

    MCD:大型视觉语言模型多模态上下文学习的因果蒸馏

    arXiv:2609.39920v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap…

  3. arXiv cs.AI TIER_1 English(EN) · Yifei Dong, Fengyi Wu, Yilong Dai, Lingdong Kong, Guangyu Chen, Yetong Sha, Qiyu Hu, Feng Liu, Siyu Huang, Qi Dai, Zhi-Qi Cheng ·

    语言条件化世界模型用于视觉导航

    arXiv:2603.26741v2 Announce Type: replace-cross Abstract: Goal-conditioned visual navigation has been a long-standing testbed for embodied AI. We study a natural language-conditioned variant, language-conditioned visual navigation (LCVN), in which an embodied agent must follow a …

  4. arXiv cs.AI TIER_1 English(EN) · Minrui Liu, Jingke Wang, Yuehao Huang, Hao Su, Jiajun Lv, Yukai Ma, Yong Liu ·

    AVERT-VLN:视觉-语言导航的感知错误恢复与训练

    arXiv:2609.39579v1 Announce Type: new Abstract: Deploying vision-and-language navigation (VLN) agents in unseen environments remains challenging because unfamiliar layouts and visual conditions can cause execution to go off track. Rather than relying on continuous human supervisi…

  5. arXiv cs.AI TIER_1 English(EN) · Rafi Ibn Sultan, Xiangyu Zhou, Md. Sajid Alam Chowdhury, Chengyin Li, Prashant Khanduri, Marco Brocanelli, Dongxiao Zhu ·

    SpatialCORE:大型视觉-语言模型中置信度感知的接地空间推理

    arXiv:2609.38716v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spat…

  6. arXiv cs.CL TIER_1 Italiano(IT) · Zhe Qian, Ziyang Gong, Zhongxing Xu, Hehan Li, Zhonghua Wang, Fei Luo, Mingxuan Wang, Xue Yang, Shiwei liu, Yanbiao Ma, Junchi Yan, Jungong Han ·

    LoopVL:循环视觉智能

    arXiv:2609.38426v1 Announce Type: cross Abstract: We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through sh…

  7. arXiv cs.AI TIER_1 English(EN) · Hyunjong Ok, Seunggu Kang, Jaeho Lee ·

    视觉-语言模型基准的层次化压缩

    arXiv:2609.37515v1 Announce Type: cross Abstract: Thorough evaluation of vision-language models (VLMs) has become prohibitively expensive, as benchmarks span an ever-broader spectrum of capabilities and new models arrive at a relentless pace. Benchmark compression methods that pr…

  8. arXiv cs.AI TIER_1 English(EN) · Yang Li, Sijia Zhang, Yihan Li, Aming WU, Zihao Zhang, Ziju Han, Yahong Han ·

    用于测试时自适应视觉-语言导航的信用引导策略改进

    arXiv:2609.37591v1 Announce Type: cross Abstract: Test-time adaptation for vision-language navigation (TTA-VLN) enables pretrained policies to adapt online to unseen environments using only test-time observations and interaction history. However, distribution shifts can distort l…

  9. arXiv cs.AI TIER_1 English(EN) · Gujie Shao, Zixun Xie, Xuechun Xing, Ruixiang Wang, Ziyun Lan, Yanlin Qi, Gangyi Zhang, Yuxin Yang, Dawei Li, Haiming Tang ·

    EyeVQA:眼科视觉-语言模型从识别到空间定位的基准测试

    arXiv:2609.32352v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging remain insufficiently characterized. Existing ophthalmic datasets are typically de…

  10. Hugging Face Daily Papers TIER_1 Italiano(IT) ·

    LoopVL:循环视觉智能

    We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through…

  11. Hugging Face Daily Papers TIER_1 English(EN) ·

    SCOPD:用于高效视觉语言模型的稀疏上下文策略内蒸馏

    Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of ta…

  12. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Salman Khan ·

    面向视觉语言模型的程序验证式自进化

    Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24\% of majority-vote l…

  13. arXiv cs.CV TIER_1 English(EN) · Moseli Mots'oehli, Thulani Babeli ·

    Front-to-Back:面向不对称跨视图车辆重识别的视觉-语言模型基准测试

    arXiv:2609.39492v1 Announce Type: new Abstract: Matching the same vehicle across front and rear cameras is difficult because the cameras do not share a view and the vehicle's appearance changes substantially. We introduce Front2Back-ReID, a benchmark of 500 manually verified vehi…

  14. arXiv cs.CV TIER_1 English(EN) · Yuanhao Sun, Huawei Ji, Jiaxin Ding, Luoyi Fu, Xinbing Wang ·

    提炼视觉证据,而非仅仅答案:面向视觉语言模型的跨世界策略内蒸馏

    arXiv:2609.38777v1 Announce Type: new Abstract: A central goal of vision-language model (VLM) distillation is to transfer both the teacher's language capabilities and its visual understanding. However, existing methods primarily supervise the student's output, leaving visual unde…

  15. arXiv cs.CV TIER_1 English(EN) · Can Zhang, Xiaotian Han, Junyuan Shang, Yuchen Ding, Zhenyu Zhang, Shuohuan Wang, Dianhai Yu, Ruirui Li ·

    RESUME:用于高效视频语言建模的运动和残差信号的递归状态更新

    arXiv:2609.39563v1 Announce Type: new Abstract: Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes between sampled frames. Codec-aware front-ends read the motion vectors and residuals th…

  16. arXiv cs.CV TIER_1 English(EN) · Hongyuan Tao, Xinggang Wang, Lianghui Zhu, Yongkang Li, Yunchao Wei, Bin Feng, Shaoyu Chen, Qian Zhang, Chang Huang, Kai Yu ·

    多模态流:嵌入空间中语言与视觉的统一流建模

    arXiv:2609.40362v1 Announce Type: new Abstract: We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and quantized images as discrete tokens or combine discrete language prediction with c…

  17. arXiv cs.CV TIER_1 English(EN) · Shuming Liu, Zhifang Zhang, Suqin Yuan, Khin Mi Mi Aung, Zhuoyi Lin, Lei Feng ·

    选择性通道恢复用于后门视觉语言模型

    arXiv:2609.37759v1 Announce Type: cross Abstract: Vision-language models (VLMs) exhibit strong multimodal capabilities but remain vulnerable to backdoors implanted through poisoned fine-tuning data. Existing defenses often require extensive parameter updates during fine-tuning or…

  18. arXiv cs.CV TIER_1 English(EN) · Zizhao Li, Chengyi Cai, Mohammed Yaqoob Ansari, Feng Liu, Joseph West, Kourosh Khoshelham ·

    通过结构化提示重参数化重编程视觉语言模型

    arXiv:2609.36680v1 Announce Type: new Abstract: Visual reprogramming adapts pretrained models to downstream tasks by modifying their input and output interfaces while keeping the backbone fixed. In vision-language models, existing methods mainly rely on intra-class prompt aggrega…