PulseAugur
EN
LIVE 08:31:09
中文(ZH) 别再乱调图像塔了!浙大 IJCAI 论文揭露 VLM 非对称性真相,给 CLIP 微调「踩刹车」

Zhejiang University unveils adaptive adapter to improve VLM generalization

Researchers from Zhejiang University and Swansea University have developed an Adaptive Asymmetric Adapter (A3B2) to improve the fine-tuning of Vision-Language Models (VLMs). Their findings indicate that aggressive fine-tuning of the image encoder can degrade generalization capabilities, especially on out-of-distribution data. The A3B2 system introduces a mechanism that automatically 'brakes' the image encoder's updates when the model's confidence is low, preserving the pre-trained model's robustness while allowing for efficient adaptation. AI

IMPACT This research offers a novel approach to VLM fine-tuning, potentially improving performance on real-world, out-of-distribution data by preventing over-adaptation.

RANK_REASON Academic paper detailing a new method for fine-tuning VLMs. [lever_c_demoted from research: ic=1 ai=1.0]

Read on 雷峰网 (Leiphone) →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Zhejiang University unveils adaptive adapter to improve VLM generalization

COVERAGE [1]

  1. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Stop Adjusting Image Towers Randomly! Zhejiang University IJCAI Paper Reveals the Truth of VLM Asymmetry, Putting the Brakes on CLIP Fine-tuning

    <section style="text-align: center; margin: 0px 16px; line-height: 1.75em; display: block;"><img class="rich_pages wxw-img" src="https://static.leiphone.com/uploads/new/images/20260817/6a8291614dca4.jpg?imageMogr2/quality/90" style="width: 100%; display: inline-block; text-align:…