Researchers have developed a new framework called PromptCCZSL to address the challenge of continually adapting vision-language models to new attributes and objects without losing previously learned information. This method utilizes session-aware compositional prompts and recency-weighted multi-teacher distillation to fuse multimodal features for novel compositions. The framework also incorporates specific losses, such as Cosine Anchor Loss, Orthogonal Projection Loss, and Intra-Session Diversity Loss, to maintain semantic consistency, ensure distinct embeddings, and promote varied representations. Experiments on UT-Zappos and C-GQA benchmarks show that PromptCCZSL significantly outperforms existing baselines in compositional zero-shot learning. AI
IMPACT Enhances the adaptability and knowledge retention of vision-language models, potentially improving their performance in dynamic environments.
RANK_REASON Academic paper detailing a new framework and methodology for vision-language models. [lever_c_demoted from research: ic=1 ai=1.0]
- C-GQA
- Cosine Anchor Loss
- Intra-Session Diversity Loss
- Orthogonal Projection Loss
- PromptCCZSL
- Sara Nadeem
- UT-Zappos
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →