Researchers have developed a new method called Cen-Prune to improve the efficiency of Large Vision-Language Models (LVLMs) by optimizing how visual tokens are pruned. Standard diversity-based pruning relies on cosine similarity, but raw visual token similarities are too concentrated to effectively distinguish redundant tokens. While centering token features before similarity calculation reveals a richer structure, it can degrade performance by losing information about globally distinctive tokens. Cen-Prune addresses this by using centered cosine similarity for diversity measurement while also retaining raw-space distinctiveness, leading to improved performance across various benchmarks and LVLM architectures with minimal computational overhead. AI
IMPACT Optimizes LVLM inference efficiency by improving visual token pruning, potentially reducing computational costs.
RANK_REASON The cluster contains a research paper detailing a new method for optimizing LVLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →