Researchers have developed 3DZip, a novel three-stage framework designed to compress tokens for 3D vision-language models (3D VLMs). This method addresses the significant computational and memory overhead generated by the thousands of tokens typically produced per scene in 3D VLMs. By employing coarse voxelization, feature-space diversity selection using a Determinantal Point Process, and spatial constraint merging, 3DZip effectively reduces token count while preserving geometric coherence. Experiments show that 3DZip can maintain 94.7% of original performance with only 128 tokens, leading to a 1.92x faster inference speed on 3D question answering benchmarks. AI
IMPACT Reduces computational costs for 3D vision-language models, enabling faster inference and broader application in spatial reasoning tasks.
RANK_REASON The cluster describes a new method published in a research paper on arXiv.
Read on Hugging Face Daily Papers →
- 2D visual features
- 2D VLMs
- 3D Question Answering with Scene Graph Reasoning
- 3D vision-language models
- 3DZip
- arXiv
- Determinantal point process
- Hugging Face
- ECCV 2026
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →