Researchers have developed a de-biased protocol using vision-language models (VLMs) to evaluate the quality of 3D meshes generated from single images. This protocol, which involves using distinct VLM judges for training and evaluation and implementing position-bias correction, aims to provide a more reliable assessment than traditional proxies like CLIP similarity or geometry validity. While the protocol proved effective in identifying failure modes and was used to adapt a generator called TRELLIS, the adaptation methods did not surpass the performance of the base model when trained on public data. The study suggests that exceeding base performance requires more than lightweight parameter-efficient fine-tuning on public datasets, and the VLM-judge protocol itself is reusable for evaluation. AI
IMPACT Establishes a new benchmark for evaluating 3D generation quality, potentially guiding future research and development in the field.
RANK_REASON The cluster contains two arXiv papers detailing a new protocol for evaluating 3D mesh generation quality using VLMs.
- Google Scanned Objects
- VLM-Judge
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- InternVL3 8B
- Qwen2.5-VL-7B
- ScienceCast
- TRELLIS
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →