PulseAugur
EN
LIVE 07:25:25

New framework AffordAny enables 3D affordance grounding from single images

Researchers have developed AffordAny, a novel framework for open-world 3D affordance grounding using monocular RGB images. This system constructs large-scale text-conditioned 3D part supervision and grounds affordances with a guided decoder, improving generalization through pseudo-label self-training. AffordAny creates a benchmark of over 5,000 objects and 10,000 part-level samples, significantly expanding categorical diversity compared to previous methods. The framework demonstrates effectiveness and robustness, achieving strong performance on unseen objects and categories. AI

IMPACT This research advances open-world 3D understanding from single images, potentially improving robotics and AR/VR applications.

RANK_REASON The cluster describes a new research paper detailing a novel framework for 3D affordance grounding. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New framework AffordAny enables 3D affordance grounding from single images

COVERAGE [1]

  1. arXiv cs.CV TIER_1 English(EN) · Junqi Wu, Kaihua Tang, Xuanwen Chen, Hongzhi Li, Jianqiang Huang, Xian-Sheng Hua ·

    AffordAny: Open-World 3D Affordance Grounding from Monocular RGB Images via Vision-Language-Guided Geometric Reasoning

    arXiv:2608.20720v1 Announce Type: new Abstract: Open-world 3D affordance grounding requires localizing functional object parts in 3D given free-form language queries. Existing methods typically assume pre-built object-centric 3D geometry and closed affordance ontologies, limiting…