PulseAugur
实时 14:48:23
English(EN) CityLoc: 6DoF Pose Distributional Localization for Text Descriptions in Large-Scale Scenes with Gaussian Representation

CityLoc 方法为基于文本的 3D 场景定位生成姿态分布

研究人员开发了 CityLoc,一种用于在大规模 3D 场景中定位文本描述的新颖方法。该方法通过生成以文本为条件的相机姿态分布来解决此类任务中的固有歧义,从而能够对广泛定义的概念进行更鲁棒的推理。该系统利用了基于扩散的架构,并与 CLIP 集成以实现文本-姿态关联,并通过 3D 高斯泼溅进行视觉推理以纠正错位的样本,进一步增强了其性能。在五个大规模数据集上的实验表明,CityLoc 的性能优于标准的分布估计方法。 AI

影响 这项研究可以改进 AI 系统根据自然语言描述理解和交互复杂 3D 环境的方式。

排序理由 该集群包含一篇详细介绍计算机视觉新方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

CityLoc 方法为基于文本的 3D 场景定位生成姿态分布

报道来源 [1]

  1. arXiv cs.CV TIER_1 English(EN) · Qi Ma, Runyi Yang, Bin Ren, Nicu Sebe, Ender Konukoglu, Luc Van Gool, Danda Pani Paudel ·

    CityLoc:高斯表示法在大规模场景下基于文本描述的6DoF位姿分布定位

    arXiv:2501.08982v3 Announce Type: replace Abstract: Localizing textual descriptions within large-scale 3D scenes presents inherent ambiguities, such as identifying all traffic lights in a city. Addressing this, we introduce a method to generate distributions of camera poses condi…