NYUv2
PulseAugur coverage of NYUv2 — every cluster mentioning NYUv2 across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
YouTube-Occ learns 3D semantic occupancy from web videos
Researchers have developed YouTube-Occ, a novel method for predicting 3D semantic occupancy from internet videos, addressing the scarcity of annotated 3D indoor data. The system utilizes a pipeline that processes raw we…
-
PixelUp enhances vision models with zero-shot feature upsampling
Researchers have developed PixelUp, a novel zero-shot method for upsampling features from Vision Foundation Models (VFMs). This technique aims to improve the accuracy of fine-grained vision tasks like semantic segmentat…
-
GeoStereo framework unifies stereo geometry estimation with diffusion priors
Researchers have introduced GeoStereo, a novel framework that unifies stereo geometry estimation for both disparity and surface normal prediction. This approach leverages diffusion priors to enhance performance in chall…
-
DPNeXt framework boosts multi-task dense prediction with efficient ViT fusion
Researchers have introduced DPNeXt, a novel framework designed to enhance multi-task learning for dense prediction tasks in robotics perception. This lightweight system efficiently fuses multi-scale features from Vision…
-
LingBot-Vision uses masked boundary modeling for self-supervised pretraining
Researchers have introduced LingBot-Vision, a new self-supervised pretraining method that focuses on masked boundary modeling. This approach aims to improve performance by forcing the model to reconstruct specific bound…
-
Ant Group's Lingbo releases suite of embodied AI models, including world action and video generation
Ant Group's Lingbo Technology has released a suite of new models aimed at advancing embodied AI and robotics. LingBot-VA 2.0 is presented as the first embodiment-native world action model, designed from the ground up fo…
-
New framework ReLiF improves fairness evaluation in multi-task learning
Researchers have developed a new framework called ReLiF to address issues in evaluating Lipschitz fairness within multi-task learning (MTL). The framework introduces fixed-delta auditing, which uses a shared reference t…
-
SA4Depth improves self-supervised monocular depth estimation
Researchers have introduced SA4Depth, a novel approach to enhance self-supervised monocular depth estimation. This method focuses on improving the alignment between the scale estimates from separate depth and pose netwo…
-
Open-source image editors show surprising zero-shot vision capabilities
Researchers have evaluated three open-source image-editing models—Qwen-Image-Edit, FireRed-Image-Edit, and LongCat-Image-Edit—for their zero-shot vision learning capabilities without any fine-tuning. The study found tha…