Grounding DINO
PulseAugur coverage of Grounding DINO — every cluster mentioning Grounding DINO across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
New GCRL method uses dynamic object masks for visual goal representation
Researchers have developed a novel approach for goal-conditioned reinforcement learning (GCRL) that utilizes dynamic object masks as goal representations. This method bypasses the need for privileged state or position i…
-
AI scientist defines 'world models' as action-conditioned, crucial for embodied AI
Zhang Lei, a leading AI scientist and founder of Visionary Future, argues that the current boom in "world models" within the AI industry is driven by the need to overcome the limitations of imitation learning in embodie…
-
New ABRA method transfers object detection knowledge across domains
Researchers have introduced ABRA (Aligned Basis Relocation for Adaptation), a novel method designed to transfer knowledge from labeled source domains to target domains lacking annotated data for open-vocabulary object d…
-
GeoAI tutorial details building footprint extraction using U-Net, DINO, SAM, and Mask R-CNN
This tutorial details a GeoAI workflow for extracting building footprints from aerial imagery using a combination of deep learning models. It covers setting up the geospatial environment, training a U-Net model with a R…
-
AI object detectors signal presence, not visibility, in cluttered scenes · 2 sources tracked
A new research paper reveals that open-vocabulary object detectors, widely used for tasks like grounding language and active perception, often signal the presence of an object rather than its visibility. Even when an ob…
-
User study finds improved robot interaction system perceptible to users
A new study published on arXiv explores the perceptual differences users experience when interacting with a multimodal human-robot system. The research compared a baseline system using Whisper, Florence-2, and Llama 3.1…
-
New AI framework translates radiologist speech to MRI tumor segmentation
Researchers have developed LoGSAM, a novel framework designed for parameter-efficient segmentation of brain tumors in MRI scans. This system transforms radiologist dictations into text prompts that guide foundation mode…
-
HKVLM model improves visual reasoning by separating localization from language
Researchers have developed HKVLM, a novel approach to visual reasoning that separates localization from language generation. This model utilizes a frozen language-aligned detector and a frozen language model, connected …
-
New framework NegAS boosts out-of-distribution object detection in VLMs
Researchers have introduced NegAS, a novel framework designed to enhance out-of-distribution (OOD) object detection in vision-language models (VLMs). NegAS addresses two key challenges: improving attention mechanisms to…
-
New 3D Detector CT-3GDINO Enhances Organ Localization in Abdominal CT Scans
Researchers have developed CT-3GDINO, a novel 3D object detection model designed for organ localization in abdominal CT scans. This lightweight model adapts a Grounding-DINO-style architecture, utilizing frozen pseudo-t…
-
AI pipeline automates labeling of unknown objects in images
Researchers have developed an automated pipeline to label objects in images that are not recognized by existing open-vocabulary models. This system aims to reduce the tedious manual work of creating bounding boxes for t…
-
Visionary Future develops object-centric latent world models
A Shenzhen-based AI team, Visionary Future, is developing an
-
New zero-shot tracking system excels in multi-animal studies
Researchers have developed a new zero-shot multi-animal tracking system that leverages vision foundation models, specifically adapting SAM2MOT with Grounding DINO and the Segment Anything Model 2. This method achieves s…
-
New methods improve open-vocabulary object detection robustness and adaptation
Researchers have introduced several new methods to improve open-vocabulary object detection, a field that aims to identify arbitrary objects based on human prompts. One approach, EBOD, integrates a prompt-based detector…