RefCOCO+
PulseAugur coverage of RefCOCO+ — every cluster mentioning RefCOCO+ across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New adversarial attack targets SAM3 image segmentation models
Researchers have developed Universal Concept Disruption (UCD), a novel adversarial attack specifically designed to target SAM3 image segmentation models. UCD learns a single image perturbation that can disrupt the model…
-
New Hi-Token method enhances visual grounding accuracy in AI models
Researchers have developed Hi-Token, a novel method for generative visual grounding that improves the accuracy of bounding-box predictions by tokenizing coordinates hierarchically. This approach encodes digits for hundr…
-
StepX-Edge: On-Device UI Vision-Language Model Achieves High Accuracy
Researchers have developed StepX-Edge, a 0.9 billion parameter vision-language model designed for on-device UI understanding. This model addresses the trade-off between accuracy and efficiency on mobile devices through …
-
New SVCR Framework Enhances Weakly Supervised Referring Expression Comprehension
Researchers have developed a new framework called Structured Visual Compositional Representation (SVCR) to improve referring expression comprehension (REC) in weakly supervised settings. This framework explicitly models…
-
HKVLM model improves visual reasoning by separating localization from language
Researchers have developed HKVLM, a novel approach to visual reasoning that separates localization from language generation. This model utilizes a frozen language-aligned detector and a frozen language model, connected …
-
CoLA framework enhances multimodal AI adaptation with dual-path LoRA
Researchers have introduced CoLA (Cross-Modal Low-rank Adaptation), a novel framework designed to efficiently adapt foundation models for multimodal tasks. Unlike existing methods that adapt each modality in isolation, …
-
Researchers seek arXiv endorsement for Locate-SAM2 computer vision paper
Two independent researchers are seeking an endorsement for their paper on a new computer vision system called Locate-SAM2. This system connects NVIDIA's LocateAnything-3B with Meta's SAM 2.1 through a lightweight adapte…
-
New Framework Enhances Semi-Supervised Segmentation with LLM Priors
Researchers have introduced "Learning to Label" (L2L), a novel framework designed to improve semi-supervised referring expression segmentation (SS-RES) by treating pseudo-label generation as a learnable process. L2L uti…
-
Visual Para-Thinker introduces parallel reasoning to multimodal LLMs
Researchers have introduced Visual Para-Thinker, a novel framework for parallel reasoning in multimodal large language models (MLLMs). This approach shifts from vertical scaling of reasoning depth to a parallel strategy…