BLIP-2
PulseAugur coverage of BLIP-2 — every cluster mentioning BLIP-2 across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New LLM generates interpretable behavior descriptions for autonomous vehicles
Researchers have developed CommandLM, a novel multimodal large language model designed to generate human-readable descriptions of ego vehicle behavior from fused sensor data. This model integrates LiDAR and multi-camera…
-
HorusEye framework uses language as dynamic attention for emergency visual analysis
A new research paper introduces HorusEye, a framework designed for emergency visual analysis that treats language as dynamic attention. The study benchmarks various vision-language models (VLMs) like Gemini, Qwen2-VL, B…
-
Medical VLM benchmarks show pretraining contamination, study finds
Researchers have audited public medical vision-language benchmarks for pretraining contamination, finding measurable image-side overlap on the SLAKE-En benchmark with models like SigLIP-B-16. Text analysis revealed cano…
-
RadJEPA: Self-supervised model for chest X-ray analysis without language
Researchers have developed RadJEPA, a novel self-supervised learning framework for medical image analysis, specifically for chest X-rays. Unlike previous methods that rely on paired image-text data, RadJEPA learns from …
-
New framework improves medical image segmentation and diagnosis
Researchers have developed Rad-VLSM, a novel two-stage framework designed to enhance medical image segmentation and diagnosis. This system uses a vision-language model to identify potential lesion areas and convert them…
-
New research reveals universal adversarial attacks on VLMs are less effective than previously thought
Researchers have developed a new evaluation method, VisInject, to distinguish between general disruption and precise injection in adversarial attacks on vision-language models. Their findings indicate that while many at…