Multi-modal Large Language Model
PulseAugur coverage of Multi-modal Large Language Model — every cluster mentioning Multi-modal Large Language Model across labs, papers, and developer communities, ranked by signal.
-
New benchmark RefBench-PRO evaluates MLLM perception and reasoning
Researchers have introduced RefBench-PRO, a new benchmark designed to evaluate the perceptual and reasoning capabilities of Multi-modal Large Language Models (MLLMs) in Referring Expression Comprehension (REC). This ben…
-
AI pipeline groups vacation rental rooms and identifies bed types
Researchers have developed a machine learning pipeline to automatically discover and group similar room scenes in unstructured vacation rental image collections. This system helps travelers understand property layouts a…
-
PointVG-R model enhances visual grounding with geometric reasoning · 3 sources tracked
Researchers have developed PointVG-R, a novel reasoning-guided Multi-modal Large Language Model (MLLM) designed to improve precise pointing localization in images. This model integrates geometric-aware reasoning, Reinfo…