Large Multimodal Models (LMMs)
PulseAugur coverage of Large Multimodal Models (LMMs) — every cluster mentioning Large Multimodal Models (LMMs) across labs, papers, and developer communities, ranked by signal.
-
New benchmarks tackle hallucination in GI endoscopy AI models
Researchers have developed new benchmarks and datasets to address hallucination issues in vision-language models (VLMs) used for gastrointestinal endoscopy. One study introduces a benchmark using the Gut-VLM dataset to …
-
ReFine3D framework enhances 3D vision-language model adaptation
Researchers have developed ReFine3D, a new framework for fine-tuning 3D vision-language models. This method addresses the challenge of adapting these models to new domains with limited data, preventing overfitting and c…
-
New framework efficiently selects data for multimodal models
Researchers have developed a new framework called One-Step-Train (OST) to efficiently select high-quality synthetic data for training large multimodal models (LMMs). OST reframes data selection as an incremental optimiz…