Multimodal foundation models
PulseAugur coverage of Multimodal foundation models — every cluster mentioning Multimodal foundation models across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New 'PRAC' attack targets vision modality of AI agents
Researchers have developed a new attack called PRAC that targets the vision modality of multimodal foundation models, specifically Computer Use Agents (CUAs). Unlike previous attacks that directly manipulated model outp…
-
Visual Prompting: A New Interaction Paradigm for Multimodal AI
Visual prompting is a new paradigm for interacting with multimodal foundation models, distinct from simply providing an image as input. It involves the deliberate design of visual context to guide a model's attention, c…
-
ReFine3D framework enhances 3D vision-language model adaptation
Researchers have developed ReFine3D, a new framework for fine-tuning 3D vision-language models. This method addresses the challenge of adapting these models to new domains with limited data, preventing overfitting and c…
-
Survey paper maps Test-Time Scaling for multimodal AI models
A new survey paper details the emerging field of Test-Time Scaling (TTS) for Multimodal Foundation Models (MFMs). The paper categorizes existing TTS methods into sampling-based, feedback-based, and search-based approach…