PulseAugur
EN
LIVE 07:25:46

New frameworks enhance multimodal AI by preserving knowledge and improving generation

Researchers are developing new frameworks to enhance multimodal AI models. Rosetta introduces a composable pretraining approach that preserves core knowledge while adding new modalities non-destructively, using Momentum-Anchored Orthogonal Projection to manage gradient conflicts. COMPASS grounds composition-intent control within a unified system, improving both perception and generation by using a shared expert token. SRUM enables unified multimodal models to self-improve their generation capabilities by using their understanding module as an internal evaluator, employing a dual reward system for global and local fidelity. Additionally, ReVisIT offers a training-free method for multimodal in-context learning by using retrieved images as "visual thought" to enhance reasoning and generation. AI

IMPACT These advancements aim to create more robust and versatile multimodal AI systems capable of better understanding and generating content across different data types.

RANK_REASON The cluster contains multiple research papers detailing novel frameworks and methods for multimodal AI.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 16 sources. How we write summaries →

New frameworks enhance multimodal AI by preserving knowledge and improving generation

COVERAGE [16]

  1. arXiv cs.AI TIER_1 English(EN) · Qianyu Chen, Canran Xiao, Runxuan Tang ·

    Hidden Forgetting in Continual Multimodal Learning: When Accuracy Survives but Grounding Fails

    arXiv:2607.02020v1 Announce Type: new Abstract: Multimodal large language models must continually adapt to evolving tasks and domains, yet standard continual learning metrics mainly measure whether old answers remain correct, leaving the stability of multimodal grounding largely …

  2. arXiv cs.AI TIER_1 English(EN) · Yuanzhi Liu, Shousheng Zhao, Bo Zhou, Kongming Liang, Zhanyu Ma ·

    MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

    arXiv:2607.01813v1 Announce Type: cross Abstract: Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to temporal staleness, data contamination, and costly maintenance. We present MMBen…

  3. arXiv cs.AI TIER_1 English(EN) · Runxuan Tang ·

    Hidden Forgetting in Continual Multimodal Learning: When Accuracy Survives but Grounding Fails

    Multimodal large language models must continually adapt to evolving tasks and domains, yet standard continual learning metrics mainly measure whether old answers remain correct, leaving the stability of multimodal grounding largely unexamined. We study this overlooked failure mod…

  4. arXiv cs.AI TIER_1 English(EN) · Guanyu Jiang, Zhaochen Su, Xiaoye Qu, Yi R. Fung ·

    XSkill: Continual Learning from Experience and Skills in Multimodal Agents

    arXiv:2603.12056v3 Announce Type: replace Abstract: Multimodal agents can now tackle complex reasoning tasks with diverse tools, yet they still suffer from inefficient tool use and inflexible orchestration in open-ended settings. A central challenge is enabling such agents to con…

  5. arXiv cs.CL TIER_1 Italiano(IT) · Xiangyue Liu, Zijian Zhang, Miles Yang, Zhao Zhong, Liefeng Bo, Ping Tan ·

    Rosetta: Composable Native Multimodal Pretraining

    arXiv:2607.00293v1 Announce Type: cross Abstract: Achieving true artificial general intelligence requires foundation models capable of integrating new modalities without forgetting prior knowledge. However, accommodating continuous generative objectives alongside discrete underst…

  6. arXiv cs.LG TIER_1 English(EN) · Sanghyuk Chun, Olga Russakovsky ·

    Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning

    arXiv:2505.19614v2 Announce Type: replace Abstract: Multimodal learning has seen remarkable progress, particularly with large-scale pre-training across various modalities. Most current approaches are built on the assumption of a deterministic one-to-one alignment between modaliti…

  7. arXiv cs.CL TIER_1 Italiano(IT) · Ping Tan ·

    Rosetta: Composable Native Multimodal Pretraining

    Achieving true artificial general intelligence requires foundation models capable of integrating new modalities without forgetting prior knowledge. However, accommodating continuous generative objectives alongside discrete understanding tasks causes severe gradient conflicts. Exi…

  8. Hugging Face Daily Papers TIER_1 English(EN) ·

    Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning

    Asymmetric Mutual Variational Learning addresses train-inference mismatch in multimodal reasoning by using bidirectional calibration to prevent answer leakage and improve latent-space stability.

  9. arXiv cs.AI TIER_1 English(EN) · Ziqi Zhou, Weize Quan, Mining Tan, Zhihan Chen, Dandan Zheng, Jingdong Chen, Jun Zhou, Weiming Dong, Dong-Ming Yan ·

    COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models

    arXiv:2606.28696v1 Announce Type: new Abstract: Composition is a high-level visual intent that governs where subjects are placed and how a scene is organized, yet current unified multimodal models remain unreliable at fine-grained composition recognition and struggle to turn such…

  10. arXiv cs.CL TIER_1 English(EN) · Weiyang Jin, Yuwei Niu, Jiaqi Liao, Chengqi Duan, Aoxue Li, Shenghua Gao, Xihui Liu ·

    SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

    arXiv:2510.12784v2 Announce Type: replace-cross Abstract: Recently, remarkable progress has been made in Unified Multimodal Models (UMMs), which integrate vision-language generation and understanding capabilities within a single framework. However, a model's strong visual underst…

  11. Hugging Face Daily Papers TIER_1 English(EN) ·

    Residual-Guided Expert Specialization for Incomplete Multimodal Learning

    As real-world prediction systems often face missing modalities at inference, incomplete multimodal learning (IML) remains a practical challenge. While prior methods aim to learn representations robust to missing inputs, representations from incomplete modalities inevitably deviat…

  12. Hugging Face Daily Papers TIER_1 English(EN) ·

    PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

    PerceptionRubrics presents a rubric-based evaluation framework that identifies gaps between benchmark scores and real-world performance through atomic auditing and gated scoring mechanisms.

  13. arXiv cs.CV TIER_1 English(EN) · Bingchen Huang, Zhiling Wang, Yifu Chen, Yuanchao Du ·

    Retrieved Images as Visual Thought: Training-Free Multimodal In-Context Learning for the Open-vs-Closed Gap

    arXiv:2607.00606v1 Announce Type: new Abstract: Recent work on Thinking with Images makes vision a dynamic part of reasoning, but does so through generation: the model invokes external tools, synthesizes code, or imagines new imagery, each at the cost of a tool protocol, brittle …

  14. arXiv cs.CV TIER_1 English(EN) · Shijie Li, Yilin Gao, Siyuan Yang, Tieyuan Chen, Chaofan Gan, Zhihao He, Zicheng Zhao, Yuyu Guo, Weiyao Lin, Hang Yu ·

    Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning

    arXiv:2607.00461v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are often constrained by a language-space bottleneck, forcing complex visual reasoning into discrete tokens which can lose perceptual nuance. A promising alternative is continuous latent reas…

  15. arXiv cs.CV TIER_1 English(EN) · Yunhan Wang, Eshika Khandelwal, Edson Araujo, Walid Bousselham, Nina Shvetsova, Hilde Kuehne ·

    Decompose, Compare, and Decide: Multimodal LLMs are Implicit Few-Shot Learners

    arXiv:2607.00125v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable abilities when analyzing images, yet translating these capabilities to few-shot image classification remains challenging. To bridge this gap, we present DeCoDe, a…

  16. arXiv cs.CV TIER_1 English(EN) · Yuanchao Du ·

    Retrieved Images as Visual Thought: Training-Free Multimodal In-Context Learning for the Open-vs-Closed Gap

    Recent work on Thinking with Images makes vision a dynamic part of reasoning, but does so through generation: the model invokes external tools, synthesizes code, or imagines new imagery, each at the cost of a tool protocol, brittle code, or an expensive training pipeline. A fourt…