Researchers have developed ConFusion, a new framework for fine-grained controllable infrared and visible image fusion. This method learns a continuous fusion space using Gaussian-conditioned spatial-aware modulation, allowing for flexible integration of thermal and structural information. ConFusion utilizes a dual-branch architecture for disentangling representations and employs a multimodal large language model to interpret user intents into modulation variables for precise image fusion. Experiments indicate that ConFusion surpasses existing methods in both fusion quality and downstream task performance. AI
IMPACT Enables more precise and adaptable image fusion for downstream applications by leveraging LLMs for control.
RANK_REASON The item describes a new research paper detailing a novel framework for image fusion. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
- ConFusion
- Gaussian-conditioned spatial-aware modulation
- Mask-Guided Specific Feature Modulator
- multimodal large language model
- Text-Driven Invariant Feature Enhancer
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →