PulseAugur
EN
LIVE 09:31:54

New framework ForgeryVCR improves MLLMs for image forgery detection

Researchers have introduced ForgeryVCR, a novel framework designed to enhance the accuracy of Multimodal Large Language Models (MLLMs) in detecting and localizing image forgeries. Unlike previous text-centric approaches that often lead to hallucinations due to the limitations of linguistic modalities in capturing fine-grained pixel inconsistencies, ForgeryVCR integrates an efficient forensic toolbox. This toolbox materializes imperceptible traces into explicit visual intermediates, enabling Visual-Centric Reasoning. The framework employs a Strategic Tool Learning paradigm, combining supervised fine-tuning with reinforcement learning, to empower MLLMs to proactively invoke multi-view reasoning paths for detailed inspection. AI

IMPACT Enhances MLLM capabilities in image forensics, potentially improving security and trust in visual media.

RANK_REASON The cluster contains a research paper detailing a new framework and methodology for image forgery detection using MLLMs. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New framework ForgeryVCR improves MLLMs for image forgery detection

COVERAGE [1]

  1. arXiv cs.CV TIER_1 English(EN) · Youqi Wang, Shen Chen, Haowei Wang, Rongxuan Peng, Taiping Yao, Shunquan Tan, Changsheng Chen, Bin Li, Shouhong Ding ·

    ForgeryVCR: Visual-Centric Reasoning via Efficient Forensic Tools in MLLMs for Image Forgery Detection and Localization

    arXiv:2602.14098v2 Announce Type: replace Abstract: Existing Multimodal Large Language Models (MLLMs) for image forgery detection and localization predominantly operate under a text-centric Chain-of-Thought (CoT) paradigm. However, forcing these models to textually characterize i…