arXiv:2608.20161v1 Announce Type: new Abstract: Instruction-based image editing uses a planner-renderer pipeline: a vision-language model (VLM) first converts the instruction into an edit plan, and a diffusion model then executes that plan. Training such systems with only final-i…
The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editing. In particular, multiple editing intents must …
arXiv:2509.24900v2 Announce Type: replace-cross Abstract: The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like …
EditBridge enables efficient ultra high-resolution image editing via a diffusion bridge that translates low-resolution edits to high-resolution outputs while preserving source details through sparse attention.
GRNEdit is a lightweight two-stage framework that models video editing intent via binary semantic decisions and source evidence, achieving strong results with minimal parameters.
CPI-Bench is a comprehensive benchmark for real-world image editing that evaluates multi-image tasks, practical applications, and reasoning-based editing to better differentiate model performance.
Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly for fine-grained local deformations. This gap stems primarily from the critical shortage of high-quality, publicly available ben…
arXiv:2608.18063v1 Announce Type: new Abstract: High-resolution image editing is increasingly demanded in professional workflows, yet existing diffusion-based models remain constrained to resolutions below 1K due to quadratic attention complexity and prohibitive memory requiremen…
arXiv:2608.17566v1 Announce Type: new Abstract: The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editin…
arXiv:2608.17559v1 Announce Type: new Abstract: In this paper, we tackle the problem of performing consistent, unified modifications to a multi-shot video sequence. This task is particularly challenging because multi-shot videos consist of discontinuous temporal segments that var…
arXiv cs.CV
TIER_1English(EN)·Long Cui, Xiaoqian Liu, Qi Qin, Yi Xin, Tao Lin, Jianguo Li, Linfeng Zhang·
arXiv:2608.16812v1 Announce Type: new Abstract: Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. However, extending this paradigm to image editing highlights two inherent discrepancies, specifically, the insufficient …
arXiv cs.CV
TIER_1English(EN)·Zhefan Rao, Bin Zou, Haoxuan Che, Xuanhua He, Chong Hou Choi, Yanheng Li, Rui Liu, Qifeng Chen·
arXiv:2608.14740v1 Announce Type: new Abstract: Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-only conditioning and creation-only training do not …
arXiv:2608.14790v1 Announce Type: new Abstract: Instruction-based video editing is commonly built on video-pretrained generative backbones: a video diffusion transformer is adapted, at considerable cost, to condition on a source video and an editing instruction. In this report we…
arXiv:2608.16328v1 Announce Type: new Abstract: Instruction-based general video editing seeks to unify diverse editing operations within a single, intuitive interface. Existing approaches often rely on resource-intensive conditioning, using either heavyweight branches or costly s…
arXiv cs.CV
TIER_1English(EN)·Niki Foteinopoulou, Ignas Budvytis, Stephan Liwicki·
arXiv:2602.20839v3 Announce Type: replace Abstract: Training-free image editing with diffusion models is highly desirable yet is complex and remains a significant challenge. While recent optimisation-based methods achieve strong zero-shot edits from text, they still struggle to p…
arXiv cs.CV
TIER_1English(EN)·Mustafa Ak{\i}n Y{\i}lmaz, Ahmet Bilican, Burak Can Biner, Ahmet Murat Tekalp·
arXiv:2601.03391v3 Announce Type: replace-cross Abstract: Image restoration has traditionally required training specialized models on thousands of paired examples per degradation type. Large pre-trained text-conditioned image editing models encode rich priors about image structur…
arXiv:2608.14546v1 Announce Type: new Abstract: With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existin…
<table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1vppl41/minimax_h3_as_a_multiref_image_editor_a/"> <img alt="MiniMax H3 as a multi-ref image editor + a Tamagotchi judging your generations" src="https://external-preview.redd.it/ZmQyY2lwM2hsb2poMRyHKKKtU…