PulseAugur
EN
LIVE 22:28:14

New datasets and frameworks advance image and video editing capabilities

Researchers have developed several new datasets and frameworks to advance image and video editing capabilities. OpenGPT-4o-Image, a large-scale dataset, uses a novel methodology with hierarchical task taxonomy and automated generation to create instruction-image pairs for improved multimodal AI performance. CPI-Bench offers a comprehensive benchmark for real-world image editing, evaluating multi-image tasks, practical applications, and reasoning-based editing to better differentiate model performance. SI-Edit introduces a dataset and framework for sketch-instruction guided local image editing with pixel-level precision, while Concept Distillation Sampling (CDS) provides a training-free framework for multi-concept image editing by leveraging LoRA adapters. Additionally, new methods like Edit2Restore demonstrate efficient few-shot image restoration by adapting pre-trained editing models, and Qwen-Video-Edit repurposes an image editing model for video editing by operating on VAE latents. AI

IMPACT These advancements in datasets and frameworks are expected to accelerate the development and improve the performance of AI models for image and video editing tasks.

RANK_REASON Multiple research papers introducing new datasets, benchmarks, and frameworks for image and video editing.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 20 sources. How we write summaries →

New datasets and frameworks advance image and video editing capabilities

COVERAGE [20]

  1. arXiv cs.AI TIER_1 English(EN) · Haoxiang Cao, Jiajiong Cao, Xuanpu Zhang, Changqian Yu, Chaoqun Wang ·

    DARS: Dual-Level Credit Assignment RL with Structured Reasoning for Instruction-Based Image Editing

    arXiv:2608.20161v1 Announce Type: new Abstract: Instruction-based image editing uses a planner-renderer pipeline: a vision-language model (VLM) first converts the instruction into an edit plan, and a diffusion model then executes that plan. Training such systems with only final-i…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

    The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editing. In particular, multiple editing intents must …

  3. arXiv cs.AI TIER_1 English(EN) · Zhihong Chen, Xuehai Bai, Yang Shi, Chaoyou Fu, Huanyu Zhang, Haotian Wang, Xiaoyan Sun, Zhang Zhang, Liang Wang, Yuanxing Zhang, Pengfei Wan, Yi-Fan Zhang ·

    OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing

    arXiv:2509.24900v2 Announce Type: replace-cross Abstract: The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like …

  4. Hugging Face Daily Papers TIER_1 English(EN) ·

    EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing

    EditBridge enables efficient ultra high-resolution image editing via a diffusion bridge that translates low-resolution edits to high-resolution outputs while preserving source details through sparse attention.

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

    A new dataset, benchmark, and 22B model enable compositional instruction-guided video editing with multi-region attention and temporal coherence.

  6. Hugging Face Daily Papers TIER_1 English(EN) ·

    GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks

    GRNEdit is a lightweight two-stage framework that models video editing intent via binary semantic decisions and source evidence, achieving strong results with minimal parameters.

  7. Hugging Face Daily Papers TIER_1 English(EN) ·

    CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing

    CPI-Bench is a comprehensive benchmark for real-world image editing that evaluates multi-image tasks, practical applications, and reasoning-based editing to better differentiate model performance.

  8. Hugging Face Daily Papers TIER_1 English(EN) ·

    SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision

    Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly for fine-grained local deformations. This gap stems primarily from the critical shortage of high-quality, publicly available ben…

  9. arXiv cs.CV TIER_1 English(EN) · Jiayi Song, Shijie Huang, Fangtai Wu, Yubo Huang, Zhenxiong Tan, Songhua Liu, Jiaming Liu, Ruihua Huang ·

    EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing

    arXiv:2608.18063v1 Announce Type: new Abstract: High-resolution image editing is increasingly demanded in professional workflows, yet existing diffusion-based models remain constrained to resolutions below 1K due to quadratic attention complexity and prohibitive memory requiremen…

  10. arXiv cs.CV TIER_1 English(EN) · Fuchen Long, Cong Wang, Zitao Gao, Wenhao Zhong, Yu Cheng, Xiaolu Hou, Yan Li, Xiao Cao, Xinlong Sun, Xi Chen, Yu Liu ·

    CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

    arXiv:2608.17566v1 Announce Type: new Abstract: The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editin…

  11. arXiv cs.CV TIER_1 English(EN) · Kunyu Feng, Yue Ma, Bingyuan Wang, Yuefeng Wang, Zhiyuan Qin, Hao Cheng, Hao Li, Qifeng Chen, Zeyu Wang ·

    MSEditor: Toward Consistent Multi-Shot Video Editing

    arXiv:2608.17559v1 Announce Type: new Abstract: In this paper, we tackle the problem of performing consistent, unified modifications to a multi-shot video sequence. This task is particularly challenging because multi-shot videos consist of discontinuous temporal segments that var…

  12. arXiv cs.CV TIER_1 English(EN) · Long Cui, Xiaoqian Liu, Qi Qin, Yi Xin, Tao Lin, Jianguo Li, Linfeng Zhang ·

    Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision

    arXiv:2608.16812v1 Announce Type: new Abstract: Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. However, extending this paradigm to image editing highlights two inherent discrepancies, specifically, the insufficient …

  13. arXiv cs.CV TIER_1 English(EN) · Zhefan Rao, Bin Zou, Haoxuan Che, Xuanhua He, Chong Hou Choi, Yanheng Li, Rui Liu, Qifeng Chen ·

    From Dense Prediction to Visual Editing: Structured Supervision for Unified Image and Video Creation

    arXiv:2608.14740v1 Announce Type: new Abstract: Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-only conditioning and creation-only training do not …

  14. arXiv cs.CV TIER_1 English(EN) · Yunpeng Bai, Yossi Gandelsman, Micha\"el Gharbi, Qixing Huang ·

    Qwen-Video-Edit: Instruction-Based Video Editing by Repurposing an Image Editing Model

    arXiv:2608.14790v1 Announce Type: new Abstract: Instruction-based video editing is commonly built on video-pretrained generative backbones: a video diffusion transformer is adapted, at considerable cost, to condition on a source video and an editing instruction. In this report we…

  15. arXiv cs.CV TIER_1 English(EN) · Feng Xie, Jiagao Hu, Fuhao Li, Zepeng Wang, Yuxuan Chen, Dahua Gao, Fei Wang, Daiguo Zhou ·

    GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks

    arXiv:2608.16328v1 Announce Type: new Abstract: Instruction-based general video editing seeks to unify diverse editing operations within a single, intuitive interface. Existing approaches often rely on resource-intensive conditioning, using either heavyweight branches or costly s…

  16. arXiv cs.CV TIER_1 English(EN) · Niki Foteinopoulou, Ignas Budvytis, Stephan Liwicki ·

    Training-Free Multi-Concept Image Editing

    arXiv:2602.20839v3 Announce Type: replace Abstract: Training-free image editing with diffusion models is highly desirable yet is complex and remains a significant challenge. While recent optimisation-based methods achieve strong zero-shot edits from text, they still struggle to p…

  17. arXiv cs.CV TIER_1 English(EN) · Mustafa Ak{\i}n Y{\i}lmaz, Ahmet Bilican, Burak Can Biner, Ahmet Murat Tekalp ·

    Edit2Restore:Few-Shot Image Restoration via Parameter-Efficient Adaptation of Pre-trained Editing Models

    arXiv:2601.03391v3 Announce Type: replace-cross Abstract: Image restoration has traditionally required training specialized models on thousands of paired examples per degradation type. Large pre-trained text-conditioned image editing models encode rich priors about image structur…

  18. arXiv cs.CV TIER_1 English(EN) · Qinye Zhou, Jun Zheng, Yongchao Du, Yuan Wang, Zhengrui Chen, Zuan Gao, Taihang Hu, Chao Lin, Yefeng Shen, Xingjian Wang, Zhao Wang, Zhengtao Wu, Xiaoli Xu, Zhengze Xu, Hao Yan, Denghui Yang, Yuhang Yu, Huayu Zhang, Mingzhou Zhang, Mengting Chen ·

    CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing

    arXiv:2608.14546v1 Announce Type: new Abstract: With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existin…

  19. r/StableDiffusion TIER_2 English(EN) · /u/RobbaW ·

    MiniMax H3 as a multi-ref image editor + a Tamagotchi judging your generations

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1vppl41/minimax_h3_as_a_multiref_image_editor_a/"> <img alt="MiniMax H3 as a multi-ref image editor + a Tamagotchi judging your generations" src="https://external-preview.redd.it/ZmQyY2lwM2hsb2poMRyHKKKtU…

  20. r/StableDiffusion TIER_2 English(EN) · /u/Patient_Ratio4177 ·

    H3 as a single-image edit model

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1vo1ab3/h3_as_a_singleimage_edit_model/"> <img alt="H3 as a single-image edit model" src="https://preview.redd.it/zaljlfnztajh1.png?width=140&amp;height=104&amp;auto=webp&amp;s=5c50778c2cb92e2b382c8ce2215…