PulseAugur
实时 22:31:22
English(EN) CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

新数据集和框架推动图像和视频编辑能力发展

研究人员开发了多个新的数据集和框架,以推动图像和视频编辑能力的进步。OpenGPT-4o-Image 是一个大规模数据集,采用新颖的方法,结合分层任务分类和自动生成,创建指令-图像对,以提高多模态 AI 性能。CPI-Bench 为真实世界的图像编辑提供了一个全面的基准,评估多图像任务、实际应用和基于推理的编辑,以更好地区分模型性能。SI-Edit 引入了一个用于草图指令引导的局部图像编辑的数据集和框架,具有像素级精度,而 Concept Distillation Sampling (CDS) 通过利用 LoRA 适配器提供了一个无需训练的多概念图像编辑框架。此外,Edit2Restore 等新方法通过调整预训练的编辑模型来演示高效的少样本图像恢复,Qwen-Video-Edit 则通过在 VAE 潜在空间上操作来重新利用图像编辑模型进行视频编辑。 AI

影响 这些在数据集和框架方面的进步有望加速图像和视频编辑任务的 AI 模型开发并提高其性能。

排序理由 多篇研究论文介绍了用于图像和视频编辑的新数据集、基准和框架。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 20 个来源。 我们如何撰写摘要 →

新数据集和框架推动图像和视频编辑能力发展

报道来源 [20]

  1. arXiv cs.AI TIER_1 English(EN) · Haoxiang Cao, Jiajiong Cao, Xuanpu Zhang, Changqian Yu, Chaoqun Wang ·

    DARS:基于指令的图像编辑的双层信用分配强化学习与结构化推理

    arXiv:2608.20161v1 Announce Type: new Abstract: Instruction-based image editing uses a planner-renderer pipeline: a vision-language model (VLM) first converts the instruction into an edit plan, and a diffusion model then executes that plan. Training such systems with only final-i…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    CoinVE-200K:用于组合式指令引导视频编辑的大规模高质量数据集

    The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editing. In particular, multiple editing intents must …

  3. arXiv cs.AI TIER_1 English(EN) · Zhihong Chen, Xuehai Bai, Yang Shi, Chaoyou Fu, Huanyu Zhang, Haotian Wang, Xiaoyan Sun, Zhang Zhang, Liang Wang, Yuanxing Zhang, Pengfei Wan, Yi-Fan Zhang ·

    OpenGPT-4o-Image:用于高级图像生成和编辑的综合数据集

    arXiv:2509.24900v2 Announce Type: replace-cross Abstract: The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like …

  4. Hugging Face Daily Papers TIER_1 English(EN) ·

    EDITBRIDGE:迈向忠实且高效的超高分辨率图像编辑

    EditBridge enables efficient ultra high-resolution image editing via a diffusion bridge that translates low-resolution edits to high-resolution outputs while preserving source details through sparse attention.

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    CoinVE-200K:用于组合式指令引导视频编辑的大规模高质量数据集

    A new dataset, benchmark, and 22B model enable compositional instruction-guided video editing with multi-region attention and temporal coherence.

  6. Hugging Face Daily Papers TIER_1 English(EN) ·

    GRNEdit:生成式精炼网络中基于新二元证据视角的高效通用视频编辑

    GRNEdit is a lightweight two-stage framework that models video editing intent via binary semantic decisions and source evidence, achieving strong results with minimal parameters.

  7. Hugging Face Daily Papers TIER_1 English(EN) ·

    CPI-Bench:一个全面、实用且智能的真实世界图像编辑基准测试

    CPI-Bench is a comprehensive benchmark for real-world image editing that evaluates multi-image tasks, practical applications, and reasoning-based editing to better differentiate model performance.

  8. Hugging Face Daily Papers TIER_1 English(EN) ·

    SI-Edit:迈向像素级精度的草图-指令引导的局部图像编辑

    Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly for fine-grained local deformations. This gap stems primarily from the critical shortage of high-quality, publicly available ben…

  9. arXiv cs.CV TIER_1 English(EN) · Jiayi Song, Shijie Huang, Fangtai Wu, Yubo Huang, Zhenxiong Tan, Songhua Liu, Jiaming Liu, Ruihua Huang ·

    EDITBRIDGE:迈向忠实且高效的超高分辨率图像编辑

    arXiv:2608.18063v1 Announce Type: new Abstract: High-resolution image editing is increasingly demanded in professional workflows, yet existing diffusion-based models remain constrained to resolutions below 1K due to quadratic attention complexity and prohibitive memory requiremen…

  10. arXiv cs.CV TIER_1 English(EN) · Fuchen Long, Cong Wang, Zitao Gao, Wenhao Zhong, Yu Cheng, Xiaolu Hou, Yan Li, Xiao Cao, Xinlong Sun, Xi Chen, Yu Liu ·

    CoinVE-200K:用于组合式指令引导视频编辑的大规模高质量数据集

    arXiv:2608.17566v1 Announce Type: new Abstract: The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editin…

  11. arXiv cs.CV TIER_1 English(EN) · Kunyu Feng, Yue Ma, Bingyuan Wang, Yuefeng Wang, Zhiyuan Qin, Hao Cheng, Hao Li, Qifeng Chen, Zeyu Wang ·

    MSEditor:迈向一致的多镜头视频编辑

    arXiv:2608.17559v1 Announce Type: new Abstract: In this paper, we tackle the problem of performing consistent, unified modifications to a multi-shot video sequence. This task is particularly challenging because multi-shot videos consist of discontinuous temporal segments that var…

  12. arXiv cs.CV TIER_1 English(EN) · Long Cui, Xiaoqian Liu, Qi Qin, Yi Xin, Tao Lin, Jianguo Li, Linfeng Zhang ·

    通过概念缩放和密集监督解锁图像编辑的潜力

    arXiv:2608.16812v1 Announce Type: new Abstract: Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. However, extending this paradigm to image editing highlights two inherent discrepancies, specifically, the insufficient …

  13. arXiv cs.CV TIER_1 English(EN) · Zhefan Rao, Bin Zou, Haoxuan Che, Xuanhua He, Chong Hou Choi, Yanheng Li, Rui Liu, Qifeng Chen ·

    从密集预测到视觉编辑:用于统一图像和视频创作的结构化监督

    arXiv:2608.14740v1 Announce Type: new Abstract: Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-only conditioning and creation-only training do not …

  14. arXiv cs.CV TIER_1 English(EN) · Yunpeng Bai, Yossi Gandelsman, Micha\"el Gharbi, Qixing Huang ·

    Qwen-Video-Edit:通过重用图像编辑模型实现基于指令的视频编辑

    arXiv:2608.14790v1 Announce Type: new Abstract: Instruction-based video editing is commonly built on video-pretrained generative backbones: a video diffusion transformer is adapted, at considerable cost, to condition on a source video and an editing instruction. In this report we…

  15. arXiv cs.CV TIER_1 English(EN) · Feng Xie, Jiagao Hu, Fuhao Li, Zepeng Wang, Yuxuan Chen, Dahua Gao, Fei Wang, Daiguo Zhou ·

    GRNEdit:生成式精炼网络中基于新二元证据视角的高效通用视频编辑

    arXiv:2608.16328v1 Announce Type: new Abstract: Instruction-based general video editing seeks to unify diverse editing operations within a single, intuitive interface. Existing approaches often rely on resource-intensive conditioning, using either heavyweight branches or costly s…

  16. arXiv cs.CV TIER_1 English(EN) · Niki Foteinopoulou, Ignas Budvytis, Stephan Liwicki ·

    无训练多概念图像编辑

    arXiv:2602.20839v3 Announce Type: replace Abstract: Training-free image editing with diffusion models is highly desirable yet is complex and remains a significant challenge. While recent optimisation-based methods achieve strong zero-shot edits from text, they still struggle to p…

  17. arXiv cs.CV TIER_1 English(EN) · Mustafa Ak{\i}n Y{\i}lmaz, Ahmet Bilican, Burak Can Biner, Ahmet Murat Tekalp ·

    Edit2Restore:通过预训练编辑模型的参数高效适应实现少样本图像修复

    arXiv:2601.03391v3 Announce Type: replace-cross Abstract: Image restoration has traditionally required training specialized models on thousands of paired examples per degradation type. Large pre-trained text-conditioned image editing models encode rich priors about image structur…

  18. arXiv cs.CV TIER_1 English(EN) · Qinye Zhou, Jun Zheng, Yongchao Du, Yuan Wang, Zhengrui Chen, Zuan Gao, Taihang Hu, Chao Lin, Yefeng Shen, Xingjian Wang, Zhao Wang, Zhengtao Wu, Xiaoli Xu, Zhengze Xu, Hao Yan, Denghui Yang, Yuhang Yu, Huayu Zhang, Mingzhou Zhang, Mengting Chen ·

    CPI-Bench:一个全面、实用且智能的真实世界图像编辑基准测试

    arXiv:2608.14546v1 Announce Type: new Abstract: With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existin…

  19. r/StableDiffusion TIER_2 English(EN) · /u/RobbaW ·

    MiniMax H3 作为多参考图像编辑器 + 一个评价你生成内容的 Tamagotchi

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1vppl41/minimax_h3_as_a_multiref_image_editor_a/"> <img alt="MiniMax H3 as a multi-ref image editor + a Tamagotchi judging your generations" src="https://external-preview.redd.it/ZmQyY2lwM2hsb2poMRyHKKKtU…

  20. r/StableDiffusion TIER_2 English(EN) · /u/Patient_Ratio4177 ·

    H3 作为单一图像编辑模型

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1vo1ab3/h3_as_a_singleimage_edit_model/"> <img alt="H3 as a single-image edit model" src="https://preview.redd.it/zaljlfnztajh1.png?width=140&amp;height=104&amp;auto=webp&amp;s=5c50778c2cb92e2b382c8ce2215…