Researchers have introduced Goku, a large-scale dataset and benchmark for instruction-based video editing, designed to overcome the limitations of existing datasets that focus on single-task appearance edits. Goku comprises 2 million high-quality video editing pairs, enabling multi-task and structural manipulations like precise subject movement control. The accompanying Goku-Edit model utilizes a multimodal large language model for instruction comprehension and a dual-branch design for structural and appearance editing. A benchmark, Goku-Bench, with 1,000 human-verified cases and 7 new metrics, was also released, showing Goku-Edit achieving up to an 8% improvement in instruction following over other open-source models. AI
IMPACT Advances capabilities in instruction-based video editing, potentially enabling more complex and creative user-driven video manipulations.
RANK_REASON The cluster describes a new dataset, benchmark, and model for video editing, published as a research paper.
Read on Hugging Face Daily Papers →
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Goku
- Goku-Bench
- Goku-Edit
- Gotit.pub
- Hugging Face
- multimodal large language model
- ScienceCast
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →