PulseAugur
EN
LIVE 18:48:39

HOMIE framework enhances video personalization with MLLM integration

Researchers have developed HOMIE, a new framework for human-object centric video personalization (HOCVP). This method aims to improve subject fidelity and interaction accuracy in videos, even with abstract concepts like logos. HOMIE integrates multimodal large language models (MLLMs) to better understand relationships between subjects and objects, and it also incorporates a modality-reference embedding to distinguish between MLLM features and other visual tokens. The framework is designed to handle both inter-subject and intra-subject personalization scenarios effectively. AI

IMPACT This research could lead to more sophisticated and accurate subject-driven video generation, improving realism and control in personalized video content.

RANK_REASON The cluster describes a new research paper detailing a novel framework for video personalization.

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

HOMIE framework enhances video personalization with MLLM integration

COVERAGE [3]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

    Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high su…

  2. arXiv cs.CV TIER_1 English(EN) · Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, Cheng Chen, Wen Zhou, Zhenbang Sun, Wei Xue, Wenhan Luo, Yike Guo ·

    HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

    arXiv:2607.18217v1 Announce Type: new Abstract: Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization st…

  3. r/StableDiffusion TIER_2 English(EN) · /u/switch2stock ·

    HOMIE - Human-object centric video personalization | Qwen3-VL-2B + Wan2.1 | R2V

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1v2h68v/homie_humanobject_centric_video_personalization/"> <img alt="HOMIE - Human-object centric video personalization | Qwen3-VL-2B + Wan2.1 | R2V" src="https://external-preview.redd.it/anA0dWVwbXl1a2Vo…