arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HOMIE:通过多模态智能增强实现以人为对象的视频个性化

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, Cheng Chen, Wen Zhou, Zhenbang Sun, Wei Xue, Wenhan Luo, Yike Guo

arXiv 2607.18217首次发表:更新:

发表机构

Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究以人为对象的视频个性化问题,提出HOMIE框架,统一处理主体间和主体内输入设置。通过更好的MLLM集成策略、自注意力中的全局多模态引导及模态参考嵌入,在多任务中达最优性能。

AI 中文摘要

以人为对象的视频个性化(HOCVP)是主题驱动视频生成中的核心任务。现有方法存在两个关键局限。多数关注主体间个性化的方法难以在高主体保真度与人和多样物体间准确交互模式间平衡,尤其物体为抽象概念如图标时。虽主体内参考有望增强保真度,但多数现有工作缺乏理解潜在对应关系的机制。为应对这些挑战,我们提出HOMIE框架,统一处理主体间和主体内输入设置。与先前方法相比,HOMIE提出更好的MLLM集成策略,在不影响文本编码器可控性或产生高成本重新对齐的情况下提取参考级关系知识。具体而言,我们在自注意力中引入全局多模态引导,更好地对齐MLLM语义特征与VAE令牌;还提出模态参考嵌入来区分MLLM特征和VAE令牌并关联主体内参考图像令牌。大量实验验证了该方法在各种HOCVP任务中达到了当前最优性能。

英文摘要

Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high subject fidelity and accurate interaction patterns between humans and diverse objects, especially when objects represent abstract concepts such as logos. Second, while intra-subject references (e.g., OCR maps, multi-view inputs) are expected to enhance subject fidelity, most existing works lack mechanisms to understand such latent correspondence. To address both challenges, we propose HOMIE, an HOCVP framework that tackles both inter- and intra-subject input settings in a unified manner. Compared to previous approaches, HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment. Specifically, we introduce global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens. Furthermore, we propose modality-reference embedding to differentiate tokens from MLLM features and VAE tokens and associate intra-subject reference image tokens. Extensive experiments validate that our method achieves state-of-the-art performance across various HOCVP tasks. Project Page: https://yiyangcai.github.io/homie-page.github.io/

Comments28 pages, 14 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑