发表机构
Kling AI Research; University of Chinese Academy of Sciences; National Cheng Kung University(克林人工智能研究院; 中国科学院大学; 国立成功大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有HOI视频生成方法不适用于实时交互应用的问题,提出StreamHOI框架。通过分析Transformer块的历史记忆偏好,进行离线分析与偏差引导训练,并引入内存距离缩放模块,实现高效长时HOI视频生成,兼具多种优势且效率高。
AI 中文摘要
现有的人类-物体交互(HOI)视频生成方法大多局限于具有复杂驱动条件的离线短视频生成,不适用于实时交互应用。我们提出了StreamHOI,一种用于长时HOI视频生成的低延迟流式框架。我们研究图像到视频的流式生成器应如何组织历史记忆以在有限延迟下保留交互。发现标准的汇聚点本地内存设计在流式HOI生成中面临权衡,不同的Transformer块对HOI区域和周围区域有不同的历史记忆偏好。为此,StreamHOI进行离线HOI感知块分析并应用偏差引导的内存专用训练,还引入内存距离缩放模块。与长视频基线和近期HOI生成方法的广泛比较表明,StreamHOI具有很强的交互合理性、物体逼真度、人类质量和效率,首块延迟0.75秒时可达17.6 FPS。
英文摘要
Existing human--object interaction (HOI) video generation methods are largely limited to offline short-video generation with complex driving conditions, making them unsuitable for real-time interactive applications. We present \emph{StreamHOI}, a low-latency streaming framework for long-duration HOI video generation. Instead of converting heavily conditioned HOI pipelines into streaming systems, we study how an image-to-video streaming generator should organize historical memory to preserve interactions under bounded latency. We find that the standard sink-local memory design faces a trade-off in streaming HOI generation, and different transformer blocks show different historical-memory preferences for HOI regions and surrounding regions. To match memory composition with block behavior, StreamHOI performs offline HOI-aware block profiling and applies bias-guided memory-specialized training to adapt the generator to block-specific memory layouts. We further introduce a memory distance scaling module to strengthen long-range access to early interaction states. Extensive comparisons with both long-video baselines and recent HOI generation methods demonstrate that StreamHOI achieves strong interaction plausibility, object fidelity, human quality and efficiency, reaching 17.6 FPS with 0.75s first-chunk latency.
CommentsCode and models are available at https://github.com/KlingAIResearch/StreamHOI