arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35728cs.CV

FlowAct-R2:超越对话式虚拟人——通过流式多模态参考与主动式智能体规划

FlowAct-R2: Beyond Talking Avatar via Streaming Multimodal References and Proactive Agent Planning

Ziyao Huang, Zhengkun Rong, Shiyang Qin, Shuang Liang, Wentao Hu, Yuxuan Luo, Yuan Zhang, Mingyuan Gao

首次发表
浏览论文内容

中文总结 AI 辅助

FlowAct-R2通过流式多模态参考扩散Transformer和主动式交互智能体,实现实时720p、小时级流式交互式人形视频生成,支持娱乐直播、购物等场景。

中文摘要 AI 辅助

我们提出了FlowAct-R2,一个用于交互式人形视频生成的框架,它将连续多模态控制与主动式智能体规划相结合。我们的方法由两个耦合组件构成。首先,一个流式多模态参考扩散Transformer将预训练的Seedance 2.0 Mini参考到视频骨干网络适配为接受滚动动作提示、流式音频以及动态更新的图像、音频和视频参考。视频驱动的旋转位置嵌入将参考块与生成时间线对齐,而参考加图像条件化以及部分加噪的历史运动帧保留了外观并避免了累积漂移。其次,一个主动式交互智能体将在线前规划与在线调度和响应分离:它提前准备角色设定、长期议程和可复用的多模态技能,然后在直播会话中自主调度行为、响应观众输入并处理中断。FlowAct-R2支持实时720p生成和小时级流式传输,适用于娱乐直播、直播购物、视频聊天和实时视频博客。

英文摘要

We present FlowAct-R2, a framework for interactive humanoid video generation that combines continuous multimodal control with proactive agent planning. Our method consists of two coupled components. First, a Streaming Multimodal Reference Diffusion Transformer adapts the pretrained Seedance 2.0 Mini reference-to-video backbone to accept rolling action prompts, streaming audio, and dynamically updated image, audio, and video references. Video-driven rotary positional embeddings align reference chunks with the generation timeline, while reference-plus-image conditioning and partially noised historical motion frames preserve appearance and avoid accumulated drift. Second, a Proactive Interaction Agent separates pre-online planning from online scheduling and response: it prepares a persona, a long-horizon agenda, and reusable multimodal skills in advance, then autonomously schedules behaviors, responds to audience input, and handles interruptions during a live session. FlowAct-R2 supports real-time 720p generation and hour-scale streaming across entertainment streaming, live shopping, video chatting, and live vlogging.

发表机构

  • Bytedance Intelligent Creation(字节跳动智能创作)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑