发表机构
Institute of Trustworthy Embodied AI, Fudan University; ByteDance TikTok; Shanghai Innovation Institution; Zhejiang University(复旦大学可信具身智能研究院; 字节跳动TikTok; 上海创新研究院; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本工作提出Live Assistant框架,将直播协助形式化为是否、何时、向谁及沟通什么四个决策,通过自回归策略和两阶段训练,在真实直播数据上实现高准确率,建立选择性参与的协助范式。
AI 中文摘要
直播是持续时间较长的交互环境,其中音视频内容、观众活动、主播行为及平台信号共同演化,产生源自直播本身的协助需求。我们提出Live Assistant,一个用于混合主动、角色条件化协助的框架,将直播交互形式化为四个耦合决策:是否行动、何时行动、向谁发言以及沟通什么。在每个10秒间隔内,一个自回归策略消费原生音频和视频,以及同步的评论、礼物、观众动态和房间元数据,然后选择OBS、MEM或ANS。OBS保持沉默,MEM记录私有语义更新,ANS指定接收者、任务和基于事实的消息。为支持该任务,我们构建了一个轨迹引擎,将真实直播会话重建为结构化因果监督,产生超过320小时的优化轨迹和一个人工审核的基准,包含275个片段和13,812个决策间隔。我们使用标记感知的多轮监督微调(MA-MSFT)训练策略,该微调强化稀疏的结构化决策,随后使用流式多轮GSPO(SM-GSPO)优化自生成轨迹,并带有回合级和轨迹级信用。在保留基准上,Live Assistant达到71.14的状态准确率、72.67的接收者准确率和58.41的任务准确率,相对于代表性的流式和多模态基线有一致的提升。综合来看,该公式、基准和训练框架将直播协助确立为对共享社交流的选择性参与。
英文摘要
Livestreams are long-lasting interactive environments where audiovisual content, viewer activity, host behavior, and platform signals evolve together, creating assistance needs that emerge from the stream itself. We introduce \liveassistant, a framework for mixed-initiative, role-conditioned assistance that formulates livestream interaction as four coupled decisions: \textbf{whether to act, when to act, whom to address, and what to communicate}. At each 10-second interval, one autoregressive policy consumes native audio and video with synchronized comments, gifts, viewer dynamics, and room metadata, then selects \textsc{OBS}, \textsc{MEM}, or \textsc{ANS}. \textsc{OBS} remains silent, \textsc{MEM} records a private semantic update, and \textsc{ANS} specifies a recipient, task, and grounded message. To support this task, we build a trajectory engine that reconstructs real livestream sessions into structured causal supervision, yielding over 320 hours of optimization trajectories and a human-reviewed benchmark of 275 clips and 13,812 decision intervals. We train the policy with Marker-Aware Multiturn Supervised Fine-Tuning (MA-MSFT), which strengthens sparse structured decisions, followed by Streaming Multiturn GSPO (SM-GSPO), which optimizes self-generated trajectories with turn- and trajectory-level credit. On the held-out benchmark, \liveassistant reaches 71.14 state accuracy, 72.67 recipient accuracy, and 58.41 task accuracy, with consistent gains over representative streaming and general multimodal baselines. Together, the formulation, benchmark, and training framework establish livestream assistance as selective participation in a shared social stream.
Commentsunder review