TaoLive数字虚拟人智能体技术报告:训练智能体随其Harness进化
Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report
浏览论文内容
中文总结 AI 辅助
本研究针对直播电商数字虚拟人主播的实时需求,提出HAT方法,结合HSA的三阶段训练,使35B紧凑模型在低延迟下适配Harness变化,性能优于基础模型及通用大语言模型。
中文摘要 AI 辅助
直播电商中基于AI的数字虚拟人主播需实时回答产品问题、与观众互动并执行不断变化的商业策略,这要求回复低延迟、准确有效,且能快速适配更新的活动、合规及风格要求。我们开发了可进化的Harness,将技能、互动钩子、系统提示和工具与模型权重解耦,使运行时行为无需重新训练即可改变。但Harness进化会形成动态执行环境:针对某一配置微调的紧凑模型可能会记住名称、模式和提示模板,而非遵循当前提供的Harness,而更强的零样本模型速度过慢,无法满足实时需求。我们通过Harness感知训练(HAT)解决这一矛盾,该方法将Harness状态纳入训练分布。HAT对技能、工具模式、提示结构和交互约束应用任务保留型Harness状态增强(HSA),包含三个阶段:基于HSA的监督微调、用于恢复通用能力的通用在线策略蒸馏,以及在生产环境驱动的直播间模拟器中进行的基于HSA的智能体强化学习。在包含4500余案例的四个评估集上,我们的35B紧凑模型在真实直播问答中得分为94.8,而基础模型为80.3,评估的最强通用大语言模型为93.0;在Harness变体问答中得分为94.6,在IFEval中保持83.5。相比之下,固定Harness的SFT使IFEval得分降低7.7个百分点。在启用MTP的单NVIDIA H20 GPU上进行的受控完整智能体重放中,系统实现了3.407秒的P50延迟和8.114秒的P95延迟。这些结果表明,HAT生成了延迟可行的紧凑智能体,在评估的Harness变化下仍保持有效,且未牺牲通用指令遵循能力。
英文摘要
AI-powered digital avatar streamers must answer product questions, engage viewers, and execute marketing strategies in real time, demanding low latency, frequent strategy updates, and accurate yet effective responses. Evolvable Harnesses, whose Skills, Hooks, prompts, and tools can be updated independently of model weights, enable rapid iteration but expose a trade-off: large models adapt zero-shot yet are too slow, whereas compact models meet latency targets but overfit to fixed Harness configurations. We propose Harness-Aware Training (HAT), which trains compact models to adapt to changing Harnesses. Its key component, Harness-State Augmentation (HSA), applies task-preserving transformations to Skill identifiers and content, tool schemas, prompt structures, and Hook functions. Training proceeds in three stages: HSA-SFT learns reasoning and tool use from strong-model trajectories across diverse environments; General On-Policy Distillation restores generalization lost during SFT; and HSA-RL improves robustness to changing Harnesses through reinforcement learning in augmented environments. Across four evaluation sets, HAT achieves 94.8 on Live-Stream QA (base: 80.3; strongest general LLM: 93.0) and 94.6 on Harness-Variant QA (base: 75.4). Unlike Fixed-Harness SFT, which lowers IFEval by 7.7 points from the base model, HAT avoids this regression and reaches 83.5. On one NVIDIA H20 GPU, the optimized system delivers P50 and P95 latencies of 3.4 s and 8.1 s. Deployed in Taobao Live's digital-avatar service, it also yields positive online A/B test results for item-page views.
发表机构
- TaoLive(陶境科技)
机构由 AI 辅助整理,请以论文原文为准。