发表机构
Carnegie Mellon University; University of Washington; Handshake AI; Stanford University; Princeton University; University of California San Diego(卡内基梅隆大学; 华盛顿大学; 汉德shake人工智能公司; 斯坦福大学; 普林斯顿大学; 加利福尼亚大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出TAHI方法,通过人机交互将跨会话信号整合到智能体中,针对30位用户的600项任务自适应后,单独任务成功率提升4.5-20.9%,还可跨用户提升最高8.8%成功率。
AI 中文摘要
AI智能体在大规模人口级数据上进行训练,以编码涵盖众多从业者能力的广泛功能,但它们生成的成果很少能达到专业人员需要用来维护自身声誉的个人标准。在成功标准异质性强且记录不足的现实开放式任务中,个体专业能力恰恰体现在对平均水平的提升与偏离上。实践中,人机的迭代交互会呈现出用户无法预先完全指定但可跨任务重复应用的标准。我们认为,这种跨会话交互数据是缩小与个体专业能力差距的丰富且未被充分利用的信号。本研究提出通过人机交互实现测试时自适应(TAHI),将这些信号整合到智能体的上下文和权重中,并通过不断演进的 rubric 模块明确每个用户的训练与评估标准。我们在写作和视觉创作这两个高实用性领域,针对30位个体的600项任务对智能体进行自适应,结果显示,智能体仅在数十项任务内就将单独任务的成功率提升了4.5%至20.9%。同时,我们的演进 rubric 模块作为可扩展的标注工具,生成的评估 rubric 比仅由语言模型(LM)或人类生成的 rubric 多捕捉到16.0%至22.3%的失败案例。此外,当智能体针对个体进行自适应时,这些个性化智能体还能在跨用户场景中实现最高达8.8%的成功率提升。
英文摘要
AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet the artifacts they produce rarely meet the personal bar professionals need to stake their reputation on. On realistic, open-ended tasks where success criteria are heterogeneous and insufficiently documented, individual expertise lives precisely in the elevation and departure from the average. In practice, iterative human-agent interaction surfaces criteria that users cannot fully specify up front, yet apply repeatedly across tasks. We argue this cross-session interaction data is a rich, underused signal for closing the gap to individual expertise. In this work, we propose test-time adaptation through human-agent interaction (TAHI), which integrates these signals into agent context and weights, and crystallizes each user's training and evaluation criteria via an evolving rubric module. We adapt agents to 30 individuals in two high-utility domains, writing and visual creation, on a total of 600 tasks. Our agents improve solo task success by 4.5-20.9% within only tens of tasks. Meanwhile, our evolving rubric module serves as a scalable annotation tool, creating evaluation rubrics that catch 16.0-22.3% more failures than those from LMs or humans alone. While agents are adapted towards individuals, we show these personalized agents also produce improvements in success of up to 8.8% that generalize across users.