arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CUEing 用户模拟器:用于多轮基准测试的校准用户嵌入

CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking

Anjali Kantharuban, Jonas Mueller

arXiv 2610.02460首次发表:更新:

发表机构

Handshake AI; Carnegie Mellon University(Handshake AI; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出校准用户嵌入(CUE)框架,通过编码会话并采样连续表示生成人物命令,无需训练即可引导LLM模拟用户,提升多轮基准测试中结果校准与泛化能力。

AI 中文摘要

最近的基准测试依赖用户模拟器来评估AI智能体在多轮交互中的表现。虽然现有的模拟技术展现了与人类风格和行为在表面上的保真度,但生态有效的交互式基准测试还需要在模拟用户群体和真实用户群体中,智能体失败的时间和方式上保持一致。我们发现现有的模拟器缺乏结果校准:即与真实用户与同一智能体交互时观察到的成功率和失败模式的一致性。我们引入了校准用户嵌入(CUE),这是一个框架,既能编码观察到的会话,又能采样连续表示,然后将它们解码为人物命令,以引导LLM充当用户模拟器,无需训练。通过这种方式,我们评估了基于用户条件的过往会话重放,以及在为相同任务采样新人物时聚合指标的一致性。在τ²-Bench上,与其他基于人物的模拟方法相比,CUE模拟器犯下的模拟器归因错误更少,并且更忠实地再现了真实用户的智能体失败模式、聚合成功率以及特定任务-用户对的结果。这些增益与使用先前工作中建立的指标衡量的竞争性用户保真度并存。在主要适配于客户支持交互后,相同的CUE模拟器能够泛化到文档创建、数学辅导和随意对话,并且在不同模拟器LLM上无需CUE重训练即可保持有效。

英文摘要

Recent benchmarks rely on user simulators to evaluate AI agents in multi-turn interaction. While existing simulation techniques demonstrate surface fidelity to human style and behavior, ecologically valid interactive benchmarking also requires alignment in when and how agents fail across simulated and real user populations. We find that existing simulators lack outcome calibration: agreement with observed success rates and failure patterns when real users interact with the same agent. We introduce Calibrated User Embeddings (CUE), a framework that both encodes observed sessions and samples continuous representations, then decodes them into persona commands to steer LLMs to act as user simulators without training. Through this, we evaluate user-conditioned replay of past sessions and aggregate metric agreement when sampling novel personas for the same tasks. On $τ^2$-Bench, CUEd simulators commit fewer simulator-attributed errors and more faithfully reproduce real-user agent failure modes, aggregate success rates, and outcomes for specific task-user pairs than other persona-based simulation methods. These gains coexist with competitive user fidelity as measured using metrics established in prior work. After being fit to mostly customer support interactions, the same CUEd simulators generalize to document creation, math tutoring, and casual conversation, and remain effective across different simulator LLMs without CUE retraining.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑