arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于陪伴智能体评估的披露门控用户模拟

Disclosure-Gated User Simulation for Companion-Agent Evaluation

Yao Liu, Yu He

arXiv 2609.00982首次发表:更新:

AI 中文总结

针对大语言模型模拟用户过于合作的缺陷,本文提出披露门控机制并训练出符合保序、尺度稳定标准的用户模拟器,其与基准原始模拟器相关性达0.993,补充了基准缺失的关键分析内容。

AI 中文摘要

使用大语言模型模拟用户现已成为可扩展评估的标准做法,但存在一个反复被发现的缺陷:模拟用户过于合作,导致被测系统可仅通过提问数量而非让用户愿意发言来获得高分。本文提出一种披露门控机制,将信息释放与陪伴智能体的行为挂钩,其状态为五级有序门,合并为三个可观测的深度层。我们对该机制进行了规格说明、消融实验和审计,并依据该规格训练用户模拟器:门控行为从训练语料库的合成分支中学习,真实分支则提供人类的说话与反应方式;训练后,模拟器在运行时无需被告知每个样本对应的门层级。该门是环境的核心组件:在已发布的陪伴智能体基准CompanionBench的英文语料上,当训练时不再告知每个样本的门层级,12个被测系统的最大排名位移超过了用新随机种子重新运行该环境设定的噪声带,但各系统的分数未出现可检测变化。我们设定两个验收标准:排名必须保序,绝对分数必须尺度稳定;在检验的候选方案中,仅我们发布的模拟器通过了两项标准,其排行榜与基准原始模拟器的相关系数达0.993。相比之下,用前沿模型作为模拟器仅略微调整排名,却将所有分数上移——这种变化仅检查排名的人无法察觉。本文指定的环境正是该基准已使用的环境,该基准仅用约400字描述了机制,而本文补充了其缺失的规格说明、消融实验、人类研究、阴性对照及下游敏感性分析。

英文摘要

Using a large language model to play the user is now standard in scalable evaluation. It has a repeatedly diagnosed failure: the simulated user is excessively cooperative, so a system under test can score by the sheer number of questions it asks rather than by making the user willing to speak. We answer with a disclosure gate conditioning information release on the companion agent's behaviour: its state is a ladder of five ordered gates, merged onto three observable depth layers. We specify, ablate, and audit it, and train a user simulator against that specification. Gating behaviour is learned from the training corpus's synthetic branch, while the real branch supplies how people speak and react; after training, the simulator need not be told at runtime which gate each item sits behind. The gate is a load-bearing component of the environment: on the English corpus of a published companion-agent benchmark (CompanionBench), once training no longer states per example which gate each item sits behind, the largest rank displacement across 12 systems under test exceeds the noise band set by re-running that environment under a new seed, while per-system scores show no detectable change. We state two acceptance criteria: a ranking must be order-preserving, and absolute scores must be scale-stable. Of the candidates we examine, only one passes both -- the simulator we release -- and its leaderboard correlates at 0.993 with the benchmark's original simulator. By contrast, prompting a frontier model as the simulator barely moves the ranking while shifting every score upward -- a shift invisible to anyone checking the ranking alone. The environment we specify is the one that benchmark already used. That publication describes the mechanism in about four hundred words, and we supply what it lacked: specification, ablations, human studies, negative controls, and downstream sensitivity analysis.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑