arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI智能体的可靠性:评分者效应、漂移与评估项目的回归

Reliability of AI Agents: Rater Effects, Drift, and the Return to an Evaluation Program

Liu Zhang, Mark Esposito

arXiv 2610.07003首次发表:更新:

AI 中文总结

本研究利用生产环境面试数据,发现AI智能体评估可靠性受评审者效应和系统漂移影响,提出纵向评估需确保跨评审者可比性、证据时效性及运营结果印证。

AI 中文摘要

企业越来越多地使用重复人工评分来评估已部署的AI智能体,然而观察到的分数变化可能既反映了测量过程本身的变化,也反映了智能体自身的变化。企业也鲜少知道,随着已部署系统的演进,一次评估在多久内仍具有信息价值。我们利用一个已部署的语音与视频面试智能体中的2,611份基于评分标准的生产环境面试记录,以及重叠的人工评审者、智能体的变更日志和支持工单,将可靠性作为一个潜在状态进行研究。评审者的严格程度差异显著:评估相同批次的两位评审者在综合指数上的评分相差0.79个标准差,且评审者构成的变化使原始趋势变得平坦。在调整评审者效应并平滑批次级估计后,评估可靠性从2026年3月至8月增加了0.53个标准差。可靠性在记录的部署之间也发生了实质性变化。按估计的未计入变动速率,预测不确定性在大约五周后达到最大观测部署对比的幅度,尽管该时间范围估计不精确且基于少量阶跃式变化。一个独立的运营结果向同一方向变动:与面试相关的支持工单相对于技术问题工单每月下降约10%。这些数据并未识别出评估对绩效的因果效应。相反,它们表明纵向AI评估需要三项检查:跨评审者的可比性、漂移下的证据时效性,以及与重要运营结果的相互印证。

英文摘要

Firms increasingly evaluate deployed AI agents using repeated human ratings, yet observed score changes may reflect the measurement process as much as changes in the agent itself. Firms also rarely know how long an evaluation remains informative as deployed systems evolve. We study reliability as a latent state using 2,611 rubric-scored production interviews from a deployed voice-and-video interviewing agent, overlapping human reviewers, the agent's change log, and support tickets. Reviewer severity varied substantially: two reviewers assessing the same batches differed by 0.79 standard deviations on the composite index, and a shift in reviewer composition flattened the raw trend. After adjusting for reviewer effects and smoothing batch-level estimates, evaluated reliability increased by 0.53 standard deviations from March to August 2026. Reliability also shifted materially across logged deploys. At the estimated rate of unaccounted movement, forecast uncertainty reaches the magnitude of the largest observed deploy contrast after about five weeks, although this horizon is imprecisely estimated and based on a small number of step-like changes. An independent operational outcome moved in the same direction: interview-related support tickets fell by about 10% per month relative to technical-issue tickets. These data do not identify a causal effect of evaluation on performance. Instead, they show that longitudinal AI evaluation requires three checks: comparability across raters, evidence currency under drift, and corroboration with consequential operational outcomes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑