arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25408cs.AIcs.IRcs.LG

从离线代理到在线决策:面向对话式AI的分层参与度评估框架

From Offline Proxies to Online Decisions: A Layered Engagement Evaluation Framework for Conversational AI

Xuanyi Li, Vaskar Nath, Hossein Amirkhani, Jay Li, Alex Deng

首次发表
浏览论文内容

中文总结 AI 辅助

提出分层评估框架,将离线代理分解为三个对齐环节,通过区间感知决策一致性审计,在27个实验中达到81.1% F1,优于原始分数,用于优先排序对话式AI候选变更。

中文摘要 AI 辅助

在线A/B实验是衡量用户参与度的决策标准,但流量和读取时间限制了可测试的对话式AI变更数量。我们探讨一个无需处理组用户暴露即可计算的离线信号,是否与这些实验的结果一致。我们贡献了一个可复用的构建与诊断清单,将离线代理视为三个对齐环节的链条:行为标签与产品结果的对齐、学习分类器与候选助手行为的对齐、以及聚合离线信号与实验效果的对齐。配套的评估协议通过区间感知的决策一致性(比较离线与在线置信区间而非点估计)和实验内排序来审计整个复合体。我们评估的实例包括:一个固定的评估套件用于对候选行为评分,一个训练用于预测会话/提示级别参与度的参与度分类器,以及一个将样本级分数差异映射到在线模型级参与度变化的校准层。随后我们报告审计结果:来自部署的多轮助手上27个实验的489对配对离线-在线对比(一个候选臂与其对照),涵盖模型检查点到系统提示调优。我们的主要测试使用地图冻结后运行的8个实验中的113个对比:在这些对比上,复合体达到81.1%的F1分数,而其所基于的原始分类器分数仅为34.3%,并且在原始分数出现31次错误方向调用的情况下,复合体没有做出任何错误方向调用。每个离线预测都在其实验运行之前计算,以防止过拟合。证据支持在稀缺的实验流量分配之前使用复合体来优先排序候选——在我们部署的实验中,用于在训练检查点和系统提示调优之间进行选择。

英文摘要

Online A/B experiments are the decision standard for user engagement, but traffic and readout time limit how many conversational-AI changes can be tested. We ask whether an offline signal designed to be computable without treatment-arm user exposure agrees with the outcomes of those experiments. We contribute a reusable construction and diagnosis checklist that treats an offline proxy as a chain of three alignments: behavioral label to product outcome, learned classifier to candidate-assistant behavior, and aggregated offline signal to experiment effect. A companion evaluation protocol audits the whole composite by interval-aware decision agreement, which compares offline and online confidence intervals instead of point estimates, and by within-experiment ranking. The instantiation we evaluate comprises a fixed evaluation suite on which candidate behavior is scored, an engagement classifier trained to predict session/prompt level engagements, and a calibration layer mapping sample-level score differences to online model-level engagement deltas. We then report the audit: 489 paired offline-online contrasts (one candidate arm against its control) from 27 experiments on a deployed multi-turn assistant, spanning model checkpoints to system-prompt tuning. Our primary test uses the 113 contrasts from eight experiments that ran after the map was frozen: on these the composite reaches 81.1% F1, against 34.3% for the raw classifier score it is built on, and makes no wrong-direction calls where that raw score makes 31. Every offline prediction was computed before its experiment ran to prevent overfitting. The evidence supports using the composite to prioritize candidates before scarce experiment traffic is allocated---in our deployment of the experiment, selecting among training checkpoints and tuning system prompts.

发表机构

  • Meta Platforms, Inc.(Meta平台公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑