AI 中文总结
该研究提出信念更新门,分离人机交互中信念报告的未变化与变化,重新分析数据集发现大量信念未变化,修正了对合并保守性的解读,建议校准分析需区分这两种情况。
AI 中文摘要
反复人机交互常通过合并信念更新斜率分析:用户观察AI的成功与失败,会沿反馈一致的方向修正报告的信念,但平均而言表现保守。本文指出这类平均值会掩盖两个关键区别:一是引出的信念报告是否发生改变,二是改变时的变化幅度,我们将这种感知测量的分解称为信念更新门。通过重新分析包含240名参与者、7200次试次、三个任务领域的多任务人机决策数据集,我们发现报告的信念存在大量未变化的情况:67.3%的试次级信念变化恰好为零,76.4%的变化小于5个百分点。将未变化与变化的报告分离后,对合并保守性的描述性解读发生改变:整体的轨迹内斜率从0.494上升至非零变化试次中的0.949。由于后一估计基于观察到的变化,我们将其解读为描述性分解,而非接近贝叶斯潜在学习过程的证据。互补的 hurdle 风格分析(即先建模零变化与非零变化,再预测更新幅度)显示,反馈与初始信念的绝对差异会预测报告是否发生改变,而带符号的反馈差异则会预测发生改变的报告的方向与幅度。重要的是,观察到的未变化无法区分真正的潜在信念惯性与未表达的小幅更新、四舍五入或其他报告过程。这些发现表明,反复人机交互的校准分析应区分引出的信念报告中的可见未变化与基于变化条件的更新,而非将报告的信念视为单一的连续更新过程。
英文摘要
Repeated human-AI interaction is often analyzed through pooled belief-updating slopes: users observe AI successes and failures, revise reported beliefs in the feedback-consistent direction, but appear conservative on average. We show that such averages can obscure an important distinction between whether an elicited belief report changes at all and how it changes conditional on movement. We refer to this measurement-aware decomposition as the belief update gate. Reanalyzing a multi-task human-AI decision-making dataset with 240 participants, 7,200 trials, and three task domains, we find substantial non-movement in reported beliefs: 67.3% of trial-level belief changes are exactly zero, and 76.4% are smaller than five percentage points. Separating non-moving from moving reports changes the descriptive interpretation of pooled conservatism: the within-trajectory slope rises from 0.494 overall to 0.949 among rows with nonzero movement. Since this latter estimate conditions on observed movement, we interpret it as a descriptive decomposition rather than as evidence of a near-Bayesian latent learning process. Complementary hurdle style analyses (i.e., modeling zero vs. non-zero changes before predicting update magnitude) show that the absolute discrepancy between feedback and entering belief predicts whether a report changes, while the signed feedback discrepancy predicts the direction and magnitude of change among reports that move. Importantly, observed non-movement does not distinguish genuine latent belief inertia from small unexpressed updates, rounding, or other reporting processes. These findings show that calibration analyses of repeated human--AI interaction should distinguish visible non-movement in elicited belief reports from updating conditional on movement rather than treating reported beliefs as a single continuous updating process.