人机协作中的指称不确定性
Referential Uncertainty in Human--AI Collaboration
- Microsoft Research(微软研究院)
- Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过协作拼图任务揭示人机协作中的指称不确定性,发现引出信念分布优于原始概率,且外化不确定性仅在准确定位时才能有效帮助人类伙伴。
AI中文摘要:
有效的人机协作要求伙伴通过交互建立指称关系,而当描述模糊、指称对象相似或伙伴所见不同时,这种关系会变得脆弱。我们研究指称不确定性——即描述所指的候选对象的不确定性——在一个协作拼图任务中,人类助手指导AI工人放置拼图块。工人必须识别并传达其不确定性,而助手必须识别并据此采取行动。我们表明,单独引出的关于候选拼图块的信念分布比原始动作标记概率具有更好的校准性(ECE 0.15)和更好的区分正确与错误放置的能力(AUROC 0.65),而原始动作标记概率严重过度自信(平均置信度0.97,ECE 0.44)。在三个前沿视觉语言模型(GPT-4.1、GPT-5、GPT-5.5)中,这种引出的不确定性随指令模糊性可预测地上升,但不会随上下文中的竞争指称对象而上升,即使这些对象增加了错误。模型很少外化这种不确定性,仅在3.5%-16.7%的回合中请求澄清。在一项受控人类研究(N=210)中,仅获得工人默认消息的参与者接受了78%的错误放置,且无法区分正确与错误(AUC 0.50)。精确的描述,尤其是针对性强的对冲表达,将错误动作接受率降至36%,同时基本保持正确动作接受率,弥补了缺乏共享意识(如看不到工人的动作)的不足。但这种益处取决于针对性:从模型自身信念熵导出的可部署对冲表达继承了该信号的弱点,可能弊大于利。外化的不确定性只有在准确定位时才能帮助人类伙伴。
英文摘要:
Effective human-AI collaboration requires partners to establish references through interaction, which becomes fragile when descriptions are ambiguous, similar referents compete, or partners see different things. We study referential uncertainty - uncertainty over which candidate object a description refers to - in a collaborative puzzle task where a human Helper instructs an AI Worker to place pieces. The Worker must identify and communicate its uncertainty, and the Helper must recognize and act on it. We show that a separately elicited belief distribution over candidate pieces is better calibrated (ECE 0.15) and better discriminates correct from incorrect placements (AUROC 0.65) than raw action-token probabilities, which are severely overconfident (0.97 mean confidence, ECE 0.44). Across three frontier vision-language models (GPT-4.1, GPT-5, GPT-5.5), this elicited uncertainty rises predictably with instruction vagueness, but not with competing referents in context, even when those increase errors. The models seldom externalize it, asking for clarification on only 3.5-16.7% of turns. In a controlled human study (N=210), participants given only the Worker's default message accept 78% of wrong placements and cannot tell right from wrong (AUC 0.50). Precise descriptions and, especially, well-targeted hedges cut wrong-move acceptance to 36% while largely preserving correct-move acceptance, compensating for missing shared awareness such as not seeing the Worker's action. But this benefit depends on targeting: a deployable hedge derived from the model's own belief entropy inherits that signal's weakness and can do more harm than good. Externalized uncertainty helps a human partner only when it is accurately targeted.