发表机构
University of Wisconsin - Madison; Brown University; Pinterest Inc(威斯康星大学麦迪逊分校; 布朗大学; Pinterest 公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对用户纠正可能错误的问题,提出GAVA框架,通过接受、拒绝、检查世界和询问的仲裁机制,利用观察受限证据和期望损失规则,在ALFWorld中实现高准确率并降低交互成本。
AI 中文摘要
当人的纠正可能是错误时,具身智能体应如何回应?我们将接地纠正仲裁(grounded correction arbitration)表述为在接受、拒绝、检查世界和询问说话者之间的选择。GAVA通过观察受限证据、合法探针和一步期望损失规则实现这一接口。在纯文本ALFWorld中,162个检查点产生972对真假干预。完整的局部检查使GAVA和always verify达到100%的纠正准确率,确立了证据契约而非比较优势。在同集执行中,GAVA相对于always verify降低了交互成本,但在完美说话者下与成本阈值持平。一个仅用于训练的探索性物体位置先验,在340个未见场景中,相对于均匀GAVA,将交互成本和声明联合成本分别降低了0.490和0.420。在冻结策略、成本、基线和多样性计划后,这些收益在77个非重叠的已见检查点(覆盖308个场景)上复现:分别为0.595和0.517,两个95%检查点自助置信区间均排除零。联合成本也优于相同先验的固定策略,而匹配的校准无VOI比较仍无定论。语义GAVA在每个队列中产生四个事实错误,对应98.8%和98.7%的准确率,所有方法均完成每个任务。结果支持在声明成本下使用语义先验进行选择性信息收集,但未确立环境信息价值相对于澄清的普遍优势。该研究使用规范化声明、完整符号观察和受控说话者;它既未评估人类参与者、视觉输入,也未评估物理机器人。
英文摘要
How should an embodied agent respond when a person's correction may be wrong? We formulate grounded correction arbitration as a choice among accepting, rejecting, inspecting the world, and asking the speaker. GAVA implements this interface with observation-bounded evidence, legal probes, and a one-step expected-loss rule. In text-only ALFWorld, 162 checkpoints produce 972 paired true and false interventions. Complete local inspections give GAVA and always verify 100 percent correction accuracy, establishing the evidence contract rather than a comparative advantage. In same-episode execution, GAVA reduces interaction cost against always verify but ties a cost threshold under a perfect speaker. An exploratory training-only object-location prior lowers interaction and declared joint cost on 340 unseen scenarios by 0.490 and 0.420 relative to uniform GAVA. After freezing the policy, costs, baselines, and multiplicity plan, the gains replicate on 77 non-overlapping seen checkpoints, covering 308 scenarios: 0.595 and 0.517, with both 95 percent checkpoint-bootstrap confidence intervals excluding zero. Joint cost also improves over an identical-prior fixed policy, while the matched calibrated no-VOI comparison remains inconclusive. Semantic GAVA makes four factual errors in each cohort, corresponding to 98.8 percent and 98.7 percent accuracy, and all methods complete every task. Results support selective information gathering with semantic priors under declared costs, but do not establish a general advantage of environmental value of information over clarification. The study uses normalized claims, complete symbolic observations, and controlled speakers; it evaluates neither human participants, visual input, nor physical robots.