失败机器人何时应询问?基于审计传感器证据发起纠正性人机对话
When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence
浏览论文内容
中文总结 AI 辅助
针对失败机器人是否应询问人类的决策,构建模拟基准并测试视觉语言模型,发现模型行为受提示影响而非证据,提出应依据实测准确率与成本而非置信度来决定询问。
中文摘要 AI 辅助
在执行任务中失败的机器人面临纠正性对话中的第一个决策:根据自身的诊断采取行动、咨询另一个机载传感器,还是打断人类。做出好的选择需要知道机器人的传感器在多大程度上揭示了失败原因,以及机器人自身诊断的可靠性。我们构建了一个模拟基准,其中每个失败的真实原因都是已知的,因为我们注入了这些原因,并测量了每个传感器所揭示的信息,同时明确检查了数据泄漏。有些失败可以从相机图像中诊断出来;其他失败只能从机器人的力数据中诊断出来(力数据的准确率为0.99,而图像方法均不高于0.55)。然后我们测试了六个开源的视觉语言模型。它们的行为跟随提示的表面形式,而非证据:在六个扫描的模型-家族组合中,将拒绝选项从答案列表的最后移到最前,导致其中三个组合的拒绝率从78-100%骤降至0-6%。在每种提示变体下,无论有无工作示例,基于帧的准确率始终等于或低于多数类基线,且所陈述的置信度不包含关于正确性的任何信息。将相同的力数据以十行文本的形式提供给这些模型,在六个模型中的四个中产生了首次高于基线的诊断:大部分失败反映了缺失的传感器数据,而非缺失的能力。我们将该选择表述为一个三动作决策问题:行动、咨询自身传感器或询问人类,其最优策略由测量的准确率决定。模型并未遵循该策略,且它们的询问率忽略了一个四倍变化的提问成本。向人类提出一个问题仍然使它们从基线提升到大致等于回答者自身的可靠性(当它们询问时,准确率为0.70-0.81)。是否询问的决策应基于测量的准确率和明确的成本,而非模型的置信度。
英文摘要
A robot that fails at a task faces the first decision in corrective dialogue: act on its own diagnosis, consult another onboard sensor, or interrupt a person. Choosing well requires knowing how much the robot's sensors reveal about the cause and how reliable the robot's own diagnosis is. We build a simulated benchmark in which every failure's true cause is known, because we injected it, and measure what each sensor reveals, with explicit checks against data leakage. Some failures are diagnosable from camera images; others only from the robot's force data (0.99 from force data, no image method above 0.55). We then test six open vision-language models. Their behavior tracks the surface of the prompt, not the evidence: moving the refusal option from last to first in the answer list collapses refusal rates from 78-100% to 0-6% in three of the six swept model-and-family pairs. Accuracy from frames stays at or below a majority-class baseline under every prompt variant, with or without worked examples, and stated confidence carries no information about correctness. Handing the same models the force data as ten lines of text produces the first above-baseline diagnoses, in four of the six models: much of the failure reflects missing sensor data, not missing ability. We pose the choice as a three-action decision problem, act, consult your own sensors, or ask a human, whose optimal policy follows from measured accuracy. The models do not follow it, and their ask rates ignore a fourfold change in question cost. One question to a human still lifts them from that baseline to roughly the answerer's own reliability (0.70-0.81 when they ask). The decision to ask should be tied to measured accuracy and stated costs, not to the model's confidence.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Centific
机构由 AI 辅助整理,请以论文原文为准。