不确定时不要切割:用于机器人采摘中VLA策略的可拒绝且校准的决策头
Do Not Cut When Uncertain: Rejectable and Calibrated Decision Heads for VLA Policies in Robotic Harvesting
浏览论文内容
中文总结 AI 辅助
针对机器人采摘中VLA策略因遮挡导致的模糊决策问题,提出可拒绝且校准的决策头(RCDH),通过显式弃权(不执行)与校准机制,仅替换输出头即可提升分布外可靠性,证明拒绝权比模型规模更关键。
中文摘要 AI 辅助
使用行为克隆或流匹配训练的视觉-语言-动作(VLA)策略被优化以输出动作轨迹,但它们无法表达“我不知道”或“我不应该行动”。在机器人采摘中,遮挡使得单帧决策本质上具有模糊性:相同的像素既可以对应可切割的茎,也可以对应根本没有茎。现有的VLA被迫做出承诺,导致具有不可逆后果的高置信度错误。我们认为VLA的失败模式不是由骨干网络规模决定的,而是由其输出接口决定的。我们提出了可拒绝且校准的决策头(RCDH),这是一种类型化、可拒绝且校准的输出接口,可以附加到冻结的VLA骨干网络上,无需重新训练或新特征。RCDH引入了(i)具有显式拒绝和有序、条件分解的决策模式,以及(ii)用于风险感知弃权(不执行)的校准程序。我们在具有可控叶片遮挡的机器人采摘平台上评估了RCDH,比较了生成式、枚举式、校准式和可拒绝式接口。我们表明,仅替换输出头即可在遮挡下恢复分布外可用性,同时保持分布内性能。我们进一步测试了拒绝空间的有序性是否关键。我们的结果表明,拒绝的权利,而非更大的模型,是在不确定性下实现可靠操作所缺失的接口。
英文摘要
Vision-Language-Action (VLA) policies trained with behavior cloning or flow matching are optimized to output an action trajectory, but they cannot express "I don't know" or "I should not act." In robotic harvesting, occlusion makes single-frame decisions fundamentally ambiguous: identical pixels can correspond either to a cuttable stem or to no stem at all. Existing VLAs are forced to commit, leading to high-confidence errors with irreversible consequences. We argue that the failure mode of a VLA is determined not by backbone scale but by its output interface. We propose Rejectable and Calibrated Decision Heads (RCDH), a typed, rejectable, and calibrated output interface that can be attached to a frozen VLA backbone without retraining or new features. RCDH introduces (i) a decision schema with explicit rejection and ordered, conditional decomposition, and (ii) a calibration procedure for risk-aware abstention. We evaluate RCDH on a robotic harvesting platform with controllable leaf occlusion, comparing generative, enumerated, calibrated, and rejectable interfaces. We show that replacing only the output head restores out-of-distribution usability under occlusion while preserving in-distribution performance. We further test whether the ordering of the rejection space is critical. Our results suggest that the right to refuse, rather than a larger model, is the missing interface for reliable manipulation under uncertainty.