发表机构
Monash University, Indonesia; SEACrowd(蒙纳士大学印度尼西亚校区; SEACrowd)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对VLN中标准共形预测无法在依赖且变长片段上提供覆盖保证的问题,提出ENCP,通过残差置信度重缩放非一致性分数并每片段校准一个最大分数,在R2R和REVERIE上达到所有步级覆盖目标,提供模型无关的不确定性估计。
AI 中文摘要
不确定性估计对于视觉与语言导航(VLN)模型而言是一项关键任务,因为它有助于识别模糊且不可靠的预测,从而使智能体能够做出更安全的导航决策。作为最先进的不确定性估计框架之一,共形预测(CP)为VLN中的不确定性估计提供了一种有前景的方法。然而,由于VLN智能体需要执行一系列步骤,共形预测中的标准校准无法在依赖性强、长度可变的VLN片段上提供其所承诺的覆盖保证。为此,我们提出了片段归一化共形预测(ENCP),该方法通过策略的残差置信度对非一致性分数进行重新缩放,并为每个片段校准一个最大分数。在可交换的校准和测试片段条件下,这种构造能够以至少$1 - \alpha$的概率覆盖每一步的真实值,同时允许片段内各步骤之间存在依赖性。在R2R和REVERIE数据集上,针对四种VLN策略和三种非一致性分数,ENCP在从已见到未见环境的评估中达到了所有报告的实证步级覆盖目标。这些结果表明,ENCP能够提供与模型无关的不确定性估计,这可能有助于确定VLN智能体何时应弃权(不执行)并交由更强大的预测器(包括人工协助)来处理。
英文摘要
Uncertainty estimation for Vision-Language-Navigation (VLN) models is a critical task since it can help identify ambiguous and unreliable predictions, enabling agents to make safer navigation decisions. As one of the most advanced uncertainty estimation frameworks, conformal prediction (CP) offers a promising approach for uncertainty estimation in VLN. However, given that VLN agent requires a sequence of steps, standard calibration in conformal prediction fails to provide coverage guarantee it promises over a dependent, variable-length VLN episode. To this end, we propose Episode-Normalized Conformal Prediction (ENCP), which rescales a nonconformity score by the policy's residual confidence and calibrates one maximum score per episode. Under exchangeable calibration and test episodes, this construction covers the ground truth at every step with probability at least $1 - α$, while allowing dependence among steps within an episode. Across four VLN policies and three nonconformity scores on R2R and REVERIE dataset, ENCP meets all reported empirical step-coverage targets on the seen-to-unseen evaluation. These results demonstrate that ENCP can provide model-agnostic uncertainty estimates, which might be useful for determining when a VLN agent should defer to a more capable predictor, including human assistance.
Comments8 pages, 5 figures