arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36906cs.CV

SafeVantage:面向可靠具身决策的视角感知记忆

SafeVantage: Vantage-Aware Memory for Reliable Embodied Decisions

  • Massachusetts Institute of Technology(麻省理工学院)
  • Korea Advanced Institute of Science and Technology(韩国科学技术院)
  • Nanyang Technological University(南洋理工大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Sean Hardesty Lewis, Zuyi Guo, Benwang Chen, Zirui Li, Hongyi Lin, Heye Huang

AI总结:

SafeVantage提出视角感知语义记忆与主动获取框架,利用候选可观测性模型引导视角选择,在部分可观测条件下提升具身决策的可靠性与效率,显著优于基线。

AI中文摘要:

在部分可观测条件下,可靠的具身决策需要信息丰富的观测和充分的支撑证据。然而,仅凭语义分数无法揭示哪些视角能证明某个声明,也无法指出应在何处获取额外证据。我们提出了SafeVantage,一种视角感知的语义记忆与主动获取框架,该框架保留每个声明的支撑视图、相机姿态和估计的目标位置,并将正向支持与搜索覆盖区分开来。一个学习得到的候选可观测性模型利用基于声明的几何信息来预测在可达视角下目标的可见性。这些预测通过期望终端决策损失的减少来引导视角选择,同时考虑行进成本和几何上不同的佐证。一个校准头随后结合支持度、空间一致性和覆盖率,产生“是”、“否”或“弃权(不执行)”的决策。我们在一个类别存在性基准上评估了SafeVantage,该基准涵盖232个未见过的ProcTHOR房屋,每种方法和动作预算下各有7,424个配对情节。与验证集选择的等预算基线相比,SafeVantage在八次和十二次动作时分别实现了24.7%和12.0%的宏F1增益,在两种预算下具有更低的风险和更高的回答率,并且在八次动作时减少了31.7%的行进距离。等输入HM3D实验显示在固定观测下选择性风险更低,而受控的ScanNet干预表明恢复支撑视图能改善下游VLM的答案。消融研究进一步支持候选可观测性对决策质量和获取效率的贡献。结果表明,声明级视角证据对于连接语义记忆、主动获取和可靠决策制定具有价值。代码可在该https URL获取。

英文摘要:

Reliable embodied decisions under partial observability require informative observations and sufficient supporting evidence. However, semantic scores alone do not reveal which viewpoints justify a claim or where additional evidence should be acquired. We introduce SafeVantage, a vantage-aware semantic memory and active acquisition framework that retains each claim's supporting views, camera poses, and estimated target location, keeping positive support distinct from search coverage. A learned candidate-observability model uses claim-grounded geometry to predict target visibility at reachable viewpoints. These predictions guide view selection through expected reduction in terminal decision loss, accounting for travel cost and geometrically distinct corroboration. A calibrated head then combines support, spatial consistency, and coverage to produce Yes, No, or Abstain decisions. We evaluate SafeVantage on a category-presence benchmark spanning 232 unseen ProcTHOR houses and 7,424 paired episodes per method and action budget. Compared with validation-selected equal-budget baselines, SafeVantage achieves macro-F1 gains of 24.7% and 12.0% at eight and twelve actions, respectively, with lower risk and higher answer rates at both budgets and 31.7% less travel at eight actions. Equal-input HM3D experiments show lower selective risk under fixed observations, while controlled ScanNet interventions show that restoring supporting views improves downstream VLM answers. Ablations further support the contribution of candidate observability to decision quality and acquisition efficiency. Results demonstrate the value of claim-level viewpoint evidence for connecting semantic memory, active acquisition, and reliable decision-making. Code is available at https://safevantage.github.io

↑