发表机构
U.S. Army Research Laboratory; Lobster Robotics(美国陆军研究实验室; Lobster Robotics)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出风险增强行为估值框架,通过Prelec概率加权调整信息目标,使危险探索机器人在高信息与低风险路径间权衡,实验证明其能重塑信息风险前沿并减少损失。
AI 中文摘要
危险机器人探索要求机器人在绘制空间风险(如不安全地形、辐射、火灾、地雷或结构损坏)的同时,在收集信息本身可能导致失败的环境中运行。一条信息量极高的路径可能使机器人暴露于危险之中,终止执行,并阻止未来的观测。因此,危险探索不仅需要决定不确定性最大的位置,还需要决定何时降低不确定性值得承担风险。本文引入了对这一问题的估值层视角。我们保持信念更新、传感器模型、物理风险模型和有限时域信息规划器不变,仅改变用于对可行路径排序的标量目标。在此框架内,我们引入了一种基于Prelec概率加权的风险增强行为信息目标,产生了一个可解释的从保守到激进的信息风险估值系列。理论上,我们证明了估值参数在高信息/高风险路径与低信息/低风险路径之间创建了切换边界,并在可行探索策略上诱导出变换后的帕累托前沿结构。大规模失败截断的网格世界实验表明,仅估值就能重塑信息风险前沿。香农信息规划仍然是一个强大的原始信息基线,而风险感知目标可以通过避免截断未来感知的失败来减少危险暴露和机器人损失。风险增强行为估值与标准风险感知基线具有帕累托竞争力,并提供可解释的保守和中间机制。这些结果支持一个框架,在该框架中,机器人不仅推理一个动作减少了多少不确定性,还推理这种减少是否值得为获得它而承担的风险。
英文摘要
Hazardous robotic exploration requires robots to map spatial risks, such as unsafe terrain, radiation, fire, mines, or structural damage, while operating where collecting information can itself cause failure. A highly informative path may expose the robot to hazards, terminate execution, and prevent future observations. Hazardous exploration therefore requires deciding not only where uncertainty is largest, but when reducing it is worth the risk. This paper introduces a valuation-layer view of this problem. We keep the belief update, sensor model, physical risk model, and finite-horizon informative planner fixed, and change only the scalar objective used to rank feasible paths. Within this framework, we introduce a risk-augmented Behavioral Information objective based on Prelec probability weighting, yielding an interpretable family of conservative-to-aggressive information-risk valuations. Theoretically, we show that valuation parameters create switching boundaries between high-information/high-risk and lower-information/lower-risk paths, and induce a transformed Pareto-frontier structure over feasible exploration policies. Large-scale failure-truncated grid-world experiments show that valuation alone reshapes the information-risk frontier. Shannon information planning remains a strong raw-information baseline, while risk-aware objectives can reduce hazard exposure and robot losses by avoiding failures that truncate future sensing. Risk-augmented Behavioral valuation is Pareto-competitive with standard risk-aware baselines and provides interpretable conservative and intermediate regimes. These results support a framework in which robots reason not only about how much uncertainty an action reduces, but whether that reduction is worth the risk required to obtain it.
CommentsAccepted at the International Symposium of Robotics Research (ISRR) 2026