arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14952cs.ROcs.CVeess.SP

不存在的证据:当视觉失效时维持世界模型的跨模态溯因风险感知

Evidence of Absence: Cross-Modal Abductive Risk Perception to Sustain World Models When Vision Fails

Cong Xu, Ravi Sankar

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出跨模态溯因风险感知方法,在视觉失效时通过声学线索维持世界模型,可提前1.7秒预警,虚警减少42%,校准良好,能在视觉退化下保持高危险感知度。

中文摘要 AI 辅助

结构化世界状态(实体、关系、上下文及预测线索)旨在感知退化时保留预测关键内容,但它依赖观测来填充自身;当主视觉模态被遮挡或退化时,这些观测可能缺失。我们通过将预期共证据的缺失视为隐藏原因的证据,解决如何从互补模态维持世界模型的问题。该溯因框架与模态无关,本文以声学形式实例化:麦克风阵列前端估计发动机与轮胎声源的方位,并提取接近速率证据(稳定音调存在时用多普勒效应,否则用宽带逼近读数);随后“特征存在、视觉共证据缺失”的事件触发对隐藏道路使用者的溯因推理,输出校准后的风险建议而非控制指令。我们将隐藏状态的可恢复性分析为区分共享信息与模态唯一信息的可识别性问题,并将线索提示建模为明确虚警预算下的奈曼-皮尔逊检测。在盲交叉口的真实遮挡接近录音中,该方法比已发表声学基线的持续窗口变体检测率相当,但虚警减少42%,在可见后中位定位误差为3.4度,校准良好(预期校准误差0.034),在仅视觉通道崩溃至0.03的 staged 视觉退化下保持危险感知度高于0.87。我们还测量了该方法的极限:校准几乎无损失地迁移至未见过的交叉口,但特征分类器不迁移,自车移动噪声是部署的核心约束。

英文摘要

A structured world-state (entities, relations, context, and predictive cues) is designed to preserve prediction-critical content when perception degrades, but it presumes observations to populate it; when the primary visual modality is occluded or degraded, those observations may be missing. We address how to sustain the world model from a complementary modality by treating the absence of expected co-evidence as evidence of a hidden cause. The abductive framework is modality-agnostic; this article instantiates it acoustically. A microphone-array front-end estimates the bearing of engine and tire sources and extracts approach-rate evidence (Doppler when a stable tone exists, a broadband looming readout otherwise); the event "signature present, visual co-evidence absent" then triggers abductive inference of a hidden road user, emitting a calibrated risk advisory rather than a control command. Recoverability of the hidden state is analyzed as an identifiability question separating shared from modality-unique information, and cueing is cast as Neyman-Pearson detection under an explicit false-alarm budget. On real occluded-approach recordings at blind junctions, the method warns a mean 1.7 seconds before line-of-sight entry, matches the sustained-window variant of the published acoustic baseline's detection rate with 42% fewer false alarms, localizes to 3.4 degrees median once in view, is well calibrated (expected calibration error 0.034), and keeps hazard awareness above 0.87 under staged vision degradation that collapses a vision-only channel to 0.03. We also measure the method's limits: calibration transfers to an unseen junction almost losslessly, the signature classifier does not, and moving-ego noise is the binding deployment constraint.

发表机构

  • University of South Florida(南佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑