AcousticDiffusion:用于搜救辅助的语义条件音频引导扩散策略
AcousticDiffusion: Semantically Conditioned Audio-Guided Diffusion Policy for Search-and-Rescue Assistance
浏览论文内容
中文总结 AI 辅助
提出AcousticDiffusion,一种语义条件音频引导扩散策略,利用声学观测和贝叶斯置信场生成航点轨迹,在仿真和实物上显著提升搜救机器人导航至呼叫者的精度与效率。
中文摘要 AI 辅助
在视觉接触退化或被遮挡的环境中,救援机器人导航至人类呼叫者是重要的能力。我们提出AcousticDiffusion,一种语义条件、音频引导的扩散策略,用于人类定向导航。一个冻结的预训练音频识别器处理10.24秒窗口,通过语音门控和基于痛苦的优先级排序,将识别输出转换为源级导航角色。麦克风阵列的到达方向测量被递归集成到机器人中心的贝叶斯鸟瞰置信场中。自运动补偿对齐连续观测,逐步约束源位置,同时保留由方位引起的距离不确定性。语义置信度、最近的声学观测、音频特征和机器人状态共同条件化一个生成航点轨迹的扩散模型。在使用录制音频的合成导航验证集上,AcousticDiffusion实现了平均终点方位误差11.20度,91.78%的轨迹在呼叫者30度内对齐。干扰物拒绝率范围为89.20%至98.99%,策略在91.07%的窗口中选择HELP指定的呼叫者而非竞争说话者。在ZSL-1四足机器人上无需额外重训练即可在线部署,平均方位误差为64.9度,而A*为98.2度,RRT为90.4度,平均规划器计算时间为6.07毫秒。尽管声学定位不完美,报告的平均最终源距离从使用ODAS(开放嵌入式听觉系统)引导的经典规划器的3.96米减少到2.48米,改善了37.4%。这些结果证明了该框架将不确定的声学观测转化为更接近人类呼叫者的能力。
英文摘要
Navigating toward human callers is an important capability for rescue robots operating where visual contact is degraded or occluded. We present AcousticDiffusion, a semantically conditioned, audio-guided diffusion policy for human-directed navigation. A frozen pretrained audio recognizer processes 10.24 s windows, with speech gating and distress-aware prioritization converting recognition outputs into source-level navigation roles. Microphone-array direction-of-arrival measurements are recursively integrated into a robot-centric Bayesian bird's-eye-view belief field. Ego-motion compensation aligns successive observations, progressively constraining source position while preserving bearing-induced range uncertainty. The semantic belief, recent acoustic observations, audio features, and robot state condition a diffusion model that generates waypoint trajectories. On a synthetic-navigation validation set using recorded audio, AcousticDiffusion achieves a mean end-point bearing error of 11.20 degrees, with 91.78% of trajectories aligned within 30 degrees of the caller. Distractor rejection ranges from 89.20% to 98.99%, and the policy favors a HELP-designated caller over a competing speaker in 91.07% of windows. Deployed online on a ZSL-1 quadruped without additional retraining, it achieves a mean bearing error of 64.9 degrees, compared with 98.2 degrees for A* and 90.4 degrees for RRT, with a mean planner compute time of 6.07 ms. Despite imperfect acoustic localization, the reported mean final source distance is reduced from 3.96 m for the classical planners using ODAS-derived (Open embedded Audition System) guidance to 2.48 m, a 37.4% improvement. These results demonstrate the framework's ability to translate uncertain acoustic observations into closer approaches to human callers.
发表机构
- Skolkovo Institute of Science and Technology(斯科尔科沃科学技术学院)
机构由 AI 辅助整理,请以论文原文为准。