LH-AVLN:长距离视听语言导航基准测试
LH-AVLN: A Benchmark for Long-Horizon Audio-Visual-Language Navigation
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- Jilin University(吉林大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
介绍长距离视听语言导航基准LH-AVLN,结合多目标任务执行等。提出免训练参考智能体PAG-Nav,可维护语义地图并规划。实验表明现有智能体完成任务有困难,PAG-Nav提供更强诊断基线。
AI中文摘要:
具身导航正朝着长距离任务发展,但现有长距离基准测试大多无声,视听导航任务通常聚焦单一目标。我们引入LH-AVLN,一个结合多目标任务执行、异构目标规范和持续空间声学线索的长距离视听语言导航基准测试。
英文摘要:
Embodied navigation is moving toward long-horizon missions, yet existing long-horizon benchmarks are largely acoustically silent, and audio-visual navigation tasks typically focus on a single goal. We introduce LH-AVLN, a benchmark for Long-Horizon Audio-Visual-Language Navigation that combines multi-goal mission execution, heterogeneous goal specifications, and persistent spatialized acoustic cues. In LH-AVLN, an agent receives a global mission of two to four goals specified by category, language description, or reference image, and navigates with RGB-D observations, pose, and binaural audio in indoor 3D environments. The benchmark supports both ordered and unordered missions, where alternating goal-associated sounds can guide non-line-of-sight search but may also become distractors as mission progress changes. We further develop PAG-Nav, a training-free reference agent that maintains a temporal uniform semantic map and performs progressive goal-state planning, using sound for search while reserving completion for visual-semantic verification. Experiments show that existing vision-language, memory-based, and audio-visual agents struggle to complete full LH-AVLN missions, and that PAG-Nav provides a stronger diagnostic baseline while leaving substantial room for future progress.