arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.03920cs.RO

LH-AVLN:长距离视听语言导航基准测试

LH-AVLN: A Benchmark for Long-Horizon Audio-Visual-Language Navigation

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Jilin University(吉林大学)

机构由 AI 辅助整理,请以论文原文为准。

Rufeng Chen, Yue Chang, Zili Shao, Zhaofan Zhang, Li Chen, Hechang Chen, Hui Xiong, Sihong Xie

AI总结:

介绍长距离视听语言导航基准LH-AVLN,结合多目标任务执行等。提出免训练参考智能体PAG-Nav,可维护语义地图并规划。实验表明现有智能体完成任务有困难,PAG-Nav提供更强诊断基线。

AI中文摘要:

具身导航正朝着长距离任务发展,但现有长距离基准测试大多无声,视听导航任务通常聚焦单一目标。我们引入LH-AVLN,一个结合多目标任务执行、异构目标规范和持续空间声学线索的长距离视听语言导航基准测试。

英文摘要:

Embodied navigation is moving toward long-horizon missions, yet existing long-horizon benchmarks are largely acoustically silent, and audio-visual navigation tasks typically focus on a single goal. We introduce LH-AVLN, a benchmark for Long-Horizon Audio-Visual-Language Navigation that combines multi-goal mission execution, heterogeneous goal specifications, and persistent spatialized acoustic cues. In LH-AVLN, an agent receives a global mission of two to four goals specified by category, language description, or reference image, and navigates with RGB-D observations, pose, and binaural audio in indoor 3D environments. The benchmark supports both ordered and unordered missions, where alternating goal-associated sounds can guide non-line-of-sight search but may also become distractors as mission progress changes. We further develop PAG-Nav, a training-free reference agent that maintains a temporal uniform semantic map and performs progressive goal-state planning, using sound for search while reserving completion for visual-semantic verification. Experiments show that existing vision-language, memory-based, and audio-visual agents struggle to complete full LH-AVLN missions, and that PAG-Nav provides a stronger diagnostic baseline while leaving substantial room for future progress.

补充信息

↑