arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

导航员的语言:面向冻结空中视觉语言导航(VLN)智能体的轨迹基础指令翻译

Speaking the Navigator's Language: Trajectory-Grounded Instruction Translation for Frozen Aerial VLN Agents

Xi Chen, Zhe Liu, Xiaogang Xu, Jiafei Xu, Chunyi Zhou, Yuan Su, Rui Zeng, Tianyu Du, Kelu Yao, Chao Li, Shouling Ji

arXiv 2610.10635首次发表:更新:

发表机构

Zhejiang University; Zhejiang Lab(浙江大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对冻结空中VLN智能体存在的指令差距问题,提出轨迹基础指令翻译器(TGIT),通过翻译指令提升了导航成功率,实现了零样本迁移及在多个数据集上的性能改善。

AI 中文摘要

空中视觉语言导航(VLN)智能体通常在细节丰富、与轨迹对齐的指令上进行训练,而用户发出的是简短、以意图为导向的指令;在冻结的OpenFly导航器上,这种“指令差距”使成功率(SR)从31.03%降至11.33%。为扩大翻译器训练规模,我们用人类编写的风格示例提示语言模型,将原始指令转换为配对的、以意图为中心的弱指令,这使成功率达到15.27%。我们引入轨迹基础指令翻译器(TGIT),这是一个前端模块,保持导航器冻结,通过从其轨迹结果中学习,将弱输入转换为智能体可执行的指令。由此产生的经弱指令训练的翻译器将弱输入的成功率提升至37.93%,并零样本迁移至真实人类指令(从11.33%提升至32.51%);它还提升了未见过的OpenFly的性能(从4.95%提升至20.79%),并在CityNav和AirVLN上实现了性能恢复。

英文摘要

Aerial vision-and-language navigation (VLN) agents are typically trained on detail-rich, trajectory-aligned commands, whereas users issue short, intent-driven instructions; on a frozen OpenFly navigator, this \emph{instruction gap} drops success rate (SR) from $31.03\%$ to $11.33\%$. To scale translator training, we prompt a language model with human-written style examples to convert original commands into paired, intent-centered Weak commands, which yield $15.27\%$ SR. We introduce the \textbf{Trajectory-Grounded Instruction Translator (TGIT)}, a front-end that keeps the navigator frozen and translates Weak inputs into agent-executable commands by learning from its trajectory outcomes. The resulting Weak-trained translator raises Weak-input SR to $37.93\%$ and transfers zero-shot to real human instructions ($11.33\%{\rightarrow}32.51\%$); it also improves held-out OpenFly ($4.95\%{\rightarrow}20.79\%$) and yields recovery on CityNav and AirVLN.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑