发表机构
Zhejiang University; Zhejiang Lab(浙江大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对冻结空中VLN智能体存在的指令差距问题,提出轨迹基础指令翻译器(TGIT),通过翻译指令提升了导航成功率,实现了零样本迁移及在多个数据集上的性能改善。
AI 中文摘要
空中视觉语言导航(VLN)智能体通常在细节丰富、与轨迹对齐的指令上进行训练,而用户发出的是简短、以意图为导向的指令;在冻结的OpenFly导航器上,这种“指令差距”使成功率(SR)从31.03%降至11.33%。为扩大翻译器训练规模,我们用人类编写的风格示例提示语言模型,将原始指令转换为配对的、以意图为中心的弱指令,这使成功率达到15.27%。我们引入轨迹基础指令翻译器(TGIT),这是一个前端模块,保持导航器冻结,通过从其轨迹结果中学习,将弱输入转换为智能体可执行的指令。由此产生的经弱指令训练的翻译器将弱输入的成功率提升至37.93%,并零样本迁移至真实人类指令(从11.33%提升至32.51%);它还提升了未见过的OpenFly的性能(从4.95%提升至20.79%),并在CityNav和AirVLN上实现了性能恢复。
英文摘要
Aerial vision-and-language navigation (VLN) agents are typically trained on detail-rich, trajectory-aligned commands, whereas users issue short, intent-driven instructions; on a frozen OpenFly navigator, this \emph{instruction gap} drops success rate (SR) from $31.03\%$ to $11.33\%$. To scale translator training, we prompt a language model with human-written style examples to convert original commands into paired, intent-centered Weak commands, which yield $15.27\%$ SR. We introduce the \textbf{Trajectory-Grounded Instruction Translator (TGIT)}, a front-end that keeps the navigator frozen and translates Weak inputs into agent-executable commands by learning from its trajectory outcomes. The resulting Weak-trained translator raises Weak-input SR to $37.93\%$ and transfers zero-shot to real human instructions ($11.33\%{\rightarrow}32.51\%$); it also improves held-out OpenFly ($4.95\%{\rightarrow}20.79\%$) and yields recovery on CityNav and AirVLN.