无需训练,飞行更佳:用于无人机导航的测试时缩放视觉语言模型
No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation
浏览论文内容
中文总结 AI 辅助
研究如何将测试时缩放应用于无人机视觉语言导航,通过迭代细化和多标准评分函数,无需额外训练,使冻结模型自我校正,生成更准确可靠飞行计划,实现最优性能。
中文摘要 AI 辅助
测试时缩放提供了一种无需额外训练即可提高视觉语言模型(VLM)推理性能的有前景方法。现有的无人机视觉语言导航(VLN)方法通常依赖单次推理,在复杂环境中可能产生次优或不安全轨迹。本文探索了一种将测试时缩放应用于无人机VLN的简单有效方法。通过无需额外模型训练的迭代细化过程增强导航推理,引导模型重新评估初始导航计划以提高准确性和安全性。该方法先促使模型生成多个并行候选,然后进行自我校正步骤,在不改变基础模型的情况下实现更深更稳健的规划。还设计了多标准评分函数基于安全、目标对齐和前进进度评估细化候选。此简单而强大的组合使冻结的无人机导航VLM能够自我校正并生成更准确可靠的飞行计划,在此任务中实现了最优性能。
英文摘要
Test-time scaling offers a promising method to improve the inference performance of Vision-Language Models (VLMs) without additional training. Existing approaches to vision-language navigation (VLN) for Unmanned Aerial Vehicle (UAV) typically relies on a single inference pass, which can falter in complex environments by producing suboptimal or unsafe trajectories. In this paper, we explore a simple and effective approach to apply test-time scaling to VLN for UAV. We enhance navigation reasoning through an iterative refinement process that requires no extra model training, guiding the model to re-evaluate its initial navigation plan for better accuracy and safety. Our method first prompts the model to generate multiple parallel candidates and then performs a self-correction step, achieving deeper and more robust planning without changing the underlying model. To further strengthen decision-making, we design a multi-criteria scoring function to evaluate the refined candidates based on safety, goal alignment, and forward-progress. This simple yet powerful combination enables a frozen UAV navigation VLMs to self-correct and generate more accurate and reliable flight plans, achieving SOTA performance in this task.