发表机构
Waseda University; Boston University(早稻田大学; 波士顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态大语言模型智能体在辅助导航中缺乏时间感知的问题,提出带原因预测监督的简单有效修改,在开环、闭环及仿真到现实泛化中显著提升性能。
AI 中文摘要
交互式视觉与语言智能体能否不仅学会说什么,还能学会何时说?当前的语言模型很少对是否以及何时向用户实现实时响应进行规划。然而,为人类决策提供准确且及时的支持,例如在引导视障人士穿越城市环境时,需要仔细的实时响应——时机不当的响应可能会分散用户的注意力或增加不必要的认知负担。作为基于多模态大语言模型(MLLM)的智能体的机器智能挑战,我们引入了一个大规模多模态基准,用于复杂户外环境中的自我中心辅助导航任务。利用该基准,我们发现现成的MLLM在提供安全且时间敏感的导航指令方面存在根本性局限,即使使用大量数据进行模型微调也是如此。随后,我们证明了一种简单而有效的模型修改方法,包括直接监督以预测每条指令背后的原因,在开环、闭环以及仿真到现实泛化设置中均能带来显著的性能提升。然而,我们的分析突显了在时间推理、安全关键物体感知以及关系与距离理解方面的持续挑战。为了推动可扩展辅助智能体的发展,我们将发布我们的仿真环境、基准和代码(可在项目网站上获取:此HTTPS URL)。
英文摘要
Can interactive vision-and-language agents learn not just what to say but also \textbf{\textit{when}} to say it? Current language models rarely plan over whether and when to realize a real-time response to a user. However, providing accurate and timely support for human decision-making, such as when guiding visually impaired individuals through urban environments, requires careful real-time responsiveness--poorly timed responses can distract users or add unnecessary cognitive load. As a machine intelligence challenge for Multimodal Large Language Model (MLLM)-based agents, we introduce a large-scale multimodal benchmark for an egocentric, assistive navigation task in complex outdoor environments. Using this benchmark, we uncover a fundamental limitation of off-the-shelf MLLMs in delivering safe and time-sensitive navigation instructions, even with model fine-tuning on substantial amounts of data. We then demonstrate that a simple yet effective modification of the model, including direct supervision to predict the underlying reason for each instruction, yields significant performance gains across open-loop, closed-loop, and sim-to-real generalization settings. However, our analysis highlights persistent challenges in temporal reasoning, safety-critical object awareness, and relational and distance understanding. To advance the development of scalable assistive agents, we will release our simulation, benchmark, and code (available at the project website: https://timeli-icra.github.io/).
CommentsICRA 2026