arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UESF-Bench:统一的具身寻找与跟随的基准测试与探究

UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following

Kun Yu, Jianhua Yang, Yixiang Chen, Changwei Wang, Hongyuan Yu, Yan Huang, Fushuo Huo, Ya Jing, Zhumin Chen, Keji He

arXiv 2607.13621首次发表:更新:

发表机构

Shandong University; Institute of Automation, Chinese Academy of Sciences; Qilu University of Technology; Xiaomi Inc; Beijing University Of Technology; Hong Kong Polytechnic University(山东大学; 中国科学院自动化研究所; 齐鲁工业大学; 小米公司; 北京工业大学; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现有具身智能体语言引导人类跟随基准测试的局限,引入UESF-Bench基准,提出SeekFollow-VLA框架,可处理语义引导探索等任务,实验显示该框架在单人和多人环境中比基线有明显改进,为统一具身寻找与跟随建立基线。

AI 中文摘要

语言引导的人类跟随是具身智能体的一项重要能力,但现有基准测试通常假设目标人物在情节开始时是可见的。这种设置简化了问题,忽略了一个更现实的要求:智能体通常需要先找到语言描述的目标,然后在动态环境中持续跟随该目标。近期工作虽已开始研究人类搜索,但现有设置通常在特定任务场景中评估,且往往依赖更强的环境先验知识。此外,它们通常将搜索和跟随视为 separate 任务,仍缺乏用于系统评估的统一基准。为解决这些限制,我们引入了统一的具身寻找与跟随基准测试(UESF-Bench),这是一个用于具身人类寻找与跟随的大规模多样化基准测试。该基准要求智能体处理语义引导的探索、可靠的行为切换和恢复以及延迟的身份定位。为此,我们提出了 SeekFollow-VLA,这是一个具有任务驱动路由机制的视觉-语言-行动框架,用于在寻找和跟随之间进行潜在阶段推理和转换建模。实验结果表明,SeekFollow-VLA 在单人和多人环境中均比单头和双头基线有明显改进,为统一的具身寻找与跟随建立了基线。

英文摘要

Language-guided human following is an important capability for embodied agents, but existing benchmarks typically assume that the target person is visible at the start of an episode. This setting simplifies the problem and overlooks a more realistic requirement: an agent often needs to first find a language-described target and then persistently follow that target in a dynamic environment. While recent work has started to study human search, existing settings are typically evaluated in task-specific scenarios and often rely on stronger prior knowledge of the environment. Moreover, they usually treat searching and following as separate tasks and still lack a unified benchmark for systematic evaluation. To address these limitations, we introduce the Unified Embodied Seeking and Following Benchmark (UESF-Bench), a large-scale and diverse benchmark for embodied human seeking and following. The benchmark requires agents to handle semantic-guided exploration, reliable behavior switching and recovery, and delayed identity grounding. To this end, we propose SeekFollow-VLA, a vision-language-action framework with a task-driven routing mechanism for latent phase inference and transition modeling between seeking and following. Experimental results show that SeekFollow-VLA achieves clear improvements over both single-head and dual-head baselines across single-person and multi-person environments, establishing a baseline for unified embodied seek-and-follow.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑