发表机构
Karlsruhe Institute of Technology; Hunan University; Mercedes-Benz(卡尔斯鲁厄理工学院; 湖南大学; 梅赛德斯-奔驰)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RoboFind提出多智能体框架,结合智能手机教学与四足机器人搜索,通过目标教授、导航、验证和协调智能体,实现盲人个性化物体搜索,成功率85%,显著优于基线。
AI 中文摘要
盲人和低视力用户通常需要定位一个特定的个人物品,而不是同类别的任意实例。这一任务要求机器人能够在空间中移动并到达用户无法到达的视角,同时需要一个无障碍界面,让用户说出目标物体并了解是否找到了正确的物品。我们提出了RoboFind,一个多智能体框架,其中智能手机教授目标,四足机器人执行搜索。目标教授智能体通过一个包含AR引导、语音和触觉反馈以及屏幕阅读器支持的无障碍采集流程,将引导式智能手机录制转换为语义目标档案和可复用的多视角参考库,因此后续任务可以引用存储的物体而无需重复教授过程。在运行时,导航智能体探索环境并提出候选目标,验证智能体根据存储的参考检查每个候选,协调与恢复智能体完成任务或触发恢复并继续搜索。在32次真实机器人任务中,RoboFind达到了85.0%的成功率,而重建的按顺序首停基线在20次试验中针对十个目标仅为25.0%,并将错误成功从75.0%降至5.0%。在六个共享目标上,它在12次试验中成功10次,而12次独立执行的仅GPT-6 Astra试验为5/12。这些结果表明,多智能体设计符合个性化物体搜索的需求,在宣布完成之前验证物体身份是使结果可靠的关键。
英文摘要
Blind and low-vision users often need to locate a specific personal object rather than an arbitrary instance of the same category. The task calls for a robot that can move through the space and reach viewpoints the user cannot, and for an accessible interface where the user says which object is meant and learns whether the right one was found. We present RoboFind, a multi-agent framework in which a smartphone teaches the target and a quadruped robot carries out the search. A Target Teaching Agent converts guided smartphone recordings into a semantic target profile and a reusable multi-view reference bank through an accessible capture flow with AR guidance, speech and haptic feedback, and screen-reader support, so later missions refer to a stored object without repeating the teaching process. At runtime, a Navigation Agent explores the environment and proposes candidate targets, a Verification Agent checks each candidate against the stored references, and a Coordination and Recovery Agent completes the mission or triggers recovery and continued search. Across 32 real-robot missions, RoboFind reaches 85.0% success against 25.0% for a reconstructed sequential first-stop baseline over 20 trials with ten targets, and reduces false success from 75.0% to 5.0%. On six shared targets it succeeds in 10/12 trials, against 5/12 for 12 independently executed GPT-6 Astra-only trials. These results show that the multi-agent design fits the demands of personalized object search, where verifying object identity before declaring completion is what makes the outcome something a user can rely on.