发表机构
The University of Osaka; Nanyang Technological University; The University of Tokyo(大阪大学; 南洋理工大学; 东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出事件驱动的主动辅助框架,利用视觉-语言模型从人-物交互结果推断任务上下文并生成辅助动作,无需用户指令,在真实桌面任务中达到与指令驱动方法相当的性能。
AI 中文摘要
协作操作中的辅助通常由用户指令发起,使得高层推理成为请求驱动。然而,在流畅的人类团队合作中,伙伴往往从观察到的动作结果推断下一步有帮助的动作,而不是等待指令。受此启发,我们研究了一种事件驱动的主动辅助公式,其中人-物交互结果在推理时无需用户提供的任务规范即可启动辅助推理。为此,我们提出一个事件驱动框架,通过事件监视器监控工作区状态变化,并在事件完成后提取稳定的前后快照,以表征产生的状态转换。一个冻结的预训练视觉-语言模型(VLM)利用其语义先验推断任务上下文,决定辅助是否合适,并在需要时从观察到的转换生成一系列辅助动作。为了使输出可执行和可验证,我们将动作限制为一组动作原语,并通过整数引用对象。我们在三个不同的真实世界桌面协作任务上评估同一框架,无需特定任务训练或微调。事件驱动框架实现了与给定用户指令的变体相当的性能。
英文摘要
Assistance in collaborative manipulation is often initiated by user instructions, making high-level reasoning request-driven. In fluent human teamwork, however, partners often infer the next helpful step from the observed outcome of an action rather than waiting for instructions. Motivated by this, we investigate an event-driven formulation of proactive assistance, where human--object interaction outcomes initiate assistive reasoning without user-provided task specifications at inference time. To this end, we propose an event-driven framework that monitors workspace state changes with an event monitor and, upon event completion, extracts stabilized pre/post snapshots that characterize the resulting state transition. A frozen pretrained Vision-Language Model (VLM) then uses its semantic priors to infer the task context, decide whether assistance is appropriate, and, when needed, generate a sequence of assistive actions from the observed transition. To make outputs executable and verifiable, we restrict actions to a set of action primitives and reference objects via integer IDs.We evaluate the same framework across three distinct real world tabletop collaboration tasks without task-specific training or fine-tuning. The event-driven framework achieves performance comparable to variants given user instructions.