面向长视距视觉-语言-动作操作的神经符号过程推理
Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation
- Fraunhofer IPK(弗劳恩霍夫生产系统与设计技术研究所)
- Technische Universität Berlin(柏林工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对长视距VLA操作的脆弱性,提出结合VLA控制、任务图与多模态过程记忆的神经符号框架,利用伪注视指导微调,在两操作领域验证了结构化推理与视觉指导的互补作用。
AI中文摘要:
视觉-语言-动作(VLA)模型可执行短时操作技能,但在需要持续任务状态、依赖感知推理、条件决策及可靠 grounding 的长视距过程中仍表现脆弱。我们研究一种神经符号框架,结合学习到的 VLA 控制与显式任务图及多模态过程记忆。任务图编码动作依赖、有效转换及分支条件,而记忆则维护活跃步骤、已完成动作、文本上下文及任务相关视觉证据。这些结构共同指导物体选择、目标定位、子目标分配及预期状态转换验证。人类演示通过注视或显著性线索提供额外时空指导。为分离其对策略学习的影响,我们的初始研究绕过跨视角注视迁移,直接在机器人视角遥操作视频中注释伪注视。所得指导用于 VLA 微调及推理。我们研究两个长视距操作领域: workspace 清理及手术器械处理,二者均需有序执行、视觉 grounding 决策及条件分支。我们评估正确物体与目标选择、子任务完成、任务进度、步骤顺序一致性、完整任务成功率及过程或执行错误。本研究将结构化符号推理与演示衍生的视觉指导定位为可靠长视距 VLA 操作的互补机制。
英文摘要:
Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon procedures requiring persistent task state, dependency-aware reasoning, conditional decisions, and reliable grounding. We investigate a neuro-symbolic framework that combines learned VLA control with explicit task graphs and multimodal procedural memory. Task graphs encode action dependencies, valid transitions, and branch conditions, while memory maintains the active step, completed actions, textual context, and task-relevant visual evidence. Together, these structures guide object selection, destination grounding, subgoal dispatch, and verification of expected state transitions. Human demonstrations provide additional spatial and temporal guidance through gaze or saliency cues. To isolate their effect on policy learning, our initial study bypasses cross-view gaze transfer and directly annotates pseudo-gaze in robot-view teleoperation videos. The resulting guidance is used during VLA fine-tuning and inference. We study two long-horizon manipulation domains, workspace clearing and surgical-instrument handling, which require ordered execution, visually grounded decisions, and conditional branching. We evaluate correct-object and destination selection, subtask completion, task progress, step-order consistency, complete-task success, and procedural or execution mistakes. This work positions structured symbolic reasoning and demonstration-derived visual guidance as complementary mechanisms for reliable long-horizon VLA manipulation.