arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向具身控制的动作与语言条件视频评估

Action- and Language-Conditioned Video Assessment for Embodied Control

Hwanhee Kim, Jaehyun Jang, Seungmin Cha, Hyeonseo Yun, Donghoon Lee, Chang D. Yoo

arXiv 2608.08273首次发表:更新:

发表机构

Korea Advanced Institute of Science and Technology (KAIST); School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院; 韩国科学技术院电气工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出ALVA轨迹评估器,结合视觉观测、动作序列与语言指令评估具身控制任务,在模拟3D家庭环境中误报率低,作为反馈机制可提升闭环策略优化效果。

AI 中文摘要

基于视觉的具身智能体执行多步自然语言指令时,需要能评估完整轨迹任务进展的反馈机制。传统基于最终帧匹配或连续嵌入相似度的方法可能忽略判断指令是否完成所必需的中间过渡。我们提出ALVA(Action- and Language-Conditioned Video Assessment,即动作与语言条件视频评估),这是一种评估视觉观测、执行动作序列和自然语言指令的轨迹评估器。该方法在两个阶段使用预训练的视觉语言模型(VLM):首先基于执行的动作总结帧间视觉过渡,再针对指令评估生成的摘要以输出离散的轨迹级进展分数。在模拟3D家庭环境中,ALVA呈现出保守评估模式,误报率接近零;作为闭环策略优化的终端反馈时,它比被评估的静态图像和基于嵌入的视觉基线提供更有效的反馈,并缩小了与真实值神谕的性能差距。这些结果表明,动作与语言条件视频评估是可解释的反馈机制,适用于所评估的模拟具身控制任务。

英文摘要

Vision-based embodied agents executing multi-step natural language instructions require feedback mechanisms that assess task progress over complete trajectories. Conventional approaches based on final-frame matching or continuous embedding similarity may overlook intermediate transitions that are necessary for determining whether an instruction has been completed. We propose ALVA (Action- and Language-Conditioned Video Assessment), a trajectory evaluator that conditions its assessment on visual observations, the executed action sequence, and the natural language instruction. The method uses a pre-trained vision-language model (VLM) in two stages: it first summarizes frame-to-frame visual transitions conditioned on the executed actions and then assesses the generated summary with respect to the instruction to produce a discrete trajectory-level progress score. In simulated 3D household environments, ALVA exhibits a conservative assessment pattern with near-zero false-positive rates. When used as terminal feedback for closed-loop policy optimization, it provides more effective feedback than the evaluated static image and embedding-based visual baselines and reduces the performance gap to a ground-truth oracle. These results support action- and language-conditioned video assessment as an interpretable feedback mechanism for the evaluated simulated embodied-control tasks.

Comments21 pages, 7 figures. v2: updated Funding section

Journal refSensors 26(16), 5046 (2026)

DOI:10.3390/s26165046

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑