arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13924cs.RO

BICPO-VLA:用于平滑异步视觉-语言-动作控制的行为识别延续偏好优化

BICPO-VLA: Behavior-Identified Continuation Preference Optimization for Smooth Asynchronous Vision-Language-Action Control

Ming Shang, Yuchen Huang, Jiaoyang Chen, Haoyuan Hu, Han Yu, Liping Song, Luyun Feng, Shuo Bao, Wei Dong, Xinzhou Wang, Fuchun Sun

首次发表
浏览论文内容

中文总结 AI 辅助

BICPO-VLA针对异步视觉-语言-动作控制的请求交接间隙问题,通过行为识别、动作分解重建、参考相对Flow-DPO优化,实现平滑控制。

中文摘要 AI 辅助

请求到交接的间隙存在三个相互关联的来源:请求时预期行为的模糊性、动作生成过程中累积的物理状态漂移,以及新动作最终接管时的残留不兼容性。BICPO-VLA依次解决这些问题:首先,指令感知因果历史编码器识别命令和当前任务进度所支持的行为;其次,序列Haar子空间生成将每个动作块分解为互补的成对支架和残差系数,实现两个专门的生成阶段后再精确重建,通过减少原始动作空间中的迭代细化,缩短机器人在新块可用前持续移动的间隔;最后,BICPO将已知的传出动作滚动到实际交接状态,并在行为匹配的候选中应用参考相对Flow-DPO,在不改变动作预期行为的前提下,使生成的块适应剩余的请求到交接不匹配。

英文摘要

The request-to-handoff gap has three coupled sources: ambiguity about the behavior intended at request time, physical-state drift accumulated during action generation, and residual incompatibility when the new action finally assumes control. BICPO-VLA addresses them in sequence. First, an instruction-aware causal history encoder identifies the behavior supported by the command and current task progress. Second, sequential Haar subspace generation decomposes each action chunk into complementary pairwise scaffold and residual coefficients, enabling two specialized generation stages followed by exact reconstruction. By reducing iterative refinement in the original action space, it shortens the interval over which the robot continues moving before the new chunk becomes available. Finally, BICPO rolls the known outgoing actions to the actual handoff state and applies reference-relative Flow-DPO among behaviorally matched candidates, adapting the generated chunk to the remaining request-to-handoff mismatch without changing its intended behavior.

补充信息

↑