AI 中文总结
本文提出FibVLA框架,通过对数事后采样和流匹配等技术,解决了视觉-语言-动作模型时序信息捕捉与推理效率的矛盾,提升了动作性能与实时响应能力。
AI 中文摘要
视觉-语言-动作模型(VLAs)利用多模态信息认知来推理物理世界的动作,为具身AI应用提供了通用解决方案。传统VLAs通常聚焦于当前数字认知,虽有研究尝试通过捕捉时序信息增强VLAs的推理能力,但对长上下文历史进行编码会导致效率下降的问题。为协调VLAs中捕捉时序信息与保持推理效率之间的矛盾,本文提出FibVLA,这是一种具备长上下文历史时序感知的高效框架。具体而言,我们对本体感受状态和视觉帧采用对数事后采样,以最小冗余捕捉长期时序依赖;对于动作专家,引入流匹配生成动作分布,并采用斐波那契循环推理策略,基于实时闭环反馈生成长程规划步骤。实验表明,FibVLA在不重新训练大规模视觉编码器的情况下,显著提升了动作平滑度和成功率;效率分析显示,在真实世界评估中,其相比基于视频的基线方法具备更出色的实时响应能力。
英文摘要
Vision-language-action models (VLAs), which leverage the cognition of multimodal information to infer physical-world actions, provide a generalized solution for embodied AI applications. Conventional VLAs usually concentrate on current digital cognition. While some efforts are made to enhance VLAs' reasoning capabilities by capturing temporal information, encoding the long-context history causes an efficiency-decreasing issue. To reconcile the conflict between capturing temporal information and maintaining inference efficiency in VLAs, this paper introduces FibVLA, an efficient framework featuring temporal perception of long-context history. Specifically, we leverage logarithmic hindsight sampling to both proprioceptive states and visual frames to capture long-term temporal dependencies with minimal redundancy. For the action expert, we introduce the flow matching to produce action distributions, and the Fibonacci recurrent inference strategy to generate long-range planning steps based on real-time closed-loop feedback. Experiments demonstrate that FibVLA significantly improves action smoothness and success rates without retraining large-scale visual encoders. Efficiency analysis demonstrates superior real-time responsiveness compared to video-based baselines in real-world evaluations.