arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19475cs.RO

FASA:面向高效扩散式视觉-语言-动作模型的反馈感知采样自适应

FASA: Feedback-Aware Sampling Adaptation for Efficient Diffusion-Based VLA Models

Yuchen Han, Jianhan Wu, Xiaoyang Qu, Lingwei Kong, Shiyi Li, Jianzong Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出FASA,一种无需训练的运行时框架,通过利用实时多模态反馈动态调整扩散式VLA模型的采样步数,在保持成功率的同时将推理速度提升高达1.45倍,适用于资源受限平台。

中文摘要 AI 辅助

基于扩散的视觉-语言-动作(VLA)模型在具身任务中表现出强大的性能,但其迭代采样过程带来了沉重的计算和内存访问开销,阻碍了在边缘平台上的实时部署。现有的加速方法要么需要昂贵的训练(例如,蒸馏、流匹配),要么通过静态调度的剪枝和缓存来降低感知质量,忽略了机器人交互中动态的工作负载变化。本文提出了FASA(反馈感知采样自适应),一个无需训练的运行时框架,将实时多模态反馈视为去噪管道的控制信号:一个交互驱动的范围适配器根据视觉和夹爪力反馈调节全局采样步数预算,一个本体感觉感知的步数适配器在适配范围内精确定位优化步数。这种协同设计的框架使得底层硬件架构能够自适应地匹配不同执行阶段的工作负载需求。在多个基准上的比较评估表明,推理速度可提升高达1.45倍,同时保持具有竞争力的成功率,为将繁重的生成式具身AI工作负载部署到资源受限的计算平台上提供了一种新颖的动态运行时架构范式。

英文摘要

Diffusion-based Vision-Language-Action (VLA) models achieve strong performance in embodied tasks, but their iterative sampling imposes heavy computational and memory-access cost, blocking real-time deployment on edge platforms. Existing acceleration methods either require expensive training (e.g., distillation, flow matching) or degrade perception via statically scheduled pruning and caching, ignoring the dynamic workload variance of robotic interactions. This paper presents FASA (Feedback-Aware Sampling Adaptation), a training-free runtime framework that treats real-time multimodal feedback as a control signal for the denoising pipeline: an interaction-driven range adaptor modulates the global sampling-step budget based on visual and gripper-force feedback, and a proprioception-aware step adaptor pinpoints the optimized step within the adapted range. This co-designed framework allows the underlying hardware architecture to adaptively match the workload demands of different execution phases. Comparative evaluations across several benchmarks show that the inference speed can be increased by up to 1.45$\times$ while maintaining competitive success rates, providing a novel dynamic runtime architecture paradigm for deploying heavy generative embodied AI workloads onto resource-constrained computing platforms.

发表机构

  • South China University of Technology(华南理工大学)
  • Ping An Technology (Shenzhen) Co., Ltd.(平安科技(深圳)有限公司)
  • Harbin Institute of Technology(哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑