arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37530cs.ROcs.CV

RawVLA:面向机器人操作的具体化神经图像信号处理器

RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation

Shuhong Liu, Heng Zhou, Lingfeng Qian, Yuhao Fang, Xianbao Hou, Qianyu Zhou, Lin Gu, Wei Sui, Jianfei Yang, Ziteng Cui

首次发表
浏览论文内容

中文总结 AI 辅助

针对VLA模型忽略成像流水线的问题,提出RawVLA神经ISP,自适应处理RAW图像以提升操作鲁棒性,并构建RawVLA-Bench基准验证其有效性。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型通常处理由固定相机图像信号处理器(ISP)生成的RGB图像,从而将成像流水线置于学习和评估循环之外。我们系统地考察了这一被忽视的设计选择在五个基本ISP维度上的后果:增益、传感器噪声、色度响应、色调响应和位深。我们的分析表明,RAW到RGB的处理过程实质性地影响动作预测和操作成功率,且不同ISP维度产生显著不同的效应。基于这些发现,我们提出了RawVLA,一种流式神经ISP,能够为冻结的VLA策略自适应地渲染RAW观测,同时将其容量集中于与具体化行为相关的成像因素。我们进一步推出了RawVLA-Bench,一个RAW域操作基准,旨在将图像处理作为清洁和恶劣采集条件下的显式评估变量。在RawVLA-Bench上的实验表明,RawVLA在标准条件下保持性能,同时在退化成像条件下大幅提升鲁棒性,从而确立自适应RAW处理作为物理相机与具体化策略之间有效接口的地位。

英文摘要

Vision-language-action (VLA) models typically operate on RGB images produced by a fixed camera image signal processor (ISP), leaving the imaging pipeline outside the learning and evaluation loop. We systematically examine the consequences of this overlooked design choice across five fundamental ISP dimensions: gain, sensor noise, chromatic response, tonal response, and bit depth. Our analysis reveals that RAW-to-RGB processing materially shapes both action prediction and manipulation success, with different ISP dimensions exerting substantially different effects. Guided by these findings, we introduce RawVLA, a streaming neural ISP that adaptively renders RAW observations for frozen VLA policies while concentrating its capacity on the imaging factors relevant to embodied behavior. We further present RawVLA-Bench, a RAW-domain manipulation benchmark to expose image processing as an explicit evaluation variable across clean and adverse acquisition conditions. Experiments on RawVLA-Bench show that RawVLA preserves performance under standard conditions while substantially improving robustness under degraded imaging, establishing adaptive RAW processing as an effective interface between physical cameras and embodied policies.

发表机构

  • UTokyo(东京大学)
  • D-Robotics(地瓜机器人)
  • TohokuU(东北大学)
  • NTU(南洋理工大学)
  • HKUSTGZ(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

↑