arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过预测不确定性触发通用推理的双系统视觉-语言-动作模型

Triggering Generalist Reasoning via Predictive Uncertainty for Dual-System VLA

Hyemin Yang, Wooseong Jeong, Giwon Lee, Kuk-Jin Yoon

arXiv 2610.05025首次发表:更新:

发表机构

KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出TUD自适应推理框架,通过预测不确定性信号动态跳过不必要的通用模型调用,在匹配成功率下大幅减少计算开销,实现更优的成本-成功权衡。

AI 中文摘要

双系统视觉-语言-动作(VLA)模型通过将具备推理能力的慢速通用模型与快速专家动作模型配对,提升了实时机器人控制性能。然而,现有方法以固定频率调用通用模型,忽略了整个任务执行过程中决策复杂度的变化。这种静态策略在简单阶段浪费计算资源,并在场景意外变化时可能延迟重新推理。我们提出TUD(通过预测不确定性触发通用推理的双系统VLA),一种自适应推理框架,可选择性地跳过不必要的通用模型调用。TUD在缓存的通用模型上下文下,测量即将到来的块槽位上动作重新预测的跨步离散度,作为预测不确定性信号。该信号捕捉了随着新观测到达未来动作计划的变化程度,并且仅利用架构已运行的前向传播计算,无需手动阶段标签或辅助不确定性模型。在VLA-Arena上,在匹配调用预算下,TUD相比其他不确定性基线实现了更高的成功率,同时保持较低墙钟开销,并能更一致地区分成功与失败的任务轨迹。此外,TUD相比非自适应基线找到了更有利的成本-成功权衡,通过变化单一阈值绘制出完整操作曲线,并在匹配成功率下大幅减少VLM调用。同样的权衡也出现在我们的真实机器人实验中,TUD相比最强固定间隔基线减少了75%的通用模型调用,同时实现了更高的成功率。我们的结果表明,预测不确定性为高效VLA控制中的自适应推理提供了实用准则。

英文摘要

Dual-system Vision-Language-Action (VLA) models improve real-time robotic control by pairing a slow, reasoning-capable generalist with a fast specialist action expert. However, existing methods invoke the generalist at a fixed frequency, ignoring the fact that decision-making complexity varies throughout a rollout. This static strategy wastes computation in easy phases and can delay renewed reasoning when the scene changes unexpectedly. We propose TUD (Triggering generalist reasoning via predictive Uncertainty for Dual-system VLA), an adaptive inference framework that selectively skips unnecessary generalist calls. TUD measures the cross-step dispersion of action re-predictions at the upcoming chunk slot under the cached generalist context, as a predictive uncertainty signal. This signal captures how much the future action plan shifts as new observations arrive and is computed from forwards the architecture already runs, requiring neither manual phase labels nor an auxiliary uncertainty model. On VLA-Arena, it achieves a higher success rate at matched call budgets than alternative uncertainty baselines while maintaining low wall-clock overhead, and more consistently separates successful from failed rollouts. Also, TUD finds a more favorable cost-success trade-off than non-adaptive baselines, tracing an entire operating curve as a single threshold is varied, and substantially reduces VLM calls at matched success rate. The same trade-off appears in our real-robot experiments, where TUD cuts generalist calls by 75% relative to the strongest fixed-interval baseline while achieving an even higher success rate. Our results suggest that predictive uncertainty provides a practical criterion for adaptive reasoning in efficient VLA control.

CommentsAccepted by NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑