通过机制可解释性和最优控制将鲁棒性引入世界行动模型
Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究世界行动模型在分布转移下的脆弱性,利用机制可解释性通过对比激活方向和基于模型的最优控制实现WAM转向,产生WA-LQR,预测不同模型转向能力,在部分模型上提高了对多种扰动的鲁棒性。
AI中文摘要:
世界行动模型(WAMs)能实现语义和物理信息控制,但在分布转移下很脆弱。本文利用机制可解释性研究WAM激活空间中与鲁棒性相关的扰动如何表示。通过比较成功和失败展开的激活情况,发现一些WAM架构对关键鲁棒性特征具有低维线性可分性,从而推动使用对比激活方向进行无训练的WAM转向。还表明WAM激活动力学中的局部线性性可通过基于模型的最优控制实现高效反馈转向,产生世界行动线性二次调节器(WA-LQR)。通过机制评估,预测了不同模型的转向能力,在Cosmos-Policy和DiT4DiT上,WA-LQR将对比方向推广到新任务并提高了对多种扰动的鲁棒性。
英文摘要:
World Action Models (WAMs) enable semantically- and physically-informed control but are brittle under distribution shift. In this work, we use mechanistic interpretability to study how robustness-relevant perturbations are represented in WAM activation space. Comparing activations across successful and unsuccessful rollouts, we find some WAM architectures exhibit low-dimensional linear separability for robustness-critical features, while others do not. This motivates the use of contrastive activation directions for training-free WAM steering. We also show that local linearity in WAM activation dynamics enables efficient feedback steering via model-based optimal control, yielding World-Action Linear Quadratic Regulator (WA-LQR), a minimally-invasive reduced-order LQR controller. Via mechanistic evaluations, we predict strong steerability in the Cosmos-Policy and DiT4DiT models but weak steerability in LingBot-VA, consistent with steering intervention results. On Cosmos-Policy and DiT4DiT, WA-LQR generalizes contrastive directions to new tasks and improves robustness to camera, gripper, and visual-noise perturbations over unsteered and prompt steering baselines.