arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OGAM:通过面向对象的注意力监控将系统测试与运行时保障相连接,用于VLA策略

OGAM: Connecting Systematic Testing to Runtime Assurance through Object-Grounded Attention Monitoring for VLA Policies

Haki Darwish, Xiangyu Yin, Changwen Li, Rongjie Yan, Francisco Gomes de Oliveira Neto, Chih-Hong Cheng

arXiv 2610.05878首次发表:更新:

发表机构

Carl von Ossietzky University of Oldenburg; Huadian Electric Power Research Institute; Chalmers University of Technology(卡尔·冯·奥西茨基大学奥尔登堡分校; 华电电力科学研究院; 查尔姆斯理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

OGAM通过测试揭示的注意力分歧,在运行时监控VLA策略,早期阻止失败,无需失败标签训练,在四种策略中阻止87-100%失败情节。

AI 中文摘要

基准测试仅将视觉-语言-动作(VLA)策略暴露于少数规范指令,而全面的部署测试是不可能的。我们引入了面向对象的注意力监控(OGAM),将系统测试与运行时保障相连接:测试揭示了成功与失败执行之间的注意力分歧,OGAM利用这一信号在有限测试套件之外阻止失败。我们通过动作模板与对象的成对组合生成场景基础指令,并分别测试保持语义的改写。所有87个超出基准的案例在OpenVLA、OpenVLA-OFT、UniVLA和π₀.₅上均表现出问题行为:没有策略完成24个可行指令中的任何一个,而不可行或危险的请求也触发了行为替换。在每个动作查询时,我们通过对象掩码投影梯度加权视觉注意力,并按指令角色分组以跨任务和策略进行比较。动态时间规整将该过程与成功参考对齐,尽管存在速度差异;对成功情节的保形校准设置了持续偏离的早期停止阈值,名义误停目标为α=0.05。在四种策略中,OGAM在20秒预算内以中位时间5-12秒阻止了87-100%的失败情节,观察到的误停率为3-5%,且无需失败标签训练。因此,有限测试识别出注意力模式,支持在失败完全展开之前进行在线干预。

英文摘要

Benchmarks expose vision-language-action (VLA) policies to few canonical instructions, while exhaustive deployment testing is impossible. We introduce Object-Grounded Attention Monitoring (OGAM), connecting systematic testing to runtime assurance: testing reveals attention divergence between successful and failed executions, and OGAM uses this signal to stop failures beyond the finite suite. We generate scene-grounded instructions through pairwise combinations of action templates and objects, and separately test meaning-preserving paraphrases. All 87 out-of-benchmark cases reveal problematic behavior across OpenVLA, OpenVLA-OFT, UniVLA, and $π_{0.5}$: none completes any of the 24 feasible instructions, while infeasible or hazardous requests also trigger behavior substitution. At each action query, we project gradient-weighted visual attention through object masks and group it by instruction role for comparison across tasks and policies. Dynamic time warping aligns this course with a successful reference despite speed differences; conformal calibration on successful episodes sets the early-stopping threshold for sustained deviations, with a nominal false-stop target of $α=0.05$. Across four policies, OGAM stops 87-100% of failed episodes at median times of 5-12s within a 20s budget, with observed false-stop rates of 3-5%, without failure-labeled training. Finite testing thus identifies attention patterns that support online intervention before failure fully unfolds.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑