arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

何时一个学习到的命令适配器是值得的?冻结运动策略的闭环识别与反事实审计

When Is a Learned Command Adapter Worth It? Closed-Loop Identification and Counterfactual Auditing of Frozen Locomotion Policies

Zongtan Li

arXiv 2607.21867首次发表:更新:

AI 中文总结

研究向冻结运动策略添加学习到的命令适配器是否值得,通过适配器必要性审计区分多种增益并映射到决策,经闭环识别等方法及实验评估,测试可观测信号能否证明依赖状态的自适应,而非预设适配器有价值。

AI 中文摘要

仅当接口展现出可从部署时观测中真实恢复的改进时,向冻结的、命令条件化的运动策略添加学习到的适配器才是值得的。我们引入了一种适配器必要性审计,它区分全局工作点增益、同状态反事实余量、相对于交叉拟合固定动作的部署增益以及相对于频率匹配随机策略的状态分配增益。源集群学习器重新拟合将这些量和约束违规映射到一个是/否/弃权决策。闭环命令响应识别提供可选决策特征。在Go2上,一个存档的比例前缀诊断发现5.2%的同状态余量,但只有0.55%的恢复分配增益。我们的验证性审计针对由直接控制、VGCC和MPC诱导的三种查询分布,在二十个独立集群上评估直接、比例、航向和偏航干预,使用200次完整的学习器重新拟合。在1%的部署和分配阈值以及5%的违规容忍度下,直接查询返回否,而VGCC和MPC查询弃权。VGCC具有最大的平均部署增益(1.34%),但其分配下限为0.09%,违规上限为6.25%。一个具有部署代表性的二十集群H1审计也返回否,而学习器级别的综合控制返回是。因此,该审计测试可观测信号是否证明依赖状态的自适应是合理的,而不是假定适配器是有价值的。

英文摘要

Adding a learned adapter to a frozen, command-conditioned locomotion policy is worthwhile only if the interface exposes improvements that are both real and recoverable from deployment-time observations. We introduce an adapter necessity audit that separates global operating-point gain,same-state counterfactual headroom, deployment gain over a cross-fitted fixed action, and state-allocation gain over a frequency-matched randomized policy. Source-cluster learner refits map these quantities and constraint violations to a GO/NO-GO/ABSTAIN decision. Closed-loop command- response identification provides optional decision features. On Go2, an archived scale-prefix diagnostic finds 5.2% same-state headroom but only 0.55% recovered allocation gain. Our confirmatory audit evaluates direct, scale, heading, and yaw interventions on twenty independent clusters for each of three query distributions induced by direct control, VGCC, and MPC, using 200 full learner refits. At 1% deployment and allocation thresholds and a 5% violation tolerance, direct queries return NO-GO, while VGCC and MPC queries ABSTAIN. VGCC has the largest mean deployment gain (1.34%), but its allocation lower bound is 0.09% and its violation upper bound is 6.25%. A deployment-representative twenty-cluster H1 audit also returns NO-GO, whereas a learner-level synthetic control returns GO. The audit therefore tests whether observable signal justifies state-dependent adaptation rather than presuming that an adapter is valuable.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑