发表机构
Hubei University; Wuhan University; Baidu Inc.(湖北大学; 武汉大学; 百度公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出因果保持概念,通过接口分解与选择性适应(因果核心)使冻结智能体状态准确回答机制探针,实验证明其优于任务充分性。
AI 中文摘要
任务性能并不必然决定智能体保留哪种干预机制。我们研究因果保持:冻结的学习状态是否回答一个独立于训练固定的机制探针映射,包括动作、上下文、直接目标、价值和延迟。对于有限结构因果模型类,最优探针误差是贝叶斯决策风险。当且仅当每个学习接口纤维位于一个探针答案纤维内时,该误差消失;任何通过对该接口进行后处理获得的状态继承相同的下界。后验覆盖定理刻画了预算受限的重新测试,而精确编辑分解表明移位集是无错误目标更新的唯一支持。因果核心通过证据门控写入、读出过滤、时间信用、隐藏上下文设置和局部诊断更新实现这些条件。实验涵盖有限因果系统、连续模拟器、官方TD-MPC2世界模型和Qwen2.5-7B-Instruct。冻结的Qwen最后一层探针在源机制上达到0.958的平衡准确率,但在改变延迟上为0.583;门控机制状态达到1.000,并且仅接受0.056的同步读出候选。在TD-MPC2中,每个执行器的五个目标状态将效果符号准确率从0.057恢复到0.948,而不降低稳定响应。因此,因果保持不同于任务充分性和源域可解码性。
英文摘要
Task performance need not determine which intervention mechanism an agent retains. We study causal retention: whether a frozen learned state answers a mechanism-probe map fixed independently of training, including action, context, direct target, value, and delay. For finite structural causal model classes, the optimal probe error is a Bayes decision risk. It vanishes exactly when every learning-interface fiber lies within one probe-answer fiber; any state obtained by post-processing that interface inherits the same lower bound. A posterior-coverage theorem characterizes budgeted retesting, while an exact edit decomposition shows that the shifted set is the unique support of an error-free target update. Causal Core implements these conditions through evidence-gated writing, readout filtering, temporal credit, hidden-context setup, and local diagnostic updates. Experiments cover finite causal systems, continuous simulators, an official TD-MPC2 world model, and Qwen2.5-7B-Instruct. A frozen Qwen last-layer probe reaches 0.958 balanced accuracy on source mechanisms but 0.583 on changed delays; the gated mechanism state reaches 1.000 and accepts only 0.056 of synchronized-readout candidates. In TD-MPC2, five target states per actuator recover effect-sign accuracy from 0.057 to 0.948 without degrading stable responses. Causal retention is therefore distinct from task sufficiency and source-domain decodability.
Comments34 pages, 4 figures, 6 tables, and 1 algorithm