LAWFUL:面向隐层忠实使用的法律对齐见证者
LAWFUL: Law-Aligned Witness for Faithful Use of Latents
浏览论文内容
中文总结 AI 辅助
本研究针对神经网络预测物理系统时的可解释性缺口,提出LAWFUL框架,以验证神经网络是否学习并内部使用特定物理定律,相关内容在Mocap2Radar transformer上得到验证。
中文摘要 AI 辅助
当神经网络能准确预测物理系统时,它是否已将支配性定律学习为形式化结构化知识?若是如此,网络的内部计算是否在该定律的整个有效域内实际使用了这种表示?我们确定了四个可解释性缺口,这些缺口限制了对连续变量上物理定律的上述问题的回答:缺乏针对连续反事实的覆盖感知因果一致性度量、缺乏对已识别电路的有效域测试、缺乏对定律不变量和禁止行为的验证、缺乏对派生物理量如何在电路中流动的量化。我们开发了基础框架LAWFUL,它弥补了前两个缺口,并为后两个奠定了基础,我们在Mocap2Radar transformer上对其进行了说明,验证它是否从运动捕捉和雷达数据中学习并内部使用多普勒频率定律$f(t) = \frac{2 v(t)}{\lambda}$,而这些数据中既未出现$f(t)$也未出现$v(t)$。
英文摘要
When a neural network predicts a physical system accurately, has it learned the governing law as formal, structured knowledge, and if so, does the network's internal computation actually use that representation throughout the law's domain of validity? We identify four interpretability gaps that limit answering these questions for {\em physics laws over continuous variables}: the absence of a coverage-aware causal-consistency measure over continuous counterfactuals; of a domain-of-validity test for the identified circuit; of a verification of the law's invariants and forbidden behaviors; and of a quantification of how a derived physical quantity flows through the circuit. We develop a foundational framework, LAWFUL, that closes the first two and lays groundwork for the remaining two, and illustrate it on the Mocap2Radar transformer, validating whether it learns and internally uses the Doppler frequency law $f(t) = \frac{2 v(t)}λ$ from motion-capture and radar data in which neither $f(t)$ nor $v(t)$ appears.