arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

参差不齐的能力从何而来?

Where Does Jagged Competence Come From?

Ioannis Tsiokos

arXiv 2610.03831首次发表:更新:

AI 中文总结

本研究通过一个两层任务探究参差不齐能力的来源,发现网络学习高层任务时对低层特征获取不均,导致能力参差,且部分失败无法解释。

AI 中文摘要

能力强的系统常常表现出参差不齐的能力:平均误差低,但在特定输入上却会失败。我们探究这种能力在一个由两个已知层构建的任务中源自何处。较低层A从记录中计算五个每槽位和;较高层B使用槽位1的和以及从先前板卡延续下来的模式来预测下一板卡的类别,因此B无法仅从当前的A计算得出:B是A的严格扩展。我们在B上训练小型循环网络,并对照精确真值测量它们获得了多少A。核心发现是,学习B使网络获得了A的参差不齐版本,这通过探针可读性和监督输出来衡量。在B需要它的地方(槽位1,线性探针达到97-98%)它是可读的,而在B不需要的地方(其他槽位,5-10%)可读性降低;类别误读集中在槽位1的截止点附近;当A被显式训练时,它只是近似正确。网络在B上表现出参差不齐的能力:自然KL低于$3\ imes10^{-4}$比特,与构造历史上最大全变差约0.25共存。匹配实验表明,即使是一个好的A也不够:将学习到的A连接到B可将误读减少1.3到17倍,相同的精确A在独热输入(0-2)下比数值输入(10-317)产生更少的误读,一次B更新会使精确A变得不精确,除非B梯度被阻止,并且精确A仍然留下一些B失败。在这个任务中,较低理论的不均匀获取、访问和保存解释了学习构建于其上的理论的网络的部分参差不齐能力;一些失败仍未解释。在更大系统中是否同样如此,是一个有待检验的测试。

英文摘要

Capable systems often show jagged competence: low average error alongside failures on particular inputs. We ask where it comes from in a task built from two known layers. A lower layer A computes five per-slot sums from records; an upper layer B uses the slot-1 sum and a mode carried over from earlier boards to predict the next board's category, so B cannot be computed from the current A alone: B is a strict extension of A. We train small recurrent networks on B and measure, against exact ground truth, what they acquire of A. The central finding is that learning B gives the network a jagged version of A, measured through probe readability and supervised outputs. It is readable where B needs it (slot 1, 97-98% by a linear probe) and becomes less readable where B does not (the other slots, 5-10%); category misreads concentrate near the slot-1 cutoffs; and when A is trained explicitly it comes out only approximately right. The networks show jagged competence in B: natural KL below $3\times10^{-4}$ bits coexists with a maximum law TV of about 0.25 on constructed histories. Matched experiments show that even a good A is not enough: connecting a learned A to B cuts misreads 1.3 to 17-fold, the same exact A gives fewer misreads as one-hot inputs (0-2) than as numerical inputs (10-317), one B update makes an exact A inexact unless the B gradient is blocked, and exact A still leaves some B failures. In this task, uneven acquisition, access and preservation of the lower theory explain part of the jagged competence of a network that learns the theory built on it; some failures remain unexplained. Whether the same holds in larger systems is a test to run.

Comments30 pages, 10 figures. Code: https://github.com/ioannist/six-birds-ml

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑