发表机构
Istituto Superiore di Sanità(意大利高等卫生研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过双网络模型,证明从奖励中可学习非语言编码传递推断规则,编码有效性取决于信息量而非多样性,且误差梯度优于奖励信号。
AI 中文摘要
一个从示例中推断出的规则,如何能在没有共享编码的情况下,传达给从未见过这些示例的人?患有严重失语症的患者通过手势或草图来做到这一点。一个网络看到工作示例并发出八个发明的符号;第二个网络对这些示例视而不见,却将它们应用于新的输入。由于第二个网络的成功而获得奖励,第一个网络学习到一种编码,将规则传递给三步转换,这些转换是训练中从未出现的,而新的学习者能够习得这些转换。与人类发明的语言类似,这种编码有两种机制:仅在奖励作用下,说话者会漂移到一条消息,就像人类语言在纯传播下会丢失词汇一样;表达压力则保持消息的差异化。在新规则上的成功程度取决于消息对规则的描述程度,而非消息的多样性。学习信号塑造了编码:奖励将许多规则归类到少数固定标签下;听者的误差梯度为每个规则提供了相似消息的区域,从而更好地区分规则。
英文摘要
How can a rule inferred from examples reach someone who never saw them, without a shared code? Patients with severe aphasia do it by gesture or sketch. One network sees worked examples and emits eight invented symbols; a second, blind to them, applies them to a new input. Rewarded for the second's success, the first learns a code carrying rules to three-step transformations training never presents, which new learners acquire. Like invented human languages, the code has two regimes: under reward alone the speaker drifts to one message, as human languages lose words under plain transmission; expressive pressure keeps messages differentiated. Success on new rules tracks how much the message says about the rule, not how varied messages are. The learning signal shapes the code: reward sorts many rules under few fixed labels; the listener's error gradient gives each rule a region of similar messages, telling rules apart far better.
Comments14 Pages, 5 Figures