From Simulation to Enaction: Post-trained language models recognize and react to their own generations
从模拟到行动:后训练语言模型识别并回应自身生成
机构 * Institute for Advanced Study, Princeton(普林斯顿高级研究院) ; Anthropic
AI总结 本文发现后训练语言模型能够识别自身生成(on-policy)并降低输出熵,通过内部表示输入意外性来调节,且显式识别与隐式识别机制不同。
Comments Anthropic fellows project mentored by Jack Lindsey