arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SanSi:一种用于系统1.5思维的循环类型化决策模型

SanSi: A Looped Typed Decision Model for System 1.5 Thinking

Shuyu Gan, Young-Jun Lee, Dongyeop Kang

arXiv 2610.07730首次发表:更新:

发表机构

University of Minnesota(明尼苏达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SanSi将预训练循环语言模型转为类型化决策模型,通过多次循环修正隐藏状态实现系统1.5思维,在10,027个测试决策上达到72.0%准确率,并扩展可解深度,作为评判者提升生成器F1达7.7个百分点。

AI 中文摘要

类型化决策模型在不生成文本的情况下回答一个声明的问题:决策头在单次前向传播中为每个声明的选项返回一个概率。单次传播是快速的、直觉性的系统1思维。我们研究了介于单次传播和生成式推理之间的情况:循环,即相同的层在类型化读出之前被递归应用多次。每次循环都让模型在提交答案之前修正其隐藏状态,而不生成任何词元;我们称之为系统1.5思维。我们提出了SanSi,它将一个预训练的循环语言模型转变为类型化决策模型。选项概率在每次循环后被读取,并且每次循环都用适当的评分规则进行训练,因此一个模型在一次运行中即可服务于从一次循环到八次循环的任何预算。在来自59个来源的10,027个测试决策上,SanSi达到了72.0%的准确率:比采用相同训练方法、相同形状的非循环模型高出13.5个百分点,比同尺寸的更新的非循环模型高出5.3个百分点,比参数为其三倍的模型低1.8个百分点。在两个深度控制任务上,循环将可解的深度扩展到训练中见过的深度之外,而更大的单次传播模型则失败。当用作强化学习策略优化的评判者时,在没有黄金答案的情况下,SanSi将生成器的F1分数提高了7.7个百分点。

英文摘要

Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a single forward pass. A single pass is fast, intuitive System 1 thinking. We study what lies between one pass and generated reasoning: looping, in which the same layers are recursively applied several times before one typed readout. Each loop lets the model revise its hidden state before it commits to an answer, without generating a token; we call this System 1.5 thinking. We propose SanSi, which turns a pre-trained looped language model into a typed decision model. The option probabilities are read after every loop, and every loop is trained with a proper scoring rule, so that one model serves every budget from one loop to eight in a single run. On 10,027 test decisions from 59 sources, SanSi reaches 72.0% accuracy: 13.5 points above a non-looped model of the same shape trained with the same recipe, 5.3 points above a newer non-looped model of its size, and 1.8 points below one with three times the parameters. On two depth-controlled tasks, loops extend the solvable depth beyond the depths seen in training, where the larger single-pass model fails. Used as the judge for policy optimization with reinforcement learning, without gold answers, SanSi raises the generator's F1 by 7.7 points.

Comments43 pages, 15 figures, 42 tables. Project page: https://minnesotanlp.github.io/Sansi/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑