AI 中文总结
研究将预训练语言模型响应分布分解为策略条件表示的问题,提出潜在变量分解方法,引入新变分目标,通过多策略算法任务基准验证了该方法能恢复潜在代码并保留基础模型响应分布。
AI 中文摘要
在推理任务上训练的语言模型\(p_\theta(y \mid x)\)通过多种不同策略学习解决问题,但其策略是隐含且交织在模型响应分布中的。本文研究将给定预训练语言模型的响应分布分解为结构化、策略条件表示的问题。具体而言,学习潜在变量分解\(p_\theta(y \mid x) \leadsto (r_\phi(z \mid x), g_\phi(y \mid x,z))\),其中路由器\(r\)将每个输入映射到潜在策略\(z\)的分布,生成器\(g\)根据该策略生成响应。面临的关键挑战是生成器从基础模型初始化,已经能表示\(p_\theta(y \mid x)\)而不使用\(z\),标准变分推理无法激励模型通过\(z\)路由信息,可能导致后验坍塌。为解决此问题,提出一个变分目标,测量相对于基础模型响应损失的分数信息增益,并将重建压力集中在基础模型惊讶度高的 tokens 上,鼓励\(z\)编码与策略相关的响应变化。引入了多策略算法任务基准,表明该目标能恢复与不同参考策略对齐的潜在代码,同时保留基础模型的响应分布。
英文摘要
A language model $p_θ(y \mid x)$ trained on reasoning tasks learns to solve problems via multiple distinct strategies, yet these strategies are implicit and entangled within the model's response distribution. We study the problem of decomposing the response distribution of a given pretrained language model into a structured, strategy-conditioned representation. Specifically, we learn a latent-variable factorization $p_θ(y \mid x) \leadsto (r_ϕ(z \mid x), g_ϕ(y \mid x,z))$, where a router $r$ maps each input to a distribution over latent strategies $z$ and a generator $g$ produces the response conditioned on that strategy. A key challenge is that the generator, initialized from the base model, already represents $p_θ(y \mid x)$ without using $z$. Standard variational inference therefore gives the model no incentive to route information through $z$ and can yield a severe form of posterior collapse. To address this, we propose a variational objective that measures fractional information gain relative to the base model's response loss and concentrates reconstruction pressure on tokens with high base model surprisal, encouraging $z$ to encode strategy-relevant response variation. We introduce a benchmark of multi-strategy algorithmic tasks and show that this objective recovers latent codes aligned with distinct reference strategies while preserving the base model's response distribution.