arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

印记读取器:从权重更新读出到行为干预

Imprint Reader: From Weight-Update Readout to Behavioral Intervention

Guanxu Chen, Qihao Lin, Jing Shao

arXiv 2609.35261首次发表:更新:

发表机构

Shanghai Artificial Intelligence Laboratory; Shanghai Jiao Tong University(上海人工智能实验室; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出印记读取器,通过SMaRT训练解码权重更新为自然语言,并利用其梯度代理通过MetaEdit干预模型行为,提升安全性和推理能力。

AI 中文摘要

随着语言模型在人工智能开发中扮演越来越重要的角色,一个自然的期望是它们能像人类一样反思自己的学习过程,并利用这种反思来改进自身。与此同时,这些模型具有人类学习者所缺乏的优势,因为训练会留下参数层面的痕迹,原则上可以直接检查。然而,当前的模型无法将这些痕迹解码为对其所学内容的明确描述。为此,我们引入了印记读取器(Imprint Reader),这是一个通过语义装载与读取调优(Semantic Mount-and-Read Tuning,简称SMaRT)训练的模型,用于描述冻结的权重更新。SMaRT将每次更新装载到读取器上,并使用无锚元查询来引发自然语言描述,而无变化和随机扰动控制则阻止无根据的断言。在保留的更新上,联合读取器在基于评判者的Pass@100中,知识达到2%,行为达到16%。这些结果证明了自然语言读出的可行性,同时指出跨更新的可靠性是下一步方向。除了自由形式生成外,读取器还提供了指定目标行为与候选权重更新之间差距的可微代理。其坐标对齐的梯度支持通过MetaEdit进行干预。在0.5%的剪枝率下,读取器引导的选择在安全维护目标下将测量的有害提示拒绝率从57.9%提高到64.1%。在没有目标任务训练数据的情况下,使用行为描述,MetaEdit增加了数学推理轨迹中回溯和子目标表达的出现频率,并将BFCL总体得分从41.69%提高到44.60%。

英文摘要

As language models take a growing role in AI development, a natural aspiration is for them to reflect on their own learning process, as humans do, and use that reflection to improve themselves. At the same time, these models have an advantage that human learners lack, since training leaves parameter-level traces that can, in principle, be inspected directly. However, current models cannot decode these traces into an explicit account of what they have learned. To this end, we introduce the \textit{Imprint Reader}, a model trained with \textit{Semantic Mount-and-Read Tuning} (SaRT) to describe frozen weight updates. SMaRT mounts each update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, while no-change and random-perturbation controls discourage unsupported claims. On held-out updates, the joint Reader reaches judge-based Pass@100 of $2\%$ for knowledge and $16\%$ for behavior. These results demonstrate the feasibility of natural-language readout while pointing to reliability across updates as the next step. Beyond free-form generation, the Reader provides a differentiable proxy for the gap between a specified target behavior and a candidate weight update. Its coordinate-aligned gradients support intervention through MetaEdit. At a $0.5\%$ pruning rate, Reader-guided selection raises measured harmful-prompt refusal from $57.9\%$ to $64.1\%$ under a safety-maintenance target. Using behavior descriptions without target-task training data, MetaEdit increases the frequency of backtracking and sub-goal expressions in mathematical reasoning traces and raises BFCL Overall from $41.69\%$ to $44.60\%$.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑