仅前向传播一次:在冻结语言模型的单次前向传播中同时完成回答与弃权(不执行)
You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model
浏览论文内容
中文总结 AI 辅助
该研究提出YOPO系统,通过训练小型网络重构残差流,在冻结Qwen2.5模型单次前向传播中同时实现回答、引导与弃权,性能优于两次前向传播基准,且在跨域等场景表现出色。
中文摘要 AI 辅助
冻结语言模型在推理任务中存在两个耦合的缺陷:其一,未充分利用其自身残差流已编码的证据;其二,无法检测输入是否不足以完成回答,从而导致生成虚假内容。本文整合了两条在同一残差流上解决上述问题的研究路线:条件引导探针在中间层写入残差流,从冻结的主干模型中恢复推理准确率;零样本充分性方向读取残差流,并在信息不足时弃权(不执行)。若将两者部署在单次前向传播中,会产生干扰:引导写入会改变方向读取的状态,在小型模型上造成跨域迁移的AUROC损失最高达8个点;若采用单独的干净前向传播,推理成本会翻倍。我们将该方向固定,训练一个小型网络从已引导的残差中重构引导前的残差——使用(引导后、干净)残差对的均方误差进行训练,无需充分性标签——并在重构结果上读取该方向。由此得到的系统YOPO(You Only Pass Once)可在冻结的Qwen2.5主干模型(1.5B/3B/7B)的单次前向传播中完成回答、引导与弃权。端到端的三向准确率较冻结基线提升一倍以上(1.5B alphaNLI上从0.375升至0.798),且在所有规模下,单次前向传播的性能均优于两次前向传播的基准(1.5B/3B/7B分别为0.798/0.830/0.893 vs 0.753/0.790/0.863),并在六个模型家族的十个主干模型上均表现出色。我们绘制了容量迁移前沿,量化了“不应在源侧训练弃权”的原则;通过源侧审计发现自身alphaNLI构造存在表面人工痕迹,因此架构主张基于原生标签复现(SQuAD2、RepLiQA、MuSiQue);在标准四域测试集上,我们贡献了首个答案或弃权基准,其中我们的门控模型在每个域内数据集上均表现最佳,且无监督方向是唯一能在跨域迁移中保留的门控家族。
英文摘要
A frozen language model on reasoning tasks has two coupled weaknesses: it under-uses evidence its own residual stream already encodes, and it fails to detect when the input is insufficient to answer, so it confabulates. This paper consolidates two research lines that address these on the same residual stream: a conditional steering probe writes the stream at mid-stack layers and recovers reasoning accuracy from a frozen backbone, and a zero-shot sufficiency direction reads the stream and abstains when information is insufficient. Deployed in one forward pass they interfere: the steering write shifts the state the direction reads, costing up to 8 AUROC points of cross-domain transfer on small models; a separate clean pass doubles inference cost. We keep the direction fixed and train a small network to reconstruct the pre-steering residual from the steered one -- mean-squared error on (steered, clean) pairs, no sufficiency labels -- and read the direction on the reconstruction. The resulting system, YOPO (You Only Pass Once), answers, steers, and abstains in one forward pass of a frozen Qwen2.5 backbone (1.5B/3B/7B). End to end, three-way accuracy more than doubles the frozen baseline (0.375->0.798 on 1.5B alphaNLI) and one pass beats the two-pass reference at every scale (0.798/0.830/0.893 vs 0.753/0.790/0.863) and on ten backbones across six model families. We chart the capacity-transfer frontier quantifying the principle that abstention should not be trained in; a source-side audit catches our own alphaNLI construction leaking a surface artifact, so architectural claims are anchored on native-label replications (SQuAD2, RepLiQA, MuSiQue); and on the standard four-domain suite we contribute, to our knowledge, the first answer-or-abstain benchmark, where our gate tops every in-domain dataset and the label-free direction is the only gate family to survive domain transfer.