Deep Noir:通过Transformer模型中的架构计时学实现自主转向发现
Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models
浏览论文内容
中文总结 AI 辅助
Deep Noir通过Logit Lens收敛和因果头级归因自主发现最优转向参数,在多个规模和任务上显著提升性能,并揭示转向带来的提示注入攻击风险。
中文摘要 AI 辅助
激活转向在推理时修改LLM行为,但确定转向位置和强度仍然是手动的。我们引入了Deep Noir,一个利用Logit Lens收敛和因果头级归因来自主发现最优转向参数的框架。在三个规模(1B x 3、2-3B x 2和7-9B x 4)上,我们的引擎在1B规模的垃圾邮件检测上实现了16.7个百分点的改进(标准差4.7;39次运行),在7-9B规模的四种架构上,改进幅度增加到21至42个百分点。在SST-2情感分析上,它在零代码更改的情况下实现了13.1个百分点的改进。机制层面的基础使得能够自动发现跨任务和架构泛化的干预点。在情感分析上,没有头掩蔽的RepE未能超过基线,而Deep Noir改进了所有模型(p小于0.01)。我们进一步表明,转向创建了一个可预测的提示注入攻击面,其脆弱性随转向幅度单调增加。这一发现与部署转向分类器的智能体系统相关。
英文摘要
Activation steering modifies LLM behavior at inference time, but identifying where and how strongly to steer remains manual. We introduce Deep Noir, a framework that uses Logit Lens convergence and causal head-level attribution to autonomously discover optimal steering parameters. Across three scales (1B x 3, 2-3B x 2, and 7-9B x 4), our engine achieves 16.7 percentage-point improvement on spam at 1B (standard deviation 4.7; 39 runs), with gains increasing to 21 to 42 percentage points at 7-9B across four architectures. On SST-2 sentiment, it achieves a 13.1 percentage-point improvement with zero code changes. Mechanistic grounding enables automated discovery of intervention points that generalize across tasks and architectures. On sentiment, RepE without head masking fails to improve over baseline, while Deep Noir improves all models (p less than 0.01). We further show that steering creates a predictable prompt-injection attack surface whose vulnerability increases monotonically with steering magnitude. This finding is relevant to agent systems deploying steered classifiers.
发表机构
- Naval Surface Warfare Center Panama City Division(海军水面作战中心巴拿马城分部)
机构由 AI 辅助整理,请以论文原文为准。