AI 中文总结
研究针对语言模型中不可信文本影响问题,开发基于来源可信度的信任模型,通过确定性管道和非模型监视器防御。经实验,钝化和级联提升真实泄漏防御率,自适应红队测试下边界成立,防御率稳定,还提高了事实归属率。
AI 中文摘要
在语言模型中,指令和数据共享一个令牌流,因此模型生成过程中没有任何内部机制能够阻止不可信文本对其产生影响。我们开发了一种信任模型,将决策权限置于模型之外的代码中:来源的可信度而非内容决定运行的操作及其行为。低可信度来源可提供信息但不能覆盖高可信度来源。未修改的模型在一个确定性管道中运行,该管道根据来源完整性对输入进行排序,一个固定的非模型监视器仅从可信输入中选择操作和任何外部动作。我们可以测量但无法证明该管道对注入的抗性;我们对其进行提示调整并报告速率。在未修改的Gemma~4 26B模型的一次性留出集上,钝化和一个包装器(级联)将真实泄漏防御率从27%提高到94%,同时清洁质量成本约为4%(相对质量$Q_{\mathrm{rel}}{=}0.96$)。在自适应红队测试下,已验证的边界无条件成立,可衡量的防御率保持在87%。级联还会归属低可信度来源的事实而非丢弃它,将归属率从0%提高到92%,并在冲突时遵循高可信度来源。
英文摘要
In a language model, instructions and data share one token stream, so nothing inside the model's generation can keep untrusted text from steering it. We develop a trust model that places the authority to act outside the model, in code: a source's standing, not its content, decides which operation runs and whether it acts. A lower-trust source may inform an answer but not override a higher one. An unmodified model runs inside a deterministic pipeline that ranks inputs by source integrity, and a fixed non-model monitor provably chooses the operation and any outside action from trusted inputs alone. We can measure but not prove the pipeline's resistance to injection; we prompt-tune it and report the rate. On a one-shot held-out set with an unmodified Gemma~4 26B model, passivation and a wrapper (the cascade) raise the genuine-leak defended rate from $27\%$ to $94\%$ at roughly a $4\%$ clean-quality cost ($Q_{\mathrm{rel}}{=}0.96$). Under adaptive red-teaming the proved boundary holds unconditionally, and the measured defense stays at $87\%$. The cascade also attributes a lower-trust source's fact rather than dropping it, raising attribution from $0\%$ to $92\%$, and follows the higher-trust source on a conflict.