arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33039cs.AI

从内部保障智能体安全:从LLM内部状态检测有害轨迹

Agent Safety From Within: Detecting Harmful Trajectories from LLM Internal States

Difan Jiao, Ashton Anderson

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有防护模型难以检测智能体轨迹风险的问题,提出TACIT方法,通过线性探针直接读取LLM内部状态,在六个基准上将宏F1从62.3提升至86.2,且训练成本极低。

中文摘要 AI 辅助

语言模型智能体现在能够通过工具和接口执行复杂的动作序列,这增加了它们可能造成的损害范围。然而,防护模型主要针对内容审核而构建,因此不适合检测这种智能体风险。为解决此问题,我们首先进行表征分析,然后利用所得见解构建解决方案。在分析中,我们聚焦于两类轨迹级智能体危害:有害内容(直接表达)和不安全的工具使用(取决于动作是否与产生该动作的交互一致)。我们研究了开源防护模型如何表征这两类危害,发现它们可在模型内部被线性读取,尽管防护模型在仅调用工具模式不同的成对样本上的预测不优于随机水平。两类危害还遵循几乎正交的内部方向,且两者都不能可靠地作为对方的代理。这些结果促使我们直接从内部状态读取轨迹安全性。我们引入TACIT,一种对冻结主干网络内部状态的读取器,不解码任何令牌。在六个轨迹安全基准上训练后,线性探针将平均宏F1分数从最强开源防护模型的62.3提升至80.7,精炼读取器达到86.2。当每个基准完全从训练中排除时,精炼读取器仍领先于最强防护模型(65.7对61.1)。在相同主干网络、训练数据和测试划分下,冻结读取器与完整安全微调性能相当,且在其之上应用时进一步改进微调模型。该探针训练的参数量约为完整微调的百万分之一,时间约为六分之一,且TACIT在我们评估的防护模型中具有最低延迟。

英文摘要

Language model agents can now perform sophisticated sequences of actions via tools and harnesses, which has increased the scope of the damage they can cause. Guard models, however, are mainly built for content moderation and thus are not well-suited to detecting this agentic risk. To address this, we proceed by first conducting a representational analysis, then use the resulting insights to build a solution. In our analysis, we focus on two types of trajectory-level agentic harms: harmful content, which is expressed directly, and unsafe tool use, which depends on whether an action is consistent with the interaction that produced it. We investigate how open-source guard models represent these two types of harm and find that they are linearly readable inside the model, even though guard models predict no better than chance on pairs that differ only in the called tool's schema. The two harm types also follow nearly orthogonal internal directions, and neither reliably serves as a proxy for the other. These results motivate reading trajectory safety directly from internal states. We introduce TACIT, a readout of a frozen backbone's internal states that decodes no tokens. Trained on six trajectory-safety benchmarks, a linear probe raises mean macro-F1 from 62.3 for the strongest open guard to 80.7, and refined readouts reach 86.2. With each benchmark held out of training entirely, the refined readouts still lead the strongest guard (65.7 vs. 61.1). With the same backbone, training data and test split, the frozen readout is on par with full safety fine-tuning, and it improves the fine-tuned model further when applied on top. The probe trains about one millionth as many parameters as full fine-tuning in about a sixth of the time, and TACIT has the lowest latency of the guards we evaluate.

补充信息

↑