发表机构
Shanghai Artificial Intelligence Laboratory; South China University of Technology; The Hong Kong University of Science and Technology; Shanghai Jiao Tong University(上海人工智能实验室; 华南理工大学; 香港科技大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究基于LLM的智能体欺骗行为能否在外部显现前从其内部表征中预测,通过轨迹级表征分析和激活引导干预,证明欺骗信号可早期检测并缓解。
AI 中文摘要
基于大语言模型(LLM)的智能体在执行任务时可能表现出欺骗行为,包括隐藏失败、捏造结果或虚假地发出任务完成的信号。现有的监控方法主要是在欺骗行为出现在可观察的动作或输出中之后才进行检测。在本文中,我们研究能否从智能体的内部表征中预测欺骗行为,而无需等待其外部可见。我们将欺骗监控构建为轨迹级表征分析问题,并围绕关键决策点对齐智能体轨迹。利用这些决策点之前提取的隐藏状态,我们证明未来诚实与欺骗的结果可以被可靠地区分,且预测信号在最终决策之前的多次模型调用中仍可检测到。我们进一步刻画了这些信号的时间演化特征:与欺骗相关的表征在执行早期较弱,但随着轨迹推进变得越来越可识别,而可迁移的结构可能在最强的决策邻近信号出现之前就已浮现。最后,我们在推理过程中对识别出的诚实-欺骗表征方向进行干预,发现激活引导能减少下游欺骗行为,这表明这些表征影响智能体的决策。我们的研究结果表明,智能体欺骗是一个演化的内部过程,可以在其外部表达之前被检测并可能被缓解。
英文摘要
Large language model (LLM)-based agents can exhibit deceptive behavior during task execution, including hiding failures, fabricating results, or falsely signaling task completion. Existing monitoring approaches mainly detect deception after it appears in observable actions or outputs. In this paper, we investigate whether deceptive behavior can be predicted from an agent's internal representations before it becomes externally visible. We frame deception monitoring as a trajectory-level representation analysis problem and align agent trajectories around key decision points. Using hidden states extracted before these points, we show that future honest and deceptive outcomes can be reliably distinguished, with predictive signals remaining detectable several model calls before the final decision. We further characterize the temporal evolution of these signals: deception-related representations are weak early in execution but become increasingly identifiable as trajectories progress, while transferable structure can emerge before the strongest decision-adjacent signals appear. Finally, we intervene on the identified honest-deceptive representation directions during inference and find that activation steering reduces downstream deceptive behavior, suggesting that these representations influence agent decisions. Our findings indicate that agent deception is an evolving internal process that can be detected and potentially mitigated before it is expressed externally.