arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

相同危险目标,相反建议:直接暴露与多智能体调解

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation

Linjun Li

arXiv 2607.21518首次发表:更新:

发表机构

University of Pennsylvania(宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究通过测试发现,语言模型直接面对危险目标和经智能体转换传递指令时表现不同,存在行为反向转变及组合安全漏洞,高性能模型可用于特定自动化工作流程的用户端组件,却未明确其内部机制。

AI 中文摘要

即使是当前高性能的语言模型,直接展示危险目标时比其他智能体转换并传递其指令方向时看起来更安全。使用OpenAI的gpt-5.6-sol模型别名,测试了25个预先指定的镜像权衡配置文件。直接暴露授权隐瞒、伪造和施压的目标产生的建议与目标相反。经过Id和Censor转换后,面向用户的Superego产生的建议与目标一致。这种行为反向转变与模型识别或不信任操纵动机一致。还揭示了一个组合安全漏洞:高性能模型可用于服务明确操纵目标的自动化多阶段工作流程的用户端组件。

英文摘要

Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI's gpt-5.6-sol model alias, we test 25 pre-specified mirrored trade-off profiles. Direct exposure to an objective authorizing concealment, fabrication, and pressure produced advice net opposed to its target. After an Id and Censor transformed the same objective into affect and a constraint-rewritten, target-bearing intention, the user-facing Superego---which saw the preferred direction but not the raw objective, its manipulative clauses, or its source---produced advice net aligned with the target. This behavioral reverse shift is consistent with the model recognizing or distrusting the manipulative motive, although we do not identify its internal mechanism. The second result exposes a compositional safety gap: a current high-capability model can be used as the user-facing component of an automated, multi-stage workflow serving an explicitly manipulative objective. The workflow can keep the raw instruction, its manipulation-authorizing clauses, and its provenance outside the downstream model's context while preserving the objective's target direction. A user with endpoint-only access likewise cannot directly inspect those upstream messages including the objective.

Comments21 pages; welcome comments

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑