模型催眠:通过叠加阈下效应对AI进行强控制
Model Hypnosis: Strong control of AI via additive subliminal effects
- University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究发现AI模型普遍存在模型催眠现象,可通过组合提示中微弱无关线索强控制模型行为,该现象跨模型家族且具迁移性,对AI安全与可解释性构成挑战。
AI中文摘要:
我们证明AI模型普遍易受一种称为模型催眠的现象影响,即提示中单独微弱且看似无关的线索可被系统组合以强控制模型行为。模型催眠存在于各类模型家族及规模中,包括前沿推理模型,且催眠提示可在模型间迁移。由于模型受释义、拼写错误等不显眼的文本选择控制,模型催眠给AI安全带来新挑战与机遇,也是AI可解释性的主要障碍。
英文摘要:
We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irrelevant cues in the prompt can be systematically combined to strongly control model behavior. Model hypnosis occurs across model families and scales, including in frontier reasoning models, and hypnotic prompts can transfer between models. Because the model is controlled by inconspicuous textual choices, such as paraphrases and typos, model hypnosis presents new challenges and avenues for AI safety, and is a major hurdle for AI interpretability.