ClaimMirage:域名中的自我声明如何改变大语言模型的威胁判断
ClaimMirage: When Self-Claims in Domain Names Change LLM Threat Judgments
- Tokyo Metropolitan University(东京都立大学)
- NTT Security Holdings Corporation & NTT, Inc.(NTT安全控股公司及NTT公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究域名中的自我安全声明(如“非钓鱼”)如何操纵大语言模型的威胁判断,通过大规模实验发现此类声明可显著改变警报率,并强调需独立验证域名安全性。
AI中文摘要:
诸如“非钓鱼”或“官方”之类的简短声明,可以在没有明确提示注入命令的情况下,改变大语言模型(LLM)对域名的判断。我们将这种操纵行为研究为ClaimMirage:被检查的名称声称自身安全或获得批准。我们分析了64个品牌和5个大语言模型中的622,080个判断,在构造的品牌类名称中,将十种声明与长度和连字符匹配的对照组进行比较。自我声明可以大幅减少或增加警报,具体取决于大语言模型和输入设置。在一种设置中,即使有基本防护措施——提示提供可能被冒充的品牌及其官方域名以供比较——可注册名称中的风险否认术语仍将警报减少了45.3个百分点。在没有这些参考的情况下,同一大语言模型中该位置的认可术语将警报增加了65.6个百分点。参考和组件注释消除了部分警报减少,但保留了其他部分或使其更大。这些发现促使在将被检查的域名视为安全或授权之前,测试对自我声明的抵抗力并寻求独立证据。
英文摘要:
Short claims such as not-phishing or official can change how a large language model (LLM) judges a domain name, without explicit prompt-injection commands. We study this manipulation as ClaimMirage: a name under inspection claims its own safety or approval. We analyze 622,080 judgments across 64 brands and five LLMs, comparing ten claims with length- and hyphen-matched controls in constructed brand-like names. Self-claims can substantially reduce or increase alerts, depending on the LLM and input setting. In one setting, risk-denial terms inside the registrable name reduce alerts by 45.3 percentage points even with a basic safeguard: the prompt supplies the potentially impersonated brand and its official domain for comparison. Without these references, endorsement terms at that position increase alerts by 65.6 points in the same LLM. References and component annotation remove some alert reductions but leave others or make them larger. These findings motivate testing resistance to self-claims and seeking independent evidence before treating a domain name under inspection as safe or authorized.