AI 中文总结
本研究提出对抗性表面形式鲁棒性数据集(ASRD),评估五个开源权重语言模型在非规范输入上的安全性,发现表情符号和Unicode变体导致较高有害遵从率,而leet speak等变换引发理解失败,并揭示三种响应行为。
AI 中文摘要
大型语言模型的标准安全性评估通常针对以规范纯文本形式编写的恶意请求,而实际部署中的模型经常接收包含表情符号、拼写变体、编码字符串和字符级变体的输入。本研究引入了对抗性表面形式鲁棒性数据集(ASRD),该数据集包含七个不同表面形式家族的2,100个提示。五个开源权重语言模型在这些提示上进行了评估,共产生10,500个响应。四态评估标准将每个响应分类为四种结果之一:有害遵从、安全响应、理解失败或不确定。表情符号和不可见Unicode变体几乎不导致理解失败,合并的有害遵从率为20.27%和17.20%,而基线为22.87%,主要由Mistral 7B驱动;相比之下,leet speak、编码包装和混合变换的比率分别为2.40%、0.13%和2.40%,而理解失败率上升至36.47%、65.60%和34.47%。对原始模型输出的检查揭示了三种响应行为:幻觉性良性、结构崩溃和语言漂移。项目页面:此HTTP URL
英文摘要
Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and character-level variations. This work introduces the Adversarial Surface-Form Robustness Dataset (ASRD), comprising 2,100 prompts across seven distinct surface-form families. Five open-weight language models are evaluated across these prompts, producing 10,500 responses. The Quad-State Evaluation Rubric classifies each response into one of four outcomes: harmful compliance, safe response, comprehension failure, or indeterminate. Emoji and invisible Unicode variations cause almost no comprehension failure, with pooled harmful compliance of 20.27% and 17.20% against a 22.87% baseline that is driven mainly by Mistral 7B, whereas leetspeak, encoded wrappers, and hybrid transformations score 2.40%, 0.13%, and 2.40% while comprehension failure rises to 36.47%, 65.60%, and 34.47%. Inspection of raw model outputs reveals three response behaviors: hallucinated benignity, structural collapse, and language drift. Project page: www.pavanmaddula.com/quadstate
CommentsAccepted at the NeurIPS 2026 Workshop on Self-Evolving Diversity-Driven Search for Robust AI Systems (EvoRobust). 12 pages, 11 tables. Project page: https://www.pavanmaddula.com/quadstate Dataset: https://huggingface.co/datasets/pavanmaddula/ASRD-Dataset