arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SinoGlyphBench:面向语言模型审核的中文字形级混淆诊断基准

SinoGlyphBench: A Diagnostic Benchmark for Chinese Glyph-Level Obfuscation in Language-Model Moderation

Yifan Wang, Zimu Wang, Suliu Qin, Changyu Zeng, Tong Chen, Siqi Chen, Yijie Lin, Lingyu Jiang, Jionglong Su, Yushan Pan, Haiyang Zhang, Wei Wang, Qiaoyu Tan

arXiv 2609.05843首次发表:更新:

发表机构

East China Normal University; New York University Shanghai; University of Liverpool; Singapore University of Technology and Design; Eastern Institute of Technology (Ningbo)(华东师范大学; 上海纽约大学; 利物浦大学; 新加坡科技设计大学; 宁波东方理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SinoGlyphBench诊断基准,通过锚点扰动区分审核证据破坏与表面变化,评估12个模型发现字形混淆显著增加有害内容漏报和误报,模型对非规范字形中文仍脆弱。

AI 中文摘要

字形级混淆可使有害中文内容对人类可读,同时削弱自动化审核。我们提出SinoGlyphBench,一个诊断基准,用于识别标签关键语义锚点,并创建文本和图像模态下匹配的原始与字形混淆输入。通过扰动锚点、背景上下文或两者,该设计区分了审核相关证据的破坏与一般表面变化。在对12个LLM和MLLM的176,916次配对评估中,混淆使有害假阴性和假阳性率分别增加6.1和4.7个百分点,并将四路准确率降低5.0个百分点。模型保留了在匹配原始输入上正确决策的75.7%。全范围扰动导致最大退化,仅锚点扰动比仅背景扰动更具破坏性,跨文字替换在文本模态中尤为困难。对结构化输出的分析识别出可见形式阅读、预期信息恢复和最终安全判断中的可观察不匹配。因此,所评估的模型对使用非规范字形书写的中文内容仍然脆弱。资源可在https URL获取。

英文摘要

Glyph-level obfuscation can leave harmful Chinese content readable to humans while degrading automated moderation. We introduce SinoGlyphBench, a diagnostic benchmark that identifies label-critical semantic anchors and creates matched original and glyph-obfuscated inputs in text and image modalities. By perturbing anchors, background context, or both, this design distinguishes corruption of moderation-relevant evidence from general surface variation. Across 176,916 paired evaluations of 12 LLMs and MLLMs, obfuscation increases harmful false-negative and false-positive rates by 6.1 and 4.7 percentage points, respectively, and reduces four-way accuracy by 5.0 points. Models retain 75.7% of the decisions that were correct on the matched original inputs. Full-scope perturbations cause the largest degradation, anchor-only perturbations are more damaging than background-only perturbations, and cross-script substitution is particularly difficult in the text modality. Analysis of structured outputs identifies observable mismatches in visible-form reading, intended-message recovery, and final safety judgment. The evaluated models, therefore, remain brittle to Chinese content written with non-canonical glyphs. Resources are available at https://github.com/fengshun124/SinoGlyphBench.

Comments24 pages, 6 figures, 16 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑