使用、提及,还是谴责?针对代码混合印地英语厌女检测中“使用-提及”区分的受控对比集诊断
Used, Mentioned, or Condemned? A Controlled Contrast-Set Diagnostic for the Use-Mention Distinction in Code-Mixed Hinglish Misogyny Detection
浏览论文内容
中文总结 AI 辅助
针对代码混合印地英语厌女检测,揭示词典模型无法区分侮辱词的使用与提及,提出对比集诊断、配对一致性指标及生成器,验证前沿LLM可达满分,而最强经典基线仍存在显著缺陷。
中文摘要 AI 辅助
基于词典的厌女检测器在构造上无法区分针对女性的侮辱性使用与在反言论中提及同一侮辱词(如“别那样叫她”)——然而正是这一区分决定了内容审核是保护还是压制讨论虐待行为的人群。我们在代码混合的印地英语(Hinglish)中研究此问题,并做出三项贡献。首先,我们在一个公开可用的脱敏语料上诊断出两个评估伪影:类别编码的匿名化占位符泄露标签(一个无学习规则得分为1.000);即使将这些占位符中和后,厌恶性与良性评论在词汇上仍处于不相交的语域,因此词袋模型在随机交叉验证下宏F1接近1.00,但在模板不相交评估下性能崩溃。其次,我们发布了Hinglish-MGY-Diag,一个确定性生成器及一个包含416个项目/163个最小对比对的对比集诊断,覆盖五个基于语言学的类别,其中侮辱词存在与性别化语域在构造上与标签去相关。第三,我们引入一个严格的配对一致性指标,仅当最小对比对的两个成员都被正确标注时才给予模型分数。在构造不相交的五折交叉验证下评估的五个从零开始的经典基线显示,最强模型在最干净的使用-提及子集上达到0.93的准确率,但一致性仅为0.82——它仍在大约五分之一的反对言论对比对上错误标注。作为作者模型上限的前沿大语言模型在所有指标上达到1.000,既作为独立的标签验证,也确认该基准是一个能力梯度而非对抗性壁垒。我们发布所有代码、数据、生成器以及一个保持距离的LLM测试框架,用于复现每个数字。
英文摘要
Lexicon-driven misogyny detectors cannot, by construction, distinguish a slur used against a woman from the same slur mentioned in counter-speech ("don't call her that") -- yet exactly this distinction governs whether moderation protects or silences the people discussing abuse. We study this problem in code-mixed Hinglish and make three contributions. First, we diagnose two evaluation artifacts on a publicly available redacted corpus: category-encoding anonymization placeholders leak the label (a no-learning rule scores 1.000), and even after they are neutralized misogynistic and benign comments occupy lexically disjoint registers, so bag-of-words reaches macro-F1 approximately 1.00 under random cross-validation but collapses under template-disjoint evaluation. Second, we release Hinglish-MGY-Diag, a deterministic generator and a 416-item / 163-minimal-pair contrast-set diagnostic across five linguistically motivated categories in which slur presence and gendered register are decorrelated from the label by construction. Third, we introduce a strict pair-consistency metric that credits a model only when both members of a minimal pair are correctly labelled. Five from-scratch classical baselines evaluated under construction-disjoint five-fold cross-validation reveal that the strongest model reaches 0.93 accuracy on the cleanest use-mention subset but only 0.82 consistency -- it still mislabels roughly one counter-speech pair in five. A frontier LLM used as an author-model ceiling attains 1.000 on all metrics, doubling as independent label validation and confirming the benchmark is a capability gradient rather than an adversarial wall. We release all code, data, the generator, and an arms-length LLM harness for reproducing every number.
发表机构
- Manipal University Jaipur(马尼帕尔大学斋浦尔分校)
- BITS Pilani, Hyderabad Campus(比拉理工学院海得拉巴校区)
机构由 AI 辅助整理,请以论文原文为准。