发表机构
University of Technology Sydney(悉尼科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对多语言不满情绪标注,指出词级词典存在问题,提出用上下文阅读模型取代术语匹配,在非循环基准上评估,结果显示阅读上下文能更准确测量不满情绪,在未选文本上测试更诚实,还发布了代码和基准。
AI 中文摘要
不满情绪是分析师评估暴力威胁时寻找的警示信号之一。目前常通过在线文本大规模测量,多使用如不满情绪词典这样的词级词典,通过匹配加权术语计分。但这种匹配无法解决术语是被断言、引用、否定还是谴责的问题,且词典评估常基于其自身检索的示例池,导致高分部分反映与词典自身选择规则的一致性。研究一个五种语言、2000项的评估池发现,词典自身几乎完美地将其分成两半,其表观宏AUROC为0.686降至固定的0.500下限。研究保留词典的22结构本体,用上下文阅读模型取代术语匹配,并在非循环基准上评估,该基准将五种语言的无条件随机、词典正向和词典负向层次分开。阅读完整帖子而非仅目标句子在词典沉默时帮助最大,提高了词典负向文本的平均精度,从0.14提高到0.20,在引用、隐含和跨句子不满情绪上收益最大。结果表明,通过阅读周围上下文能更忠实地测量不满情绪,在词典未选择的文本上测试时更诚实。
英文摘要
Grievance is one of the warning signs analysts look for when assessing threats of violence. It is increasingly measured at scale from online text, most often with word-level lexicons like the Grievance Dictionary that score by matching weighted terms. Such matching is a fast and transparent proxy, but it cannot resolve whether a term is asserted, quoted, negated, or condemned. These lexicons are also often evaluated on pools enriched with the very examples they retrieve, so a high score partly reflects agreement with the lexicon's own selection rule. Examining a five-language, 2{,}000-item evaluation pool, we find its halves separated almost perfectly by the lexicon itself: every item labeled ``random'' is in fact lexicon-negative, so the lexicon's apparent macro-AUROC of 0.686 collapses to a 0.500 floor fixed by construction. We keep the dictionary's 22-construct ontology but replace term matching with context-reading models, evaluated on a non-circular benchmark that separates unconditional-random, lexicon-positive, and lexicon-negative strata across five languages. Reading the full post rather than the target sentence alone helps most where the lexicon is silent, raising average precision on lexicon-negative text from 0.14 to 0.20, with the largest gains on quoted, implicit, and cross-sentence grievance. Together, these results show that grievance is measured more faithfully by reading the surrounding context, and more honestly when tested on text the lexicon did not select. We release our code and benchmark at https://github.com/behavioral-ds/multilingual_grievance.
Comments12 pages, 1 figure, 9 tables