发表机构
University of Technology Sydney(悉尼科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有道德基础检测系统跨域泛化差、成本高的问题,提出CHARM框架,其基于轻量微调LLM,在多数据集上优于基线,可用于研究COVID-19相关道德化错误信息传播。
AI 中文摘要
道德语言在塑造在线认可与信息传播中发挥核心作用,但现有道德基础检测系统常存在跨域泛化能力差、依据不足且依赖成本高昂的基于提示的大语言模型(LLM)的问题。我们提出CHARM框架,即感知MAC与仇恨言论、依据对齐的道德基础检测框架,该框架基于轻量微调LLM构建,整合互补的道德依据、依据对齐以及感知极性的仇恨言论信号,以支持更鲁棒、更可信的道德预测。与之前将计算与心理理论解耦的基于词典、微调或提示的检测器不同,CHARM的每个组件——MAC交叉注意力、依据对齐和仇恨言论调制——均对应一个独特的心理构念。使用MFTC、MFRC和News训练池的30%子样本,结合MFTCXplain中更丰富的监督信号,CHARM在域内将AUC提升了15.3%,在所有跨域数据集上的AUC和F1均优于监督基线,且为基于提示的LLM检测器提供了可扩展、低成本的替代方案。我们进一步将CHARM应用于Twitter上的大规模COVID-19论述,发现道德价值对齐与在线认可行为密切相关。通过实现道德框架的规模化可测量,CHARM为研究道德化错误信息的传播提供了实用工具。
英文摘要
Moral language plays a central role in shaping online endorsement and the diffusion of information, yet existing moral foundation detection systems often suffer from poor cross-domain generalization, weak rationale grounding, and reliance on costly prompting-based large language models (LLMs). We introduce CHARM, a MAC- and Hate-speech-Aware Rationalealigned Moral foundation detection framework built on a lightweight fine-tuned LLM, which integrates complementary moral grounding, rationale alignment, and polarity-aware hate speech signals to support more robust and faithful moral prediction. Unlike prior dictionary-, fine-tune-, or prompt-based detectors, which decouple computation from psychological theory, CHARM is built so that each component -- MAC cross-attention, rationale alignment, and hate-speech modulation -- operationalizes a distinct psychological construct. Using a 30\% subsample of the MFTC, MFRC, and News training pools together with the richer supervision in MFTCXplain, CHARM improves AUC by up to 15.3\% in-domain, surpasses the supervised baselines on every out-of-domain dataset in both AUC and F1, and offers a scalable, low-cost alternative to prompting-based LLM detectors. We further apply CHARM to large-scale COVID-19 discourse on Twitter and show that moral value alignment is strongly associated with online endorsement behavior. By making moral framing measurable at scale, CHARM offers a practical tool for studying the spread of morally charged misinformation. Code and additional materials: https://github.com/HuixiangF/CHARM/.
CommentsAccepted to the EMNLP 2026 Main Conference