arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29001cs.CL

礼貌但不一致:评估LLM礼貌判断与人类语用规范的契合度

Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms

发表机构蒂宾根大学 · 同济大学
查看机构详情
  • University of Tübingen(蒂宾根大学)
  • Tongji University(同济大学)

机构由 AI 辅助整理,请以论文原文为准。

Rong Wang, Kun Sun, Yadong Guo

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估七个LLM的礼貌判断与人类语用规范的契合度,发现模型间一致性高于模型与人类的一致性,且模型存在系统性中性压缩偏差,强调需超越总体指标考察分歧方向性模式。

中文摘要 AI 辅助

尽管大型语言模型(LLM)在标准基准测试中表现强劲,但其评估社会语用能力是否与人类判断一致仍不明确。我们使用两个具有互补标注格式的英语数据集评估LLM的礼貌判断:连续人类评分和三分类标签。在评估的七个模型中,我们发现模型间一致性高于模型与人类的一致性。策略层面分析表明,模型与人类的契合度与显性语言线索相关,而某些建立融洽关系的策略在不一致案例中出现频率更高。在分类任务中,模型预测表现出系统性中性压缩,其特征为过度生成“中性”标签和不足预测“不礼貌”标签。当以专家共识作为诊断子集的参考时,该模式依然存在。我们的研究结果强调,语用评估需要超越总体一致性指标,通过考察不同人类参考下模型与人类分歧的方向性模式来进行。

英文摘要

Despite strong performance on standard benchmarks, it remains unclear whether large language models (LLMs) evaluate social pragmatics in ways that align with human judgments. We evaluate LLM politeness judgments using two English-language datasets with complementary annotation formats: continuous human ratings and three-way categorical labels. Across the seven evaluated models, we find that inter-model agreement is stronger than model--human agreement. Strategy-level analyses suggest that model--human alignment is associated with explicit linguistic cues, while some rapport-building strategies occur more frequently in misaligned cases. In the categorical task, model predictions exhibit systematic neutral compression, characterized by the overproduction of Neutral labels and the underprediction of Impolite labels. This pattern persists when expert consensus is used as the reference on a diagnostic subset. Our findings highlight the need for pragmatic evaluations that go beyond aggregate agreement metrics by examining directional patterns of model--human disagreement across different human references.

补充信息

↑