发表机构
Stony Brook University; University of Virginia(石溪大学; 弗吉尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出首个基于语用学理论的LLM拒绝分类法,通过分析16个LLM在14类有害请求下的拒绝回复,发现其拒绝的特点及存在的问题,呼吁开展兼顾情境适应性与社会问责的对齐评估。
AI 中文摘要
在语用学中,拒绝常被视为威胁面子的行为,因为它们可能挑战请求者在社会层面宣称的自我形象。大语言模型(LLM)正接受越来越多的训练以拒绝不安全和不恰当的请求,而当模型未能妥善处理这种互动成本时,这些拒绝可能会伤害用户。现有研究主要将LLM的不配合视为安全对齐的结果,但未提供一种方法来评估LLM在不同有害情境下的拒绝是否恰当。为研究这一问题,我们提出(据我们所知)首个基于语用学理论的LLM拒绝分类法。将该分类法应用于16个现代LLM在14个有害类别下的回复,我们发现,尽管模型的拒绝方式存在差异,但总体而言它们的拒绝是明确的且带有强烈道德评价性,互动修复主要通过提供或给出更安全的替代方案实现,而非人际面子维护。这种模式在敏感有害情境中尤为重要,在此情境下,过度使用负面框架可能会让用户感到羞愧或被激怒,从而破坏安全不配合的目的。因此,我们呼吁开展对齐评估,不仅要考虑模型是否拒绝有害请求,还要考虑其拒绝方式是否具有情境适应性,并对说“不”的互动后果承担社会层面的责任。
英文摘要
Refusals are often treated as face-threatening acts in pragmatics because they can challenge the requester's socially claimed self-image. Large language models (LLMs) are increasingly trained to refuse unsafe and inappropriate requests, and these refusals may harm users when models fail to manage this interactional cost properly. While existing work has mainly approached LLM non-compliance as a safety-alignment outcome, it does not provide a way to evaluate whether LLMs refuse appropriately across different harmful contexts. To study this question, we propose (to our knowledge) the first taxonomy of LLM refusals that is grounded in pragmatic theory. Applying this taxonomy to responses from 16 modern LLMs across 14 harm categories, we find that although models differ in how they refuse, their refusals are overall explicit and strongly morally evaluative, with interactional repair occurring mainly through offering or providing safer alternatives instead of interpersonal facework. This pattern is especially consequential in sensitive harm contexts, where overuse of negative framing may make users feel shamed or provoked, undermining the purpose of safe non-compliance. We therefore call for alignment evaluation that considers not only whether models refuse harmful requests, but also whether they refuse in ways that are contextually adaptive and socially accountable for the interactional consequences of saying no.
CommentsTo appear in the Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)