You Shouldn't Have Asked: A Pragmatics-Inspired Taxonomy for Evaluating LLM Refusals
你本不该问:一种受语用学启发的评估大语言模型(LLM)拒绝行为的分类法
机构 * Stony Brook University(石溪大学) ; University of Virginia(弗吉尼亚大学)
专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL
AI总结 该研究提出首个基于语用学理论的LLM拒绝分类法,通过分析16个LLM在14类有害请求下的拒绝回复,发现其拒绝的特点及存在的问题,呼吁开展兼顾情境适应性与社会问责的对齐评估。
Comments To appear in the Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)