Where Norms and References Collide: Evaluating LLMs on Normative Reasoning
规范与参照的碰撞:评估大语言模型在规范推理中的能力
AI总结 本文提出SNIC测试平台,评估大语言模型在基于规范的参照解析任务中的能力,发现现有模型在处理隐含或冲突规范时存在显著不足。
Comments Accepted to the 40th AAAI Conference on Artificial Intelligence (AAAI-26)