arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.29685cs.AI

NICE:一个基于理论的LLM社交智能诊断基准

NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs

  • Department of Psychology and Behavioral Sciences, Zhejiang University(浙江大学心理学与行为科学系)
  • College of Artificial Intelligence, Zhejiang University(浙江大学人工智能学院)
  • Human Machine Interaction Lab, Huawei Technologies Co., Ltd.(华为技术有限公司人机交互实验室)
  • Zhejiang Key Laboratory of Neurocognitive Development and Mental Health(浙江省神经认知发展与心理健康重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

Yunjin Qi, Zhaojun Jiang, Xuan Wu, Hanxi Pan, Yixuan Wang, Yanfang Liu, Xiang Ji, Churu Yu, Chunyuan Zheng, Yingze Chen, Jie He, Liuqing Chen, Zaifeng Gao

更新

AI总结:

本文通过构建基于社会理论的社交智能框架,提出诊断基准NICE,用于细粒度评估大语言模型在社交交互中的能力弱点。

AI中文摘要:

随着大语言模型(LLM)在情感陪伴和客户服务等社交场景中的广泛应用,衡量其社交智能对人工智能交互的质量与安全性变得至关重要。然而,现有的社交智能基准缺乏统一框架来组织社交能力,因此无法进行细粒度诊断。为了构建首个基于社会理论的整体诊断评估,我们首先通过文献综述和多阶段专家验证(遵循心理测量学原则)构建了一个社交智能框架。该框架包括4个类别和11个维度,每个维度进一步由细粒度的能力方面指定。基于此框架,我们提出了NICE(规范、交互、认知、体验),一个包含137个项目的诊断基准,通过代表性中文情境进行操作化。在5个前沿LLM和一个人类参考组中,模型在总体准确率上得分较高,但在沟通方面表现出持续的弱点,框架将其定位到三个具体能力方面:多轮沟通、非语言沟通和同步性。因此,NICE将社交智能评估重新定义为对LLM中具有社会后果的弱点的基于理论的诊断。

英文摘要:

As large language models (LLMs) are increasingly applied in social contexts such as emotional companionship and customer service, measuring their social intelligence has become critical to the quality and safety of human-AI interaction. However, existing social intelligence benchmarks lack a unified framework that organizes social abilities into a unified structure, and therefore cannot enable fine-grained diagnosis. To build the first holistic diagnostic evaluation grounded in social theory, we first construct a social intelligence framework through a literature review and multi-stage expert validation guided by psychometric principles. The resulting framework includes 4 categories and 11 dimensions, each further specified by fine-grained capability facets. Building on this framework, we introduce NICE (Norm, Interaction, Cognition, Experience), a diagnostic benchmark of 137 items operationalized through representative Chinese contexts. Across 5 frontier LLMs and a human reference group, models score higher in aggregate accuracy yet show a consistent weakness in Communication, which the framework localizes to 3 specific capability facets: multi-turn communication, nonverbal communication, and synchrony. NICE thus reframes social intelligence evaluation toward theory-grounded diagnosis of socially consequential weaknesses in LLMs.

↑