arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22425cs.HCcs.CLcs.CY

四大领先的大语言模型在与经人格验证的合成求助者交流时,说得多听得少

All four leading LLMs talk more than they listen to personality-verified synthetic help-seekers

Pablo A. Fonseca, Raquel Rodríguez-Carvajal, Rafael A. Calvo

首次发表
浏览论文内容

中文总结 AI 辅助

该研究构建人格感知评估,发现四大领先大语言模型在为合成求助者提供危机建议时,存在说多听少、未充分探索情况就解决问题的共性,且评估覆盖了大五人格外的心理倾向。

中文摘要 AI 辅助

在人们陷入困境时,越来越多地会向大语言模型求助,但单轮基准测试既未测试持续的交流,也未区分不同用户。我们构建了一种人格感知评估,其中四个广泛使用的模型为几名合成求助者提供建议,每个求助者都有经心理测量学明确的急性危机档案:一名照料者得知亲属被诊断为痴呆症。对档案提示不知情的审核员仅通过对话就以极高的一致性从对话中恢复了指定的量表,每种工具的组内相关系数(ICC(2,4))为0.91;按工具划分的区间为0.79-0.96;区间得分相关系数r为0.78,这符合大五人格的预期,也适用于应对风格、应对自我效能、心理韧性和逆反性,而这些是词汇方法从未涵盖的。因此,此类评估超越了大五人格模型,覆盖了动机、调节和自我评估倾向。四个模型在情绪稳定方面没有差异,且表现相似,共享三种模式:冗长、说听比高于1,以及在未充分探索情况前就进行问题解决。

英文摘要

Large language models are increasingly consulted at moments of distress, yet single-turn benchmarks neither test sustained exchanges nor distinguish between users. We built a personality-aware evaluation in which four widely used models advised several synthetic help-seekers, each given a psychometrically specified profile, in an acute crisis: a caregiver learning of a relative's dementia diagnosis. Auditors blind to the profile prompt recovered the specified bands from dialogue alone with high agreement on every instrument (ICC(2,4) = 0.91; 0.79-0.96 by instrument; band-score r = 0.78), as expected for the Big Five but equally for coping style, coping self-efficacy, resilience and reactance, which the lexical approach never covered. Such evaluation therefore reaches beyond the Five Factor Model to motivational, regulatory and self-appraisal dispositions. The four models were not distinguishable on emotion stabilisation and failed alike, sharing three modes: verbosity, a talk-to-listen ratio above one, and problem-solving before the situation had been explored.

发表机构

  • Dyson School of Design Engineering, Imperial College London(帝国理工学院戴森设计工程学院)
  • Universidad Autónoma de Madrid(马德里自治大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑