arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21857cs.CLcs.AI

个性调优的大语言模型能成为更好的社交智能体吗?

Do Personality-Tuned LLMs Make Better Social Agents?

Tim Krabbe, Xiaodan Shi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过微调两个小型开源大语言模型,评估个性感知微调是否优于指令提示以提升社交模拟中的个性一致性与可控性,结果显示微调模型未优于基线,但提高了语言多样性。

中文摘要 AI 辅助

大语言模型越来越多地被用于社交模拟,以构建社交互动智能体和机器人,相比基于规则的系统提供了更大的灵活性。然而,尽管它们能很好地模仿人类行为,但始终存在一种难以消除的异质感。本研究探讨了与仅使用指令提示相比,个性感知的微调是否能通过提高个性条件对话生成的一致性和可控性来缩小这一差距。我们使用一个结合了带个性标签的社交媒体帖子和对话的语料库,对两个小型开源大语言模型Qwen2.5-7B-Instruct和Ministral-8B-Instruct进行微调,以构建用于社交模拟的基于个性的对话引擎。所得到的模型在多个社交互动场景中,使用三个独立的大语言模型评判员进行评估,这些评判员评估个性保真度并提供基于证据的行为解释。我们还量化了评判员间的一致性和生成对话的词汇特征。结果表明,微调后的模型在扮演不同个性方面并不优于其各自的基线模型。然而,较低的评判员间一致性限制了这些结果的可信度。关于生成文本的质量,微调后的模型大多与基线相当,其中微调提高了Qwen模型的语言多样性。虽然结果总体上看起来可用,且基线模型提供了最佳的整体性能,但未来的研究应更加重视训练数据的质量和领域对齐,以实现准确的个性角色扮演。

英文摘要

LLMs are increasingly used in social simulations for socially interactive agents and robots, offering more flexibility than rule-based systems. However, even though they mimic human behaviour very well, there is a persistent alienness to them. This work investigates whether personality-aware fine-tuning can reduce this gap by improving the consistency and controllability of personality-conditioned dialogue generation compared with instruction prompting alone. We fine-tune two small open-weight LLMs, Qwen2.5-7B-Instruct and Ministral-8B-Instruct, using a corpus that combines personality-labelled social media posts and dialogues to create a personality-based dialogue engine for social simulation. The resulting models are evaluated across multiple social interaction scenarios using three independent LLM judges, which assess personality fidelity and provide evidence-based behavioral interpretations. We additionally quantify inter-rater agreement and lexical characteristics of the generated dialogue. Results indicate that fine-tuned models are not better at role-playing different personalities than their respective baseline models. However, low inter-rater agreement limits the confidence with which these results can be interpreted. Concerning the quality of generated texts, fine-tuned models are mostly comparable to the baselines, with fine-tuning improving the linguistic diversity of the Qwen models. While the results appear generally usable and the baseline models offer the best overall performance, future studies should place greater emphasis on the quality and domain alignment of training data for accurate personality role-playing.

发表机构

  • Stockholm University(斯德哥尔摩大学)

机构由 AI 辅助整理,请以论文原文为准。

↑