arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18691cs.AIcs.CL

语义原素作为大语言模型中情感的解释因素

Semantic Primes as Explanans for Emotion in Large Language Models

Frank Xing

首次发表
浏览论文内容

中文总结 AI 辅助

研究探讨大语言模型中情感的解释问题,通过在四个指令微调的LLMs上实验,发现自然语义元语言的语义原素是可恢复的内部元素,能更有效控制情感,且与相应情感可互换,是比其他选项更好的情感解释因素。

中文摘要 AI 辅助

在理解大语言模型(LLMs)的情感机制方面已取得进展。然而,如何解释LLMs中的情感,甚至什么构成好的解释,尚不清楚。情感表征、成分和回路可广泛恢复,但作为模型自身计算的解释是循环的;情感空间维度往往是任意且无终止的。一个紧迫问题是更原始的内部变量集——自然语义元语言(NSM)的语义原素——是否能起作用。在四个指令微调的LLMs(Llama - 1B、Gemma - 2B、Gemma - 9B、OLMo - 7B)上的实验表明,NSM原素是可恢复的内部元素;在参考模型上,基于原素的方向干预对情感的控制强度约为基于最佳评估方向的三倍,选择性为两倍;模型将基于原素的解释与相应情感视为可互换的。这些证据表明,根据科学解释标准,NSM原素似乎比许多其他选项更适合作为LLMs中情感的解释因素。

英文摘要

Progresses have been made on understanding emotion mechanisms of large language models (LLMs). However, how to explain emotion in LLMs, or even what constitutes good explanations, are less clear. Emotion representations, components, circuits are widely recoverable, but as explanations of a model's own computation they are circular; the emotion space dimensions tend to be arbitrary and non-terminating. A pressing question to ask is whether a more primitive set of internal variables does the work: the semantic primes of the Natural Semantic Metalanguage (NSM). Across four instruction-tuned LLMs (Llama-1B, Gemma-2B, Gemma-9B, OLMo-7B), experiments show that the NSM primes are (1) recoverable internal elements; and (2) on the reference model, intervening with a prime based direction controls emotion about three times as strongly, and twice as selectively, as the best appraisal based direction; and (3) the model treats a prime based explication as interchangeable with the corresponding emotion. These evidences suggest that NSM primes seem to be better explanans for emotion in LLMs than many alternative options according to scientific explanations criteria.

补充信息

↑