arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RELATE:用于衡量大型语言模型关系导向的评估框架

RELATE: An Evaluation Framework for measuring Relational Orientation of Large Language Models

Shivam Shukla, Jihye Kim, Shubham Gaur, Mahnaz Roshanaei, Magy Seif El-Nasr

arXiv 2610.09569首次发表:更新:

发表机构

University of California, Santa Cruz; Stanford University(加州大学圣克鲁兹分校; 斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出关系导向概念及RELATE评估框架,通过句子级分析衡量LLM在多轮对话中的内向型与外向支架型语言,实验发现模型随时间更倾向内向导向,且对间接用户较少鼓励现实联系。

AI 中文摘要

大型语言模型(LLM)越来越多地被用于提供情感支持,这引发了担忧,即持续使用可能会将用户从他们的现实人际关系中拉离。然而,现有的评估主要关注回复的安全性、同理心或有用性,却很少审视一个关系性问题:模型将用户导向何处以获取持续支持?为了解决这个问题,我们引入了关系导向这一概念,并通过两个非排他性维度将其操作化:内向型(IF)语言,即将AI定位为用户持续支持的来源;以及外向支架型(OS)语言,即鼓励现实世界中的人际联系。基于心理学和社会学文献,我们形式化了一个关系导向的分类体系,并提出了RELATE,一个以角色为条件的框架,用于在多轮对话中在句子级别衡量内向型和外向支架型语言。RELATE将76个改编自自然出现问题的求助情境与三种模拟用户风格配对,提供了228个评估刺激。在我们的实验中,我们使用每个包含六轮助手回复的对话评估了七个LLM,生成了1,596个对话和69,194个助手句子。我们使用一个基于评分标准的主要LLM评判器评估这些句子,并对一个子集应用次要评判器。在自动化评估下,我们发现,在第六轮助手回复中被标记为IF的句子比例高于第一轮,而被标记为OS的句子比例对于犹豫不决、间接的模拟用户而言,显著低于明确寻求安慰的用户。RELATE提供了一个可复现的框架和句子级别的信号,用于审计和引导支持型LLM的关系导向。

英文摘要

Large language models (LLMs) are increasingly used for emotional support, raising concern that sustained use may draw users away from their real-world relationships. Yet existing evaluations primarily focus on the safety, empathy, or helpfulness of responses, leaving under-examined a relational question: where does the model orient the user for continued support? To address this question, we introduce relational orientation, a property operationalized through two non-exclusive dimensions: inward-facing (IF) language, which positions the AI as the user's ongoing source of support, and outward-scaffolding (OS) language, which encourages real-world human connection. Grounded in psychological and sociological literature, we formalize a taxonomy of relational orientation and present RELATE, a persona-conditioned framework for measuring inward-facing and outward-scaffolding language at the sentence level in multi-turn dialogues. RELATE pairs 76 help-seeking situations adapted from naturally occurring questions with three simulated user styles, providing 228 evaluation stimuli. In our experiments, we evaluate seven LLMs using dialogues with six assistant turns each, yielding 1,596 dialogues and 69,194 assistant sentences. We assess these sentences using a primary rubric-based LLM judge and apply a secondary judge to a subset. Under automated evaluation, we find that the proportion of sentences labeled as IF is higher at the sixth assistant turn than at the first, while the proportion labeled as OS is substantially lower for hesitant, indirect simulated users than for explicit, reassurance-seeking users. RELATE provides a reproducible framework and a sentence-level signal for auditing and steering the relational orientation of supportive LLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑