arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

少即是多:心理咨询对话中极简回应与大语言模型(LLM)行为的实证研究

When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and the Behavior of LLMs

Zhiyang Qi

arXiv 2608.24080首次发表:更新:

发表机构

The University of Tokyo(东京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过实证研究发现,心理咨询对话中常见的极简回应在LLM生成内容中占比极低,商业LLM可按指令生成此类回应但难判断适用时机,专用模型表现差且基于LLM的评估易低估其价值。

AI 中文摘要

在心理咨询中,有效的支持并非总是通过冗长且信息丰富的回应传递。极简回应(如反馈信号和简洁的共情陈述)有助于传递专注倾听、表达共情,并鼓励来访者继续表达。然而,现有的心理咨询对话系统和评估框架往往偏向明确、内容丰富的回复,忽视了咨询师简短话语的互动价值。本文对多个心理咨询对话数据集中的极简回应开展了系统的跨语言分析,提出了一种基于话语长度和内容的两阶段过滤方法,随后利用大语言模型(LLM)进行上下文验证。分析显示,极简回应在人工收集的数据集中较为常见,但在LLM生成的回应中占比极低。我们进一步在人工整理的、人类咨询师使用极简回应的对话情境中评估当前LLM,结果表明,强大的商业LLM在明确指示时能够生成极简回应,但仍难以判断此类回应的适用时机;基于合成数据训练的心理咨询专用模型表现尤其差,倾向于生成更长、信息更丰富的回应。此外,基于LLM的回应质量评估可能会低估极简回应的价值,即便它们在互动上是恰当的。

英文摘要

In psychological counseling, effective support is not always delivered through long, information-rich responses. Minimal responses, such as backchannel cues and concise empathic statements, help convey attentive listening, express empathy, and encourage clients to continue expressing themselves. However, existing counseling dialogue systems and evaluation frameworks often favor explicit, content-rich replies, overlooking the interactional value of brief counselor utterances. This paper presents a systematic cross-lingual analysis of minimal responses across multiple counseling dialogue datasets. We develop a two-stage filtering method based on utterance length and content, followed by contextual verification using a large language model (LLM). Our analysis shows that minimal responses are common in human-collected datasets but substantially underrepresented in LLM-generated ones. We further evaluate current LLMs in manually curated dialogue contexts where human counselors used minimal responses. The results show that strong commercial LLMs are capable of generating minimal responses when explicitly instructed, but still struggle to determine when such responses are appropriate. Counseling-specific models trained on synthetic data perform particularly poorly, tending instead to produce longer and more information-rich responses. Moreover, LLM-based response-quality evaluation may undervalue minimal responses, even when they are interactionally appropriate.

CommentsAccepted to EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑