语言承载专家印象:基于工具锚定的LLM裁判迁移咨询质量评估并超越领域内训练
Language Carries the Expert's Impression: Instrument-Anchored LLM Judges Transfer Counseling-Quality Assessment and Beat In-Domain Training
浏览论文内容
中文总结 AI 辅助
本研究证明在咨询质量评估中,跨领域训练优于领域内训练,基于专家评分工具锚定的LLM裁判特征显著提升预测性能,且语言内容比非语言信号更关键。
中文摘要 AI 辅助
二元咨询对话中沟通质量的自动评估受数据瓶颈制约:专家评分的语料库规模小且扩充成本高昂。我们研究了专家总体印象预测在三个德语模拟咨询语料库(两个全科医疗、一个学校相关家长-教师;n=195个专家评分会话,一个语料库经量表等值化处理)上的跨领域迁移。在其他领域上训练优于在目标领域内训练:留一领域外迁移达到嵌套Spearman ρ=0.54,而目标领域内仅为≤0.48,配对会话级差距为+0.15,当训练集规模匹配时仍保持+0.12,因此这并非单纯的数据量效应。决定性特征来自读取双说话者转录文本的小型开放权重LLM生成的会话级构念分数,这些构念主要源自专家的评分工具:基于工具的特征组将单一裁判从0.32提升至0.41(相对于通用对话质量),来自三个模型家族的裁判集成达到0.51(仅语言),非语言二元模块额外增加+0.03,在此样本量下与噪声不可区分。我们还评估了录音设置:一个语料库丢失了其每说话者音频,其16%的分段说话者标注错误,修复该问题在该语料库上价值+0.07。在实际可达到的语料库规模下,专家的总体印象由所说内容承载,且更多由其他沟通项目的数据而非自身数据承载。
英文摘要
Automatic assessment of communication quality in dyadic counseling conversations is bottlenecked by data: expert-rated corpora are small and expensive to grow. We study cross-domain transfer of expert overall-impression prediction across three German corpora of simulated counseling (two general-practice medical, one school-related parent-teacher; $n=195$ expert-rated sessions, one corpus after scale equating). Training on the other domains beats training in-domain: leave-one-domain-out transfer reaches nested Spearman $ρ= 0.54$ against $\le 0.48$ within the target domain, a paired session-level gap of $+0.15$ that holds at $+0.12$ when the training-set sizes are matched, so it is not simply data volume. The decisive features are session-level construct scores from small open-weight LLMs reading the two-speaker transcript, with the constructs largely derived from the experts' rating instruments: the instrument-derived battery lifts a single judge from $0.32$ to $0.41$ over generic dialogue qualities, judges from three model families ensemble to $0.51$ language-only, and a nonverbal-dyadic block adds $+0.03$ more, not separable from noise at this sample size. We also price the recording setup: one corpus lost its per-speaker audio, 16% of its diarised segments carry the wrong speaker, and repair is worth $+0.07$ there. At practically attainable corpus sizes, the expert's overall impression is carried by what is said, and by other communication programs' data more than by one's own.
发表机构
- University of Augsburg(奥格斯堡大学)
机构由 AI 辅助整理,请以论文原文为准。