arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

检测≠可靠控制:可解码的共情方向在自动共情评分中仅产生最多部分偏移

Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

Haoran Jisun

arXiv 2608.24901首次发表:更新:

发表机构

University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对EPITOME的认知与情感共情维度,在三个指令微调LLM中发现,可解码的共情方向仅能部分控制自动共情评分,认知维度的控制不可靠,需对认知共情主张做测量敏感性检验。

AI 中文摘要

可解码的“共情”方向常被视为因果杠杆,混淆了可解码性、自动度量控制与人类感知变化。本研究针对EPITOME衍生的两个维度——认知维度的识别(Recognition)与情感维度的共鸣(Resonance),在三个指令微调大语言模型(LLM)中进行测试,采用两个LLM评判器及判别式EPITOME分类器对每个干预措施评分,各评分工具均由情感vs中性的阳性控制项门控。所有自动工具对情感维度的控制均通过,但各工具对认知维度的控制范围不一致。对句子嵌入衍生的表面得分进行残差化处理后,两个维度仍可解码,且引导可大幅改写文本。然而添加共鸣方向仅部分提升情感评分:Qwen提升0.29(约为自然差距的26%)。方向间直接对比证实,Qwen和Llama(非Gemma)的偏移具有维度特异性,但未确立对应的人类感知变化。累加认知引导无显著变化,但域内控制显示认知工具过于粗糙,无法分辨此类引导会产生的差异——属于不可测量,而非明确的零结果。相比之下,Gemma的识别维度消融后,即使调整响应长度,分类器的认知评分仍降低。研究表明,全局干预下检测并不意味着可靠控制,认知共情相关主张需进行明确的测量敏感性检验。

英文摘要

A decodable "empathy" direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change. We test this for two EPITOME-derived facets -- Recognition (cognitive) and Resonance (affective) -- in three instruction-tuned LLMs, scoring every intervention with two LLM judges and a discriminative EPITOME classifier, each gated by an emotional-vs-neutral positive control. The control passes for the affective facet across all automated instruments, but cognitive range is inconsistent across them. Both facets remain decodable after residualizing against a sentence-embedding-derived surface score, and steering can substantially rewrite the text. Yet adding the Resonance direction raises the affective score only partially -- in Qwen by +0.29 (approximately 26% of the natural gap). A direct between-direction contrast confirms the shift is facet-specific in Qwen and Llama (not Gemma); we do not, however, establish a matching human-perceived change. Additive cognitive steering produces no measurable change, but a within-domain control shows the cognitive instrument is too coarse to resolve the differences such steering would produce -- unmeasurable, not a clean null. By contrast, Gemma Recognition ablation lowers the classifier's cognitive score even after adjusting for response length. Detection does not imply reliable control under global interventions, and cognitive-empathy claims warrant an explicit measurement-sensitivity check.

CommentsUnder review at BlackboxNLP 2026 (EMNLP). 8 pages body, 10 figures/tables, plus appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑