AI 中文总结
该研究基于动机访谈准则构建双轴评估框架,通过偏好优化训练LLM咨询师,发现惩罚对抗会降低目标坚持性,惩罚妥协无显著效果,仅提示控制可提升关系调谐且无目标坚持性代价。
AI 中文摘要
在动机访谈(MI)中,来访者的维持性谈话(支持现状的论点)要求咨询师顺应阻力,这一行为可能以两种相反的方式失败:妥协(放弃改变议程以维持融洽关系)或对抗(争辩或指导,凌驾于来访者的自主权之上)。我们基于动机访谈治疗完整性(MITI)准则,引入对咨询师回应的双轴评估:目标坚持性(GP)与关系调谐(RA),形成四象限框架,其中顺应阻力在两个维度上均表现良好,我们探究通过偏好优化惩罚某一失败类型,是否会教会模型顺应阻力或引发相反结果。我们从专家标注的AnnoMI语料库构建主题不相交的直接偏好优化数据,其偏好集仅在拒绝哪类失败上存在差异,采用在线策略负样本。一个经AnnoMI专家标签验证并由训练有素的人类编码员复核的自动评判器,在不相交模型家族生成、标注和评判的隔离环境下,对每个基础模型的盲法成对胜率进行评分。在Qwen和Llama家族的三个对齐指令模型中,惩罚对抗会在所有基础模型和所有种子运行中,可靠地将目标坚持性降至低于基准水平,这是一种稳健的代价;而调谐增益依赖于基础模型,在三个基础模型中的两个上存在,第三个上不存在。惩罚妥协无显著效果,因为这些模型在线策略中很少妥协,因此该权衡受每个基础模型的失败特征调控。仅提示控制可在不产生目标坚持性代价的情况下提升调谐,表明代价源于优化过程而非调谐本身。
英文摘要
In Motivational Interviewing (MI), a client's sustain talk (arguments for the status quo) calls for the counselor to roll with resistance, a move that can fail in two opposite ways: capitulation (abandoning the change agenda to preserve rapport) or confrontation (arguing or directing, overriding the client's autonomy). We introduce a two-axis evaluation of counselor responses, anchored in the Motivational Interviewing Treatment Integrity (MITI) code, Goal Persistence (GP) and Relational Attunement (RA), yielding a four-quadrant framing in which rolling with resistance is high on both, and we ask whether penalizing one failure through preference optimization teaches rolling with resistance or provokes its opposite. From the expert-annotated AnnoMI corpus we build topic-disjoint Direct Preference Optimization data whose preference sets differ only in which failure is rejected, using on-policy negatives. An automatic judge, validated against AnnoMI's expert labels and rechecked by trained human coders, scores blind pairwise win-rates against each base under a firewall in which disjoint model families generate, label, and judge. Across three aligned instruction models spanning the Qwen and Llama families, penalizing confrontation reliably lowers goal persistence below parity, on every base and in every seed run, a robust cost, whereas the attunement gain is base-dependent, present on two of the three bases but absent on the third. Penalizing capitulation is inert, because these models rarely capitulate on-policy, so the trade is gated by each base's failure profile. A prompt-only control raises attunement without the goal-persistence cost, locating the cost in the optimization rather than in attunement itself.