arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StageWell:面向积极心理学支持对话的流程对齐中文语料库

StageWell: A Process-Aligned Chinese Corpus for Positive-Psychology Support Dialogue

Yuxiong Wang, Ziwei Lin, Bo Wang, Yu Zhang, Shiguang Ni

arXiv 2608.29326首次发表:更新:

发表机构

Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建了流程对齐中文语料库StageWell及HQS协议,用于优化多轮积极心理学支持对话的监督,在多个开源大模型上实现了流程控制、回复质量与安全性的提升。

AI 中文摘要

积极心理学对话旨在支持情绪困扰并构建积极资源,要求模型不仅生成共情回复,还需在多轮支持过程中保持连贯进展。现有资源常将监督简化为轮级策略或整体偏好标签,导致流程位置、支持功能及局部修复目标隐含未明。我们提出StageWell,一种面向积极心理学对话的流程对齐中文语料库,以及用于数据构建与评估的结构化协议HQS。StageWell将支持过程组织为六个阶段,采用多智能体全对话重写工作流构建了12445个SFT实例、1849个DPO偏好对,以及包含120个专家修订对话和977个问答对的GroundTruth子集。在HQS指导下,DPO对构建为流程局部修复:将有缺陷的模型输出作为拒绝回复,在相同上下文和阶段约束下的针对性重写作为选择回复。在四个9B至14B规模的开源LLMs上,该监督信号在流程控制、回复质量和安全性方面均取得显著提升:模型平均BERTScore提升0.037,Q-Overall增加1.32分,S-exact提升0.236,H-critical率降低0.167。这些结果表明,将支持性对话建模为结构化多轮支持过程而非单轮回复生成具有重要价值。

英文摘要

Positive psychology dialogue aims to support emotional distress and positive resource building, requiring models to produce not only empathetic replies but also coherent progression through a multi-turn support process. Existing resources often reduce supervision to turn-level strategies or holistic preference labels, leaving process position, support function, and local repair targets implicit. We introduce StageWell, a process-aligned Chinese corpus for positive psychology dialogue, together with HQS, a structured protocol for data construction and evaluation. StageWell organizes support into a six-stage support process and uses a multi-agent whole-dialogue rewriting workflow to construct 12,445 SFT instances, 1,849 DPO preference pairs, and a GroundTruth subset of 120 expert-revised dialogues and 977 QA pairs. Guided by HQS, DPO pairs are built as process-localized repairs: flawed model outputs are used as rejected responses, and targeted rewrites under the same context and stage constraint are used as chosen responses. Across four 9B-14B open-source LLMs, this supervision yields robust gains in process control, response quality, and safety. Averaged across models, BERTScore improves by 0.037, Q-Overall increases by 1.32 points, S-exact increases by 0.236, and the H-critical rate decreases by 0.167. These results highlight the value of modeling supportive dialogue as a structured multi-turn support process rather than as single-turn response generation.

Comments29 pages, 20 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑