AI 中文总结
DeepSAGE是融合LLM与DRL的阶段感知咨询对话框架,经评估在阶段目标完成与对话效率上优于6种替代方案,为AI咨询提供了有前景的新路径。
AI 中文摘要
基于大语言模型(LLM)的咨询智能体可生成流畅且具支持性的回复,但常缺乏开展连贯治疗会话所需的结构化、目标导向式进展。我们提出DeepSAGE(Strategic AI Guidance Engine,战略AI引导引擎),这是一种混合LLM-深度强化学习(DRL)框架,用于基于认知行为疗法(CBT)首次会话的阶段感知咨询对话。DeepSAGE将会话表示为11个具有明确治疗目标的阶段,外部控制器确定阶段完成情况,DRL模型选择指导LLM回复生成的治疗意图。我们将DeepSAGE与6种基于检索、提示、阶段和策略的替代方案进行评估,结果显示DeepSAGE能引发更高的模拟客户参与度与开放性,在阶段结构化系统中实现了阶段目标完成度与对话效率的最佳平衡。领域专家评审进一步表明,生成的对话呈现出大致合理的情绪轨迹与可识别的CBT流程。由于评估主要依赖模拟客户与基于模型的指标,这些发现仅体现了对话控制的比较改进,而非临床有效性。这些结果表明,将阶段结构化对话与学习到的策略选择相结合是AI咨询的一种有前景的方法,不过临床有效性、安全性与实际应用价值仍需进一步的人工评估。
英文摘要
Large Language Model (LLM)-based counseling agents can generate fluent and supportive responses, but they often lack the structured, goal-directed progression required to conduct a coherent therapeutic session. We present DeepSAGE (Strategic AI Guidance Engine), a hybrid LLM--Deep Reinforcement Learning (DRL) framework for stage-aware counseling dialogue grounded in the first session of Cognitive Behavioral Therapy (CBT). DeepSAGE represents the session as eleven stages with explicit therapeutic objectives, with an external controller determines stage completion and the DRL model selects therapeutic intentions that guide LLM response generation. We evaluate DeepSAGE against six retrieval-, prompting-, stage-, and policy-based alternatives. DeepSAGE elicits higher simulated client engagement and openness and achieves the strongest balance of stage-goal completion and dialogue efficiency among stage-structured systems. Domain expert review further indicates that the generated conversations exhibit broadly plausible emotional trajectories and recognizable CBT processes. Because the evaluation relies primarily on simulated clients and model-based metrics, these findings demonstrate comparative dialogue-control improvements rather than clinical effectiveness. These results suggest that combining stage-structured dialogue with learned strategy selection is a promising approach for AI counseling, though clinical effectiveness, safety, and real-world utility require further human evaluation.