发表机构
Lyncia Lab; Tianqiao and Chrissy Chen Institute(Lyncia实验室; 天桥与克里斯·陈研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对心理健康护理短缺,提出PsyEvo框架,通过个性化技能策略和响应优化在测试时自我进化,提升心理咨询效果,在PsychEval上获得7.684总分。
AI 中文摘要
心理健康障碍影响了全球相当大比例的人口,然而训练有素的从业者持续短缺,导致大多数人得不到充分的护理。基于大语言模型(LLM)的心理咨询师为提供可扩展的对话式心理支持展现了一个有前景的方向。仅靠离线模型训练,留给适应个体客户或从测试时的持续治疗互动中学习的空间有限。我们提出了PsyEvo,一个基于LLM的咨询框架,通过三个组件实现在测试时的客户特定个性化和响应策略改进:分层贝叶斯技能策略(HBSP)通过维护从会话反馈更新的每个客户技能后验来个性化应用何种干预;会话间列表式偏好优化(LiPO)通过从跨客户偏好证据更新共享响应适配器来改进所选技能的表达方式;状态条件序数信用分配(SOCA)通过一致性检查的比较和序数投影向这两个组件提供候选偏好和轨迹信用。在模拟客户评估中,采用共享在线队列自适应,PsyEvo在PsychEval上获得7.684的总分,并在三次匹配运行中均超过每个组件变体。在共享配置下,移除单个组件会使平均总分降低0.138至0.171,支持完整框架内的条件贡献。我们的代码可在以下网址获取:此https URL
英文摘要
Mental health disorders affect a substantial proportion of the global population, yet a persistent shortage of trained practitioners leaves the majority without adequate care. Large language model (LLM)-based counselors present a promising direction for delivering scalable conversational psychological support. Offline model training alone leaves limited room to adapt to individual clients or to learn from ongoing therapeutic interaction at test time. We introduce PsyEvo, an LLM-based counseling framework that enables both client-specific personalization and response-policy improvement at test time through three components: Hierarchical Bayesian Skill Policy (HBSP) personalizes what intervention to apply by maintaining a per-client skill posterior updated from session feedback; Inter-session Listwise Preference Optimization (LiPO) improves how the selected skill is expressed by updating a shared response adapter from cross-client preference evidence; and State-conditioned Ordinal Credit Assignment (SOCA) supplies candidate preferences and trajectory credit to the two components through consistency-checked comparisons and ordinal projection. In simulated-client evaluation with shared online cohort adaptation, PsyEvo obtains 7.684 Overall on PsychEval and exceeds every component variant in each of three matched runs. Removing individual components lowers mean overall score by 0.138--0.171 under the shared configuration, supporting conditional contributions within the complete scaffold. Our code is available at https://github.com/Lingxi-mental-health/PsyEvo