arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向物质使用障碍患者对话生成的多目标对齐小型语言模型框架

Multi-Objective Aligned Small Language Model Framework for SUD Patient Dialogue Generation

Thushara Manjari Naduvilakandy, Hyeju Jang, Mohammad Al Hasan

arXiv 2610.09209首次发表:更新:

发表机构

Indiana University Indianapolis(印第安纳大学印第安纳波利斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对物质使用障碍咨询中患者对话生成,提出认知对齐的小型语言模型框架,通过知识蒸馏、偏好优化和奖励塑形,提升认知实现与对齐效果。

AI 中文摘要

物质使用障碍(SUD)咨询需要患者回应能够反映其潜在认知状态,如信念、应对策略和改变准备度。尽管大型语言模型(LLMs)能够生成流畅的文本,但它们往往无法产生认知连贯且临床逼真的患者行为,尤其是在伦理和数据稀缺的临床环境下。此外,在医疗应用中部署前沿规模的LLMs面临实际挑战,包括高计算成本、延迟、隐私问题以及在资源受限环境中部署能力有限,这促使了对认知对齐的小型语言模型(SLMs)的需求。我们提出了一种基于认知的SUD患者对话生成框架,该框架显式地建模并将潜在认知成分与患者病史和咨询师问题对齐。我们的流程包括两个阶段:认知成分检测和认知成分对齐的对话生成。为了利用较小模型实现有效学习,我们结合了来自高能力教师模型的知识蒸馏、基于人类标注的偏好优化以及注意力引导的奖励塑形。使用自动评分(如BERTScore、ROUGE、METEOR和BLEU)以及LLM作为评判者的命中指标,针对人类和教师模型参考进行的广泛评估表明,认知信息微调显著提高了认知实现和对齐程度,优于通用指令微调基线和心理健康领域特定的SLMs,尤其在开放式认知成分上取得了显著提升。

英文摘要

Substance Use Disorder (SUD) counseling requires patient responses that reflect underlying cognitive states such as beliefs, coping strategies, and readiness for change. Although large language models (LLMs) can generate fluent text, they often fail to produce cognitively coherent and clinically realistic patient behavior, especially under ethical and data-scarce clinical settings. Moreover, deploying frontier-scale LLMs in healthcare applications presents practical challenges including high computational cost, latency, privacy concerns, and limited deployability in resource-constrained environments, motivating the need for cognitively aligned small language models (SLMs). We propose a cognitively grounded framework for SUD patient dialogue generation that explicitly models and aligns latent cognitive components with patient histories and counselor questions. Our pipeline consists of two stages: cognitive component detection and cognitive component-aligned dialogue generation. To enable effective learning with smaller models, we combine knowledge distillation from high-capacity teacher models, preference optimization from human-annotations, and attention-guided reward shaping. Extensive evaluations using automatic scores like BERTScore, ROUGE, METEOR and BLEU, and LLM-as-judge hit-metrics against both human and teacher-model references show that cognitively informed fine-tuning substantially improves cognitive realization and alignment over a generic instruction-tuned baselines and mental health domain specific SLMs, with particularly strong gains for open-ended cognitive components.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑