基于大语言模型的自动化临床行为编码:使用儿童BOSCC记录的案例研究
Towards Automated Clinical Behavioral Coding with Large Language Models: A Case study Using BOSCC recordings of Children
浏览论文内容
中文总结 AI 辅助
本研究针对BOSCC编码需专家操作的问题,评估大语言模型基于不同输入表征预测BOSCC编码的能力,发现其评估言语交流表现良好但识别非典型言语模式不足,仍存在编码标准应用等挑战。
中文摘要 AI 辅助
自闭症谱系障碍(ASD)是一种神经发育疾病,特征为社交沟通差异、兴趣受限及重复行为。治疗干预常以社交沟通技能为目标,因此需要可靠的行为变化测量指标。社交沟通变化简短观察量表(BOSCC)是一种经验证的治疗反应测量工具,基于儿童与经培训的检查者之间的简短游戏及社交沟通互动。BOSCC编码过程资源密集,需经培训的专家操作,因此需实现编码自动化以提升可扩展性与可及性。本研究评估通用大语言模型(LLMs)基于不同输入表征预测与言语相关的BOSCC编码的能力,在163份内部记录中比较文本、带说话人标注的文本及目标音频三种条件。LLMs在评估言语交流方面表现良好,但识别非典型言语模式的表现较差;不同评分决策的性能差异显著,且在不同诊断组间无一致模式。对模型预测的审查显示,应用BOSCC编码标准及解释模糊言语证据仍是挑战。
英文摘要
Autism spectrum disorder (ASD) is a neurodevelopmental condition characterized by differences in social communication and by restricted interests and repetitive behaviors. Treatment interventions often target social-communication skills, creating a need for reliable measures of behavioral change. The Brief Observation of Social Communication Change (BOSCC) is a validated treatment-response measure based on brief play and social-communication interactions between a child and trained examiner. The BOSCC coding process is resource-intensive and requires trained experts, motivating the automation of coding in order to improve scalability and accessibility. In this work, we evaluate general-purpose large language models (LLMs) for predicting speech-related BOSCC codes from different input representations. We compare transcript, diarized-transcript, and targeted audio conditions across 163 in-house recordings. LLMs are able to perform well in assessing verbal exchange, but do not perform as well when identifying atypical speech patterns. Additionally, performance varies considerably across scoring decisions, with no consistent pattern across diagnosis groups. An audit of model predictions indicates that applying the BOSCC coding criteria and interpreting ambiguous speech evidence remain challenges.