arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种通过基于评分准则的剧本归纳构建的双语AI听力学家,在模拟病例盲评中表现优于人类听力学家

A bilingual AI audiologist built through rubric-guided playbook induction outperforms human audiologists in a blinded evaluation of simulated cases

Linkai Li, Changgeng Mo, Hanlin Yu, Congxi Lu, Shangqiguo Wang, Matthew B Fitzgerald, Shan X Wang

arXiv 2609.32220首次发表:更新:

发表机构

Stanford University; Orka Labs Inc.; The University of Hong Kong(斯坦福大学; 奥卡实验室公司; 香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种基于评分准则归纳的双语AI听力学家,结合大语言模型与多模态听力图解读,在盲评中58个模拟病例全部超越人类听力学家,为低数据医学领域提供实用方案。

AI 中文摘要

听力咨询需要结构化的病史采集、听力图解读和以患者为中心的沟通,然而现实世界中的病例材料十分稀缺。我们提出了一种双语AI听力学家,将通用大语言模型与基于评分准则的剧本归纳、多模态听力图解读和检索增强的接地相结合,且无需对语言模型主干进行微调。利用一份包含21个项目的评分准则和一个AI患者模拟器,我们从73个训练病例(43个英语、30个中文)中归纳出一个包含19条规则的咨询策略,并在预先指定、来源盲法的比较中,针对58个独立模拟病例(30个中文、28个英语)对系统进行了评估,与17名执业听力学家进行了对比。AI听力学家在每个病例上均优于人类听力学家(58/58;在5分制综合评分上平均配对差异Δ=+1.35,Cohen's d=1.84,P=4.5×10^-20),在21个评分准则项目中的20个上以及两种语言中均表现更优。组件消融实验表明,剧本是最大的贡献因素,为低数据医学领域中的专科咨询智能体提供了一条实用途径。

英文摘要

Audiology consultation requires structured history-taking, audiometric interpretation and patient-centred communication, yet real-world case material is scarce. We present a bilingual AI audiologist pairing a general-purpose large language model with rubric-guided playbook induction, multimodal audiogram interpretation and retrieval-augmented grounding, without fine-tuning the language-model backbone. Using a 21-item rubric and an AI patient simulator, we induced a 19-rule consultation policy from 73 training cases (43 English, 30 Chinese) and evaluated the system on 58 independent simulated cases (30 Chinese, 28 English) in a pre-specified, source-blinded comparison with 17 practising audiologists. The AI audiologist outperformed human audiologists on every case (58/58; mean paired $Δ$ = +1.35 on a 5-point composite, Cohen's d = 1.84, $P = 4.5 \times 10^{-20}$), on 20 of 21 rubric items and in both languages. Component ablation identified the playbook as the largest contributor, offering a practical route to specialist consultation agents in low-data medical domains.

Comments75 pages, including 51 pages of Supplementary Information

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑