arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

以指南为神谕:眼科电话分诊智能体的零标注训练

Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent

Chenyu Wang, Yi Liu, Baoqing Li, Min Tu, Diping Song

arXiv 2608.04772首次发表:更新:

AI 中文总结

本研究提出以指南为神谕(GAO)方法,将美国眼科学会指南转化为训练监督信号,零标注训练出GAO-Triage智能体,大幅提升眼科电话分诊的一致性与紧急案例召回率,且性能优于7个通用系统。

AI 中文摘要

为多轮医疗智能体扩充监督信号难度较大,因为专家对话标注成本高昂,且临床对话受隐私限制。我们提出以指南为神谕(Guideline-as-Oracle, GAO),将美国眼科学会指南汇编成一张70行的操作规则表,并将其作为3000条训练对话的实例级监督的唯一来源,仅保留人工标注用于评估。由于将规则转换为对话本身是一个设计问题,我们梳理了8种构建策略,包括引用行层级分配、单事实边界对、仅元数据修复和标签修复,并对每种策略的证据状态进行了刻画:标注机制、无效、混杂或仅作为整体评估。在该语料库上微调9B主干模型得到GAO-Triage,其与201个案例操作参考的一致性从61.7%提升至74.1%(精确McNemar检验p值为0.0046),紧急案例召回率从9.5%提升至69.0%;在第二个随机种子和患者模拟器上,该提升依然存在。我们测试的7个通用系统均未在两个指标上同时优于GAO-Triage,且GAO-Triage在推理时不需要前沿模型。打乱标签-对话分配会使模型退化为常规模型预测器,表明信号来自指南衍生的分配而非对话表面形式。标签修复与训练后期安全退化的消失相吻合。

英文摘要

Scaling supervision for multi-turn medical agents is difficult because expert dialogue annotation is costly and clinical conversations are privacy-restricted. We introduce Guideline-as-Oracle (GAO), which compiles American Academy of Ophthalmology guidance into a 70-row operational rule table and uses it as the sole source of instance-level supervision for 3,000 training dialogues, reserving human labeling for evaluation. Because converting rules into dialogues is itself a design problem, we catalog eight construction strategies, including cited-row tier assignment, one-fact boundary pairs, metadata-only repair, and label repair, and characterize the evidential status of each: labeling mechanism, null, confounded, or evaluated only as a package. Fine-tuning a 9B backbone on this corpus yields GAO-Triage, improving agreement with a 201-case operational reference from 61.7% to 74.1% (exact McNemar p=0.0046) and emergent-case recall from 9.5% to 69.0%; the gains persist across a second seed and patient simulator. None of the seven general-purpose systems we test dominates GAO-Triage on both metrics, and GAO-Triage requires no frontier model at inference time. Permuting label-dialogue assignments collapses the model to a constant-routine predictor, indicating that the signal lies in guideline-derived assignment rather than dialogue surface form. Label repair coincides with the disappearance of a late-training safety degradation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑