AI 中文总结
本研究通过后见之明引导蒸馏训练罕见病诊断模型,发现污染过滤(而非蒸馏本身)带来微小但显著的准确率提升,并识别出GT幻觉现象及校准差距等未来研究方向。
AI 中文摘要
我们研究了在ZebraMap数据集上针对罕见病诊断的后见之明引导蒸馏方法:一个1.5B参数的学生模型在8B参数教师模型生成的思维链轨迹上进行微调,而该教师模型在生成过程中能够观察到真实诊断结果。所有模型的绝对准确率仍然较低——在此规模下任务难度较大——但在此上限内,经过过滤的变体(StudentF)相比教师模型取得了微小但统计显著的准确率优势(p < 0.001),且该优势集中在代表性更好的疾病上。未过滤的学生模型并未显著优于教师模型(p = 0.129),这证实了污染过滤(而非单纯的后见之明蒸馏)是带来增益的关键因素。这一差距可追溯至我们称之为“GT幻觉”的伪影。标签可见的生成过程导致教师模型在其推理链中嵌入“真实结果为X”的短语;SFT(监督微调)复制了这一模式。在推理时,未过滤的学生模型在33.9%的情况下复现该短语,当幻觉标签错误时,准确率严重下降。一个正则表达式过滤器移除这些槽位后,将污染降至接近零,从而产生了观察到的增益——尽管效果仍然较小。我们精确量化了这种增益-成本权衡,记录了在RL(强化学习)训练的教师模型中缺失的频率依赖性知识迁移,并刻画了SFT未能弥合的校准差距——将这两点均确定为未来工作的方向。
英文摘要
We study hindsight-guided distillation for rare disease diagnosis on ZebraMap: a 1.5B student is fine-tuned on chain-of-thought traces from a 8B teacher that observes the ground-truth diagnosis during generation. Absolute accuracy remains low for all models - the task is hard at this scale - but within this ceiling a filtered variant (StudentF) achieves a small, statistically significant accuracy advantage over the teacher (p < 0.001), concentrated in better-represented diseases. The unfiltered student does not significantly outperform the teacher (p = 0.129), establishing that contamination filtering - not hindsight distillation alone - drives the gain. The gap traces to an artifact we term GT hallucination. Label-visible generation causes the teacher to embed "ground truth is X" phrases in its reasoning chain; SFT copies the pattern. At inference, the unfiltered student reproduces the phrase in 33.9% of cases, with severe accuracy degradation when the hallucinated label is wrong. A regex filter removing these slots reduces contamination to near-zero, producing the observed gain - though the effect remains small. We precisely quantify this gain-cost tradeoff, document frequency-dependent knowledge transfer absent from the RL-trained teacher, and characterize a calibration gap that SFT does not close - identifying both as directions for future work.
Comments15 pages, 4 figures, Github: https://github.com/joetheguide2/hindsight, Accepted at AACL-IJCNLP SRW 2026