arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DynamicDx:评估基于视频的诊断中的证据获取

DynamicDx: Evaluating Evidence Acquisition in Video-Based Diagnosis

Jiahui Li, Yutong Guo, Nan Yang, Wenzhan Song, Jin Lu, Fei Dou

arXiv 2609.32957首次发表:更新:

AI 中文总结

DynamicDx通过71次神经科会诊评估视频诊断中的证据获取,发现视频提升准确率9.9-22.5个百分点,但瓶颈在于证据获取,提供决定性检查可将准确率提升至73.2-93.0%,并通过视频描述器和文献检索干预改善。

AI 中文摘要

从视频中诊断患者不仅仅需要识别体征:视觉语言模型必须将所见转化为假设、问题和测试。DynamicDx在11个体征类别的71次神经科会诊中评估每一步,将真实患者视频与确诊诊断及基于相同病例报告构建的固定图表相连接,从而确保每个模型查询相同的证据。在五个此类模型中,视频相比盲输入将准确率提高了9.9至22.5个百分点,但仅靠识别或时间顺序无法解释这一提升:原因通常是,即使体征被识别,模型的仅视频鉴别诊断中仍缺少病因,且打乱帧顺序并未导致可靠的准确率下降。相反,轨迹回放将大部分增益追溯到视频提示的检查结果。证据获取是瓶颈:提供决定性的检查可将准确率提升至73.2%至93.0%。两种干预措施作用于该瓶颈。一个经过后训练的4B视频描述器改善了体征描述,尤其是从短且密集采样的片段中,而源清洁的文献检索扩展了初始假设;两者都使模型排序的测试更接近治疗临床医生记录的测试,并通过这些测试提高准确率。对于基于视频的诊断,更好的观察有助于引导更好的提问。

英文摘要

Diagnosing a patient from video requires more than recognizing the sign: a vision-language model must turn what it sees into hypotheses, questions and tests. DynamicDx evaluates each step in 71 neurological consultations across 11 sign categories, linking authentic patient videos to confirmed diagnoses and fixed charts built from the same case reports, so that every model queries the same evidence. Across five such models, video improves accuracy by 9.9-22.5 percentage points over blind input, but neither recognition alone nor temporal order explains the gain: the cause is usually missing from the model's video-only differential diagnosis even when the sign is recognized, and shuffling the frames produces no reliable accuracy loss. Instead, a trajectory replay traces most of the gain to the investigation results the video prompts. Evidence acquisition is the bottleneck: supplying the decisive investigations raises accuracy to 73.2-93.0%. Two interventions act on it. A post-trained 4B video describer improves sign descriptions, especially from a short, densely sampled segment, and source-clean literature retrieval expands initial hypotheses; both bring the tests a model orders closer to those the treating clinicians documented and, through them, raise accuracy. For video-based diagnosis, seeing better helps when it leads to asking better.

Comments49 pages. Code and data: https://github.com/jimmylihui/dynamicDx

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑