arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22555cs.AI

深度镜头诊断代理:代理工作流程设计使小型推理模型能够与前沿语言模型竞争

DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs

发表机构约翰·斯诺实验室
查看机构详情
  • John Snow Labs(约翰·斯诺实验室)

机构由 AI 辅助整理,请以论文原文为准。

Mahmood Bayeshi, Veysel Kocaman, Muhammed Ali Naqvi, Yigit Gul, David Talby

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对医疗诊断中前沿语言模型单步提示诊断推理脆弱的问题,提出以小型医疗推理模型和检索增强生成的深度镜头诊断代理五阶段利用管道,提升了诊断准确率,降低成本,还产生结构化工件支持高风险设置。

中文摘要 AI 辅助

医疗诊断是一个多阶段过程,前沿语言模型虽强大,但单步提示诊断推理脆弱。我们提出深度镜头诊断代理,这是一个以小型医疗推理模型和检索增强生成(RAG)为中心的五阶段利用管道。该管道执行结构化临床提取、规范检索、约束候选生成、明确证据三角测量和可审计的最终决策。在915例诊断竞技场基准测试中,该代理实现了60.14%的 top-1诊断准确率。同一模型在无代理工作流程时准确率为23.99%,工作流程设计使其提升36个百分点。该代理每例成本0.0072美元,延迟24秒,比Claude Sonnet 4.5和Gemini 3.1 Pro更便宜且性能更优。此外,管道产生结构化中间工件,支持高风险设置中的可追溯性、可重复性和可审计证据。

英文摘要

Medical diagnosis is a multi-stage process: extract facts, consult knowledge, generate a differential analysis, and select the best diagnosis with explanations. Frontier LLMs are strong generalists, but single-shot prompting often yields brittle diagnostic reasoning. We present the DeepLens Diagnosis Agent, a five-stage harnessing pipeline (combining model capabilities with disciplined process constraints) centered on a small medical reasoning model (JSL Medical Small 7B v2) and retrieval-augmented generation (RAG). The pipeline enforces structured clinical extraction, disciplined retrieval, constrained candidate generation, explicit evidence triangulation, and an auditable final decision. On the 915-case DiagnosisArena benchmark, the agent achieved 60.14% top-1 diagnostic accuracy, the highest among small and medium-sized models. The same model without the agent workflow achieved 23.99%, a +36-point gain from workflow design alone, despite 88.2% on standard medical benchmarks, showing that diagnostic reasoning under uncertainty requires more than knowledge recall. The agent costs USD 0.0072 per case (24K tokens on A100) with 24-second latency, 35-45% cheaper than Claude Sonnet 4.5 (USD 0.0110) and Gemini 3.1 Pro (USD 0.0128) while outperforming them by +9.70pp and +9.17pp. Harnessing can also correct frontier model failures; workflow constraints can outweigh parameter count or API cost. Beyond aggregate accuracy, the pipeline produces structured intermediate artifacts that make each stage inspectable and support error localization. These properties support high-stakes settings where traceability, reproducibility, and auditable evidence matter alongside benchmark performance.

补充信息

↑