arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Cura 1T:用于智能医疗保健的专用模型

Cura 1T: Healthcare Foundation Model via Recursive Self-Improvement

Haolin Chen, Leon Qi, Steve Brown, Deon Metelski, Tao Xia, Joonyul Lee, Qixuan Wang, Kevin Riley, Frank Wang, Weiran Yao

arXiv 2607.15314首次发表:更新:

发表机构

actAVA AI(actAVA人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对医疗保健中多任务需求及能力失效问题,提出通过人工门控自我进化循环训练的Cura 1T模型,以数据为中心改进模型,该模型在医疗评估套件中表现优异,在相关基准测试中保持竞争力。

AI 中文摘要

医疗保健涉及高风险沟通、专家推理和工作流程执行,但涵盖这些用例的专用语言模型仍然有限。医疗模型必须处理患者咨询、文本和图像的临床推理、交互式诊断以及电子健康记录工具的使用。这些能力会以不同方式失效,针对一项任务的狭隘更新可能会降低另一项任务的性能。我们提出了Cura 1T,这是一个通过人工门控自我进化循环训练的医疗专用语言模型。在每个进化轮次中,训练智能体规划目标能力、训练模型、评估基准轨迹,并根据观察到的失败情况优化数据混合。这个以数据为中心的循环通过有针对性的合成和精选示例来改进模型,而不是单一的通用医疗数据更新。在整个医疗评估套件中,Cura 1T在前沿基线中排名靠前或接近顶部,同时在域外推理和智能基准测试中保持竞争力。

英文摘要

Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized language models that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use. These capabilities fail in different ways, and a narrow update for one task can degrade another. We present Cura 1T, a healthcare foundation model trained through recursive self-improvement (RSI). In each RSI round, the RSI harness runs the current model on healthcare benchmarks, evaluates the trajectories to locate capability gaps, and refines the training mixture by synthesizing training data. On 6 healthcare benchmarks, Cura 1T scores highest on MedAgentBench, HealthBench Professional, HealthBench Hard, MedXpertQA text, and AgentClinic, and second on MedXpertQA multimodal. It preserves performances on out-of-domain reasoning and agentic benchmarks including AIME, GPQA-Diamond, and $τ^2$-Bench.

CommentsModel: https://actava.ai/cura; Docs: https://actava.ai/cura/docs; Github: https://github.com/actava-ai/Cura

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑