针对特定语料库的临床检索增强生成(RAG)系统在HealthBench上的表现与较新的前沿大语言模型相当或更优
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
浏览论文内容
中文总结 AI 辅助
本文评估专为中低收入环境构建的临床RAG系统VITA,其在HealthBench基准测试中与前沿LLM表现相当或更优,语料库特异性可提升模型落地性但会降低沟通流畅度。
中文摘要 AI 辅助
近期有研究报告称,通用大语言模型(LLM)在医学基准测试中的表现可与专业临床AI工具相媲美甚至超越,但这类对比仅基于有限的系统集合,且所用基准测试大多是在高收入环境下开发的。本文对VITA进行评估,VITA是一款专为印度及其他中低收入(LMIC)环境下的上下文知识检索而构建的检索增强生成(RAG)系统。VITA从经整理的特定疾病指南、印度特定抗菌药物耐药性数据、国家处方集约束条件以及资源有限护理方案构成的语料库中进行检索;其架构和语料库为专有,但基准测试、医生编写的评分标准以及我们的完整响应和评分输出均公开,可供独立验证。在4023个英文HealthBench问题(占基准测试的80.5%)上,采用GPT-4.1作为评判者评分,VITA以51.9%的可能评分标准得分位居第一,领先于GPT-5.4(46.1%)、o4-mini(44.3%)、Gemini 3.1 Pro(42.6%)和Claude Sonnet 4.6(37.3%),且在45.4%的问题上得分最高。为测试对较新模型和评判者谱系的鲁棒性,我们对500个问题的子集重新使用当前代模型(GPT-5.5、Claude Opus 4.8、Gemini 3.5 Pro、Grok 4.3)运行,并由与所有测试系统无任何谱系关联的中立开源权重评判者(DeepSeek-V4-Pro)进行评分。在此情况下,差距缩小至相当水平:VITA和GPT-5.5在每问题平均得分上无统计学差异,而VITA在加权得分上领先且赢得了最多问题。在中立评判者下,VITA的准确性和完整性优势依然存在;但其沟通得分较低。这些结果表明,一款专为特定语料库构建的临床RAG系统在开放基准测试中仍能与前沿LLM竞争,这与语料库特异性作为一种设计变量的作用一致,该变量以牺牲沟通流畅度为代价提升了模型的落地性。
英文摘要
General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings. We evaluate VITA, a retrieval-augmented generation (RAG) system purpose-built for contextual knowledge retrieval in India and other low- and middle-income (LMIC) settings. VITA retrieves from a curated corpus of disease-specific guidelines, India-specific antimicrobial resistance data, national formulary constraints, and resource-limited care protocols; its architecture and corpus are proprietary, but the benchmark, the physician-written rubrics, and our full response and scoring outputs are public for independent verification. On 4,023 English-language HealthBench questions (80.5% of the benchmark), scored with a GPT-4.1 judge, VITA ranked first with 51.9% of possible rubric points, ahead of GPT-5.4 (46.1%), o4-mini (44.3%), Gemini 3.1 Pro (42.6%), and Claude Sonnet 4.6 (37.3%), and scored highest on 45.4% of questions. To test robustness to newer models and judge lineage, a 500-question subset was re-run against current-generation models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Pro, Grok 4.3) and graded by a neutral open-weight judge (DeepSeek-V4-Pro) sharing no lineage with any system tested. Here the gap narrowed to parity: VITA and GPT-5.5 were statistically indistinguishable on mean per-question score, while VITA led on points-weighted score and won the most questions. VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower. These results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.