arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11185cs.AI

大型语言模型能否遵循医学专家逻辑?风险偏倚评估中层级逻辑一致性基准

Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment

  • Beijing University of Posts and Telecommunications(北京邮电大学)
  • Hypertension Center, Beijing Anzhen Hospital, Capital Medical University(首都医科大学附属北京安贞医院高血压中心)

机构由 AI 辅助整理,请以论文原文为准。

Jiayu Huang, Zichen Tang, Qianhui Ling, Zemin Kuang, Haihong E

AI总结:

该研究提出LogiMed-RoB基准,基于Cochrane RoB 2.0逻辑评估LLMs的层级逻辑一致性,发现高原子一致性下存在错误复合效应和证据推理差距,强调临床部署需白盒验证。

AI中文摘要:

循证医学要求严格的逻辑一致性,然而当前对大型语言模型(LLMs)的评估侧重于表面的标签匹配,而非真正的推理。我们提出了LogiMed-RoB,一个基于Cochrane风险偏倚(RoB)2.0专家逻辑的基准,包含860项随机对照试验(RCTs)和14,820个查询。该基准在层级逻辑一致性(HLC)框架下从四个维度评估模型:原子一致性、领域一致性、聚合一致性和证据忠实性。对10个最先进的LLMs进行的实验揭示了一种灾难性的错误复合效应:尽管顶尖模型的原子一致性达到98.88%,其端到端一致性却骤降至45.13%,且多个开放权重架构的一致性几乎降至0%。我们进一步发现了系统性的证据-推理差距:即使模型检索到高质量证据,在18.63%至40.05%的案例中仍无法推导出正确结果,而盲目猜测率高达48.28%。LogiMed-RoB表明,高结果准确性可能掩盖关键的推理缺陷,这强调了在临床部署中进行白盒逻辑验证的必要性。

英文摘要:

Evidence-based medicine demands strict logical consistency, yet current evaluations of large language models (LLMs) prioritize superficial label matching over genuine reasoning. We introduce LogiMed-RoB, a benchmark grounded in Cochrane Risk of Bias (RoB) 2.0 expert logic, comprising 860 randomized controlled trials (RCTs) and 14,820 queries. It evaluates models under the Hierarchical Logical Consistency (HLC) framework across four dimensions: Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness. Experiments on 10 state-of-the-art LLMs reveal a catastrophic Error Compounding Effect: despite the top model reaching 98.88% Atomic Consistency, its end-to-end consistency collapses to 45.13%, with several open-weight architectures plummeting to nearly 0%. We further uncover a systematic evidence-reasoning gap: even when models retrieve high-quality evidence, they fail to deduce correct outcomes in 18.63-40.05% of cases, while Blind Guess Rates reach 48.28%. LogiMed-RoB demonstrates that high outcome accuracy can conceal critical reasoning flaws, underscoring the necessity of white-box logical verification for clinical deployment.

补充信息

↑