arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MIRA-Ev:临床考试中颗粒证据检测和关系推理的基准

MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams

Iker De la Iglesia, Johanna Ramirez-Romero, Jose Maria Villa-Gonzalez, Irune Urroz García, Ander Barrena, Aitziber Atutxa

arXiv 2607.19201首次发表:更新:

发表机构

HiTZ Center, University of the Basque Country (EHU); Hospital Universitario de Cruces(巴斯克大学HiTZ中心; 克鲁塞斯大学医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对临床自然语言处理评估局限,引入基于西班牙MIR执照考试案例的MIRA-Ev基准,重新注释相关关系并发布多语言版本,将评估组织为证据句子检索、论证组件提取和关系分类的三层任务层次结构。

AI 中文摘要

临床自然语言处理评估仍以多项选择题回答(MCQA)为主,其仅对最终答案准确性打分,无法检测模型在基于不相关、缺失或矛盾证据得出正确诊断时的情况。我们引入了MIRA-Ev,这是一个基于西班牙内科住院医师(MIR)执照考试案例构建的临床论证挖掘基准,由专家临床医生重新注释了跨度级别的前提、主张和定向支持/攻击关系,并以西班牙语(母语)、英语和巴斯克语平行版本发布,这是首个巴斯克语临床论证资源。MIRA-Ev将评估组织成一个三层任务层次结构:证据句子检索、论证组件提取和关系分类。

英文摘要

Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct diagnosis while grounding it in irrelevant, absent, or contradictory evidence. We introduce MIRA-Ev, a clinical argument mining benchmark built on Spanish Médico Interno Residente (MIR) licensing-exam cases, re-annotated by expert clinicians with span-level premises, claims, and directed support/attack relations, and released in parallel Spanish (native), English, and Basque versions, the first clinical argumentation resource in Basque. MIRA-Ev organizes evaluation into a three-tier task hierarchy: evidence sentence retrieval, argumentative component extraction, and relation classification.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑