arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ARCagent:面向临床问答的自适应检索校准智能体

ARCagent: An Adaptive Retrieval Calibration Agent for Clinical Question Answering

Yuyan Chen

arXiv 2609.36392首次发表:更新:

AI 中文总结

ARCagent是一个针对ME/CFS临床问答的自适应检索校准智能体,通过构建含冲突登记表的知识库、冲突感知检索校准流程及LLM评判基准,达到95.3%性能,优于基础大语言模型。

AI 中文摘要

在临床指南不完整、存在争议或相互矛盾的疾病中,知识完整性和动态冲突感知综合是标准检索增强生成系统所不具备的两个安全关键属性。因此,我们提出了\sysname,一个针对ME/CFS的自适应检索校准临床问答智能体,该疾病的诊断框架并存,且主要指南在治疗方面积极相互矛盾。ARCagent贡献了三个组成部分。首先,一个包含1,706个块、10个来源的知识库,具有结构化的指南间冲突登记表,覆盖所有活跃的ME/CFS诊断框架。其次,一个冲突感知的检索校准流程,利用查询特定的焦点和冲突信号对检索到的证据进行重新排序。第三,一个由LLM作为评判者评分的基准,避免了关键词匹配导致的平均10.1个百分点的系统性低估。ARCagent达到了95.3%的性能,优于所有基础大语言模型。代码可在该https URL获取。

英文摘要

In diseases where clinical guidelines are incomplete, contested, or mutually contradictory, knowledge completeness and dynamic conflict-aware synthesis are two safety-critical properties that standard Retrieval-Augmented Generation systems do not provide. Therefore, we present \sysname, an adaptive retrieval calibration clinical question-answering agent for ME/CFS, a disease where diagnostic frameworks coexist and major guidelines actively contradict each other on treatment. ARCagent contributes three components. First, a 1,706-chunk, 10-source knowledge base with a structured inter-guideline conflict registry spanning all active ME/CFS diagnostic frameworks. Second, a conflict-aware retrieval calibration pipeline that re-ranks retrieved evidence using query-specific focus and conflict signals. Third, a benchmark scored by LLM-as-Judge, avoiding systematic underestimation averaging 10.1 percentage points caused by keyword matching. ARCagent achieves 95.3%, outperforming all base LLMs. Code is available at https://github.com/Yukyin/ARCagent.

Comments13 pages, 7 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑