arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ClinAgent:基于ReAct的临床试验信息对话式访问智能体

ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information

Antonino Vaccarella, Riccardo Cantini, Domenico Talia, Paolo Trunfio, Marianna Talia, Rosamaria Lappano, Marcello Maggiolini

arXiv 2609.13860首次发表:更新:

发表机构

University of Pisa; University of Calabria; National Research Council(比萨大学; 卡拉布里亚大学; 国家研究委员会)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ClinAgent是一个基于ReAct的对话式智能体,利用检索增强生成和多种工具,实现临床试验信息的自然语言查询与综合,实验表明Gemini性能最佳,DeepSeek思考模式规划质量最优。

AI 中文摘要

查询临床试验注册库仍然是一个手动且容易出错的过程,要求研究人员在没有自然语言交互或跨来源综合支持的情况下,浏览大量半结构化数据。为了解决这个问题,我们引入了ClinAgent,一个基于智能体检索增强生成(RAG)的对话系统,使临床医生和研究人员能够用自然语言查询临床试验信息,并在多轮交互中获得有依据的、最新的回答。该系统以一个遵循ReAct范式的大型语言模型(LLM)智能体为核心,该智能体迭代地对查询进行推理,在一组集成的工具中进行选择,并根据中间输出优化其行动。这些工具包括一个this http URL搜索界面、一个PubMed模块,以及一个基于本地缓存的临床试验结构化数据集运行的Python分析器。我们使用一个三阶段框架评估该系统,该框架评估操作有效性、规划质量、工具使用效率和专家定性判断,并比较了三个LLM后端:Gemini 3.0 Flash和DeepSeek V3.2的两个变体(思考模式和非思考模式)。结果显示了互补的优势,DeepSeek(思考模式)在规划质量方面表现出色,而Gemini实现了最高的整体性能和最强的专家评分。总的来说,我们的发现强调了智能体AI系统在提高临床试验信息的可访问性和综合方面的潜力,支持更高效和以用户为中心的生物医学研究工作流程。

英文摘要

Querying clinical trial registries remains a manual and error-prone process, requiring researchers to navigate large volumes of semi-structured data without support for natural language interaction or cross-source synthesis. To address this, we introduce ClinAgent, a conversational system based on agentic Retrieval-Augmented Generation (RAG) that enables clinicians and researchers to query clinical trial information in plain language and receive grounded, up-to-date responses across multi-turn interactions. The system centers on a Large Language Model (LLM) agent following the ReAct paradigm, which iteratively reasons over queries, selects among a set of integrated tools, and refines its actions based on intermediate outputs. These tools include a ClinicalTrials.gov search interface, a PubMed module, and a Python-based analyzer operating on a locally cached structured dataset of clinical trials. We evaluate the system using a three-phase framework assessing operational effectiveness, planning quality, tool-use efficiency, and expert qualitative judgments, comparing three LLM backends: Gemini 3.0 Flash and two variants of DeepSeek V3.2 (thinking and non-thinking). Results reveal complementary strengths, with DeepSeek (thinking mode) excelling in planning quality, while Gemini achieves the highest overall performance and strongest expert ratings. Overall, our findings highlight the potential of agentic AI systems to improve the accessibility and synthesis of clinical trial information, supporting more efficient and user-centered biomedical research workflows.

CommentsAccepted at the CIBB 2026 conference (https://cibb2026.teralab.ai/)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑