arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28624cs.CLcs.AI

MA-RAG:用于帕金森病纵向评估查询驱动式摘要的多智能体检索增强生成框架

MA-RAG: Multi-Agent Retrieval-Augmented Generation for Query-Driven Summarization of Longitudinal Parkinson's Disease Assessments

Sana Alamgeera, Denise Goberta, Muhammad Irshad, Anne H. H. Ngu

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLMs解读帕金森病纵向临床评估时的事实与时间一致性问题,提出MA-RAG多智能体检索增强生成框架,经评估其事实精度大幅提升、幻觉率显著降低,获临床专家高度认可。

中文摘要 AI 辅助

准确解读帕金森病的单次就诊及纵向临床评估耗时且依赖专业医师经验。尽管大型语言模型(LLMs)可生成自然语言摘要,但它们常缺乏特定领域的临床依据,且难以针对结构化纵向评估数据生成事实准确、时间一致的响应。为解决这些局限,我们提出MA-RAG,这是一种查询驱动的多智能体检索增强生成框架,它将临床推理分解为领域专用智能体,结合结构化事实提取,并通过最终验证阶段生成符合临床依据的摘要。该框架支持四项临床分析任务:单次就诊、轨迹、对比及队列摘要。我们采用客观指标(事实精度、幻觉率、时间保真度、语义相似度)及临床专家的主观评估来评估MA-RAG。与传统方法、仅RAG方法、单智能体RAG基线相比,MA-RAG大幅提升了事实正确性,事实精度相对提升最高达122%(从0.436升至0.990),幻觉率降低最高达98%(从0.564降至0.010),同时在条理性和临床实用性方面始终获得临床专家的最高评分。这些结果表明,领域专用多智能体推理可实现对结构化纵向临床评估数据的可靠查询驱动式摘要。

英文摘要

Accurate interpretation of single-visit and longitudinal clinical assessments for Parkinson's disease is time-consuming and often depends on specialist expertise. Although large language models (LLMs) can generate natural language summaries, they frequently lack domain-specific clinical grounding and struggle to produce factually correct and temporally consistent responses for structured longitudinal assessment data. To address these limitations, we propose MA-RAG, a query-driven multi-agent retrieval-augmented generation framework that decomposes clinical reasoning into domain-specialized agents, combines structured fact extraction, and synthesizes clinically grounded summaries through a final verification stage. The framework supports four clinical analysis tasks: single-session, trajectory, comparison, and cohort summarization. We evaluate MA-RAG using objective metrics, namely Fact Precision, Hallucination Rate, Temporal Fidelity, and Semantic Similarity, together with subjective evaluations conducted by clinical experts. Compared to Traditional, RAG-only, and Single-agent RAG baselines, MA-RAG substantially improves factual correctness, achieving up to a 122% relative increase in Fact Precision (from 0.436 to 0.990) and reducing the Hallucination Rate by up to 98% (from 0.564 to 0.010), while consistently receiving top ratings from clinical experts for organization and clinical usefulness. These results demonstrate that domain-specialized multi-agent reasoning enables reliable query-driven summarization of structured longitudinal clinical assessment data.

发表机构

  • Texas State University(德克萨斯州立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑