发表机构
DaNaA Data Co., Ltd.(DaNaA数据有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一个多智能体大语言模型框架,用于个性化健康检查结果解读,通过并行执行多个任务特定智能体,在复合查询上显著提升回答质量,但增加延迟和成本。
AI 中文摘要
个性化解读健康检查结果需要跨纵向记录、医学知识、生活方式指导和医疗导航进行推理。我们提出了一个多智能体大语言模型(LLM)系统,该系统可识别多个意图,将每个意图映射到特定任务的智能体,并行执行这些智能体,并综合其输出。我们在120个韩语复合查询(包含两到四个需求)上,使用合成健康检查记录,比较了单智能体和多智能体设置下生成的答案。多智能体将加权LLM评判分数从1.695提高到1.797(p = 0.027),另外三个LLM评判员显示出一致的改进(Δ = +0.111至+0.186,所有p < 0.05)。改进来自有用性、一致性以及对复合查询中每个需求的处理,而数值准确性和基础性仅在四个评判员之一下显著改善,医疗安全性没有差异,严重失败的发生率相似(单智能体15.0% vs. 多智能体13.3%)。两位人类评估者在66.7%和68.3%的成对比较中偏好多智能体。多智能体执行将延迟和成本分别增加了1.31倍和2.02倍。在探索性亚组分析中,改进集中在涉及个人记录查询的查询中。
英文摘要
Personalized interpretation of health checkup results requires reasoning across longitudinal records, medical knowledge, lifestyle guidance, and healthcare navigation. We present a multi-agent large language model (LLM) system that identifies multiple intents, maps each to a task-specific agent, executes them in parallel, and synthesizes their outputs. We compared answers generated in Single Agent and Multi Agent settings on 120 Korean compound queries combining two to four requirements, using synthetic health checkup records. The Multi Agent improved the weighted LLM-judge score from 1.695 to 1.797 (p = 0.027), and three additional LLM judges showed consistent improvements ($Δ$ = +0.111 to +0.186, all p < 0.05). The gains came from usefulness, consistency, and the handling of every requirement in compound queries, whereas numerical accuracy and grounding improved significantly under only one of the four judges and medical safety did not differ, and critical failures occurred at similar rates (Single Agent 15.0% vs. Multi Agent 13.3%). Two human evaluators preferred Multi Agent in 66.7% and 68.3% of pairwise comparisons. Multi Agent execution increased latency and cost by 1.31$\times$ and 2.02$\times$, respectively. In exploratory subgroup analyses, the improvement was concentrated in queries involving personal-record lookup.
Comments16 pages, 2 figures, 8 tables and Appendix