人工智能代理起草转化影响摘要的实际评估
Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries
浏览论文内容
中文总结 AI 辅助
研究针对CTSA项目人工整理学者研究影响耗时久且难扩展的问题,构建人工参与的AI代理。该代理为学者整理证据档案、起草影响摘要,经评估,在CTSA中心工作流程中表现良好,能提高效率、实现队列规模影响报告。
中文摘要 AI 辅助
引言。临床与转化科学奖(CTSA)项目必须记录学者的研究影响,但人工整理每位学者的记录估计需工作人员15小时,且无法扩展到整个队列。人工智能代理可作为跨平台和学科收集学者数据的工具。方法。我们构建了一个人工参与的人工智能代理,为每位学者整理来源证据档案,并起草单句转化科学效益模型(TSBM)影响摘要供工作人员审核。我们在一个CTSA中心的影响报告工作流程中对10名职业发展(KL2/K12)学者进行了评估。两名评估人员将所有507项结果独立编码为接受、编辑或拒绝;主要衡量指标是一致可用率,即被接受或编辑的比例。结果。两位评审人员接受或编辑了代理结果的81.7%。评审人员每位学者平均花费14分钟,取代了估计15小时的人工整理。评分者间一致性中等(关于可用与拒绝决定的科恩kappa系数为0.43)。一项概况发现研究发现代理的召回率接近人工搜索。代理的影响证据涵盖所有四个TSBM领域,约三分之一的评审结果属于常规流程易遗漏的非学术类别。评审人员在5分制上对综合准确性评分为4.5,对有用性评分为4.8。结论。人工参与的人工智能代理可作为学者影响记录的初稿作者,将工作人员从收集和撰写工作转变为审核工作,使队列规模的影响报告成为可能。
英文摘要
Introduction. Clinical and Translational Science Award (CTSA) programs must document their scholars' research impact, but assembling each scholar's record by hand takes staff an estimated 15 hours and does not scale to a full cohort. An artificial intelligence (AI) agent could serve as a tool to gather scholar data across platforms and disciplines. Methods. We built a human-in-the-loop AI agent that assembles a dossier of sourced evidence for each scholar and drafts one-sentence Translational Science Benefits Model (TSBM) impact summaries for staff review. We evaluated it in the impact-reporting workflow of one CTSA hub across 10 career-development (KL2/K12) scholars. Two evaluation staff independently coded all 507 findings as accept, edit, or reject; the primary measure was the unanimous usable rate, defined as the share both accepted or edited. Results. Both reviewers accepted or edited 81.7% of the agent's findings. Reviewers each spent a median of 14 minutes per scholar, replacing an estimated 15 hours of manual assembly. Inter-rater agreement was moderate (Cohen's kappa 0.43 on the usable-versus-reject decision). A profile discovery study found the agent's recall close to human search. The agent's impact evidence spanned all four TSBM domains, and about a third of the reviewed findings fell in non-scholarly categories that routine processes tend to miss. Reviewers rated synthesis accuracy 4.5 and usefulness 4.8 on a 5-point scale. Conclusions. A human-in-the-loop AI agent can serve as the first-pass author of a scholar's impact record, shifting staff from collecting and writing to reviewing, and making cohort-scale impact reporting feasible.
发表机构
- University of Illinois Chicago(伊利诺伊大学芝加哥分校)
机构由 AI 辅助整理,请以论文原文为准。