AI 中文总结
该研究提出PROVE多智能体框架,结合LLM与程序化验证器,在临床试验报告TFL验证中,LLM辅助语义匹配可显著提升差异检测的召回率与F1值,兼顾解读灵活性与数值准确性。
AI 中文摘要
确保临床试验表格、图形和清单(TFL)的准确性与一致性仍是监管报告中的重大挑战。独立编程与人工审核是必要的质量控制措施,但跨输出验证仍高度依赖审核人员检查,可能遗漏结构性、逻辑性或算术差异。大语言模型(LLM)可辅助解读多样的表格语言并浏览冗长研究文档,但无法替代程序化统计检查。我们推出PROVE(Programmatic Reporting and Output Verification Engine,程序化报告与输出验证引擎),这一可审计框架可选择使用LLM与检索支持进行表格解读,同时将数值与逻辑决策交由程序化验证器处理。PROVE将结果与源证据关联,支持跨输出一致性检查,并可根据研究需求启用或禁用LLM。我们使用10个通过原始数据经SDTM、ADaM及TFL输出生成的重复合成肿瘤学报告包评估PROVE,每个包包含配对的干净版本与注入差异的版本,每个重复包含15个随机注入的差异。我们测试两种表格标签设置:与验证器词汇匹配的精确标签,以及临床含义相似但措辞不同的标签。在已实施的规则类别中,所有自动化PROVE变体在精确标签设置下均实现完美分类;在标签变化设置下,与精确匹配、模糊词汇及嵌入相似度变体相比,LLM辅助的语义匹配将整体召回率从0.588提升至0.993,整体F1值从0.735提升至0.996。这些结果表明,LLM最适用于解读现实世界中TFL措辞与格式的变化,而最终数值验证应由可执行检查负责。
英文摘要
Ensuring the accuracy and consistency of clinical trial Tables, Figures, and Listings (TFLs) remains a major challenge in regulatory reporting. Independent programming and manual review are essential quality-control practices, but cross-output verification still depends heavily on reviewer inspection and may miss structural, logical, or arithmetic discrepancies. Large language models (LLMs) can help interpret varied table language and navigate lengthy study documents, but they are not reliable substitutes for programmed statistical checks. We introduce PROVE (Programmatic Reporting and Output Verification Engine), an auditable framework that uses optional LLM and retrieval support for table interpretation while reserving numerical and logical decisions for programmed validators. PROVE links findings to source evidence, supports cross-output consistency checks, and allows LLM use to be enabled or disabled based on study requirements. We evaluated PROVE using ten replicated synthetic oncology reporting packages generated from raw data through SDTM, ADaM, and TFL outputs, with paired clean and discrepancy-injected packages; each replicate included 15 randomly injected discrepancies. We examined two table-label settings: exact labels matching the validator vocabulary and labels with similar clinical meaning but different wording. Within the implemented rule classes, all automated PROVE variants achieved perfect classification in the exact-label setting. In the label-variation setting, LLM-assisted semantic matching improved overall recall from 0.588 to 0.993 and overall F1 from 0.735 to 0.996 compared with exact-match, fuzzy lexical, and embedding-similarity variants. These findings suggest that LLMs are most useful for interpreting real-world variation in TFL wording and formatting, while executable checks should remain responsible for final numerical validation.