证据先于文本:声明锁定式报告
Provenance Before Prose: Claim-Locked Reporting for Statistical Text Generation
查看机构详情
- Xidian University(西安电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对LLM生成统计报告时的数值漂移等问题,提出声明锁定式报告协议,在两类报告任务中提升了可复现性,还降低了token使用量与生成延迟。
中文摘要 AI 辅助
大型语言模型(LLM)能够流畅地将统计证据转化为自然语言,但统计报告仍可能出现数值漂移、效应方向颠倒,或将阈值化对比重述为分类效应。我们将此类失败视为一个控制问题:科学报告中承载证据的内容应由结构化统计结果确定,而非在文本生成过程中采样得到。因此,我们利用跨运行可复现性来测试报告可见的数值与声明是否在文本生成前被绑定。现有控制方法作用于文本或槽位层面;确定性混合模板在不同随机种子下仅能复现61.1%的报告可见数值内容,原因在于LLM仍会选择模板呈现的发现与数值。我们提出声明锁定式报告,这一「证据先于文本」协议会在LLM仅撰写连接性文本前,固定每项可报告声明的证据来源、数值、方向及允许的语言强度。在fMRI功能连接报告和基于Evidence Inference 2.0的随机对照试验报告中,声明锁定式报告较混合模板分别提升了37.4和20.5个百分点的可复现性。盲法人工审计支持观测到的方向保留与管控趋势;在与DeepSeek开展的fMRI成本分析中,声明锁定式报告还实现了观测到的最低token使用量和中位数生成延迟。
英文摘要
Large language models (LLMs) can fluently verbalize statistical evidence, yet statistical reports can still drift numerical values, invert effect directions, or restate thresholded contrasts as categorical effects. We frame these failures as a control problem: the evidence-bearing content of a scientific report should be fixed by structured statistical results rather than sampled during prose generation. We therefore use cross-run reproducibility to stress-test whether report-visible numbers and claims are bound before prose generation. Existing controls operate at the text or slot level; a deterministic hybrid template reproduces only 61.1% of report-visible numerical content across seeds because the LLM still selects which findings and numbers the template renders. We propose claim-locked reporting, a provenance-before-prose protocol that fixes the evidence source, numbers, direction, and allowed language strength of each reportable claim before the LLM writes only connective prose. Across fMRI functional-connectivity reporting and randomized controlled trial reporting on Evidence Inference 2.0, claim-locked reporting improves reproducibility over the hybrid template by 37.4 and 20.5 points, respectively. Blinded human audits support the observed direction-preservation and governance trends. In an fMRI cost analysis with DeepSeek, claim-locked reporting also yields the lowest observed token use and median generation latency.