AI 中文总结
本研究针对德国法定医疗保险网站,提出一种结合多步骤的大语言模型辅助审核优先级排序工作流程,经实验验证可有效分配审核工作量,区分AI来源信号与质量声明。
AI 中文摘要
背景:德国法定医疗保险(SHI)基金发布的网络内容超出了持续专业审核的能力,其内容会影响公众对健康和福利的预期。通用AI文本检测无法识别医疗、福利、法律或编辑方面的审核需求。目的:明确一种多阶段工作流程,该流程可对实质性审核需求进行优先级排序,同时将AI来源信号与质量声明区分开来。方法:我们分析了来自84个SHI网站或子网站的56198个页面。该工作流程结合了确定性筛选、模型辅助分类与深度审核、最低证据检查、时间有效性保障以及配对模型比较。该工作流程受可复现性限制,并非经过验证的检测器。生产代码为专有,可复现性依赖于冻结的衍生表和配对比较产物。300页的低优先级检查是单模型、风险富集的路由压力测试,而非人工参考评估。结果:所有页面均获得审核状态。该工作流程生成了35998条审核记录,并将21452条记录分配至案例审核。工作量集中在透明度、法律框架、医疗内容、矛盾之处以及AI相关故障模式信号上。在捕获的页面文本中,有31347条记录可定位到引用段落,确认了字面存在而非事实正确性。路由压力测试在300页样本中的100页(占比33.3%)上检测到信号。在182个匹配案例中,两个模型的一致性为75.8%(kappa=0.532;95%置信区间0.415-0.649)。结论:该工作流程生成了优先级排序后的工作量,而非错误发生率或最终的法律、医疗或保险公司层面的结论。它既不能证明AI的作者身份,也不能验证自主检测。配对模型一致性衡量的是一致性,而非正确性或足够的分类性能;公开声明需要人工裁决。
英文摘要
Background: German statutory health insurance (SHI) funds publish web portfolios that exceed continuous specialist review capacity. Their content can shape health and benefit expectations. Generic AI-text detection does not identify medical, benefit, legal, or editorial review needs. Objective: To characterize a multi-stage workflow that prioritizes substantive review needs while separating AI-provenance signals from quality claims. Methods: We analyzed 56,198 pages from 84 SHI websites or sub-sites. The workflow combined deterministic screening, model-assisted triage and in-depth review, minimum evidence checks, temporal-validity safeguards, and paired-model comparison. It is reproducibility-bounded, not a validated detector. Production code is proprietary; reproducibility rests on frozen derived tables and paired-comparison artifacts. The 300-page lower-priority check was a single-model, risk-enriched routing stress test, not a human-reference evaluation. Results: All pages received a review state. The workflow generated 35,998 review records and routed 21,452 to case review. The workload concentrated in transparency, legal framing, medical content, contradictions, and AI-related failure-mode signals. A quoted passage was locatable in captured page text for 31,347 records, confirming literal occurrence rather than factual correctness. The routing stress test surfaced a signal on 100/300 pages (33.3% within the sample). Across 182 matched cases, two models agreed in 75.8% (kappa = 0.532; 95% CI 0.415-0.649). Conclusions: The workflow produces a prioritized workload, not error prevalence or final legal, medical, or insurer-level findings. It neither proves AI authorship nor validates autonomous detection. Paired-model agreement quantifies consistency, not correctness or sufficient triage performance; public claims require human adjudication.
Comments31 pages, 5 figures. Also archived at Zenodo: https://doi.org/10.5281/zenodo.21738595