LLM答案发布中的审批完整性与恢复
Approval Integrity and Recovery in LLM Answer Publication
- Bahçeşehir University(巴赫切谢希尔大学)
- University of Turkish Aeronautical Association(土耳其航空协会大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过大规模人工标注实验,系统评估了LLM答案发布机制中的审批完整性与恢复策略,揭示了语义性错误批准、授权过期及恢复引发错误之间的差异。
AI中文摘要:
LLM系统中的发布完整性要求将已批准的内容绑定到其当前的授权上下文。我们考察了Lightcap发布强制机制中的精确内容绑定、授权新鲜度和检查点恢复。基于来自150个源任务的900条独立人工标注的RAGTruth响应,三个不同日期的Ministral模型和一个同模型直接接地基线产生了3,600项评估。使用14B实例化的生产响应-动作检查器接受了302个未支持标签答案中的291个;直接基线接受了41个。支持答案保留率分别为95.2%和66.9%。一个精确的提升-修正恒等式追踪了100条按时间顺序的3B-14B-8B-14B答案轨迹中的错误。在65个最初批准的答案中,最终的有状态重新检查-恢复策略相对于初始检查点将精确匹配错误增加了9.23个百分点(95%文章聚类区间[-1.72, 19.61])。受控的证据指纹变化暴露了发布与恢复之间不对称的新鲜度强制。一个单独的BIPIA提示注入实验在266个有效编辑器输出中记录了零次目标插入。外部Hugging Face校准实验将检索模型从ArguAna迁移到SciFact和NFCorpus,并将诊断决策规则从Thunderbird迁移到BGL,区分了概率校准与排序变化。这些测量在可执行的发布边界上区分了语义性错误批准、过期授权和恢复引起的错误。
英文摘要:
Publication integrity in LLM systems requires binding approved content to its current authorization context. We examine exact-content binding, authorization freshness and checkpoint recovery in Lightcap's publication enforcement mechanism. On 900 independently human-annotated RAGTruth responses from 150 source tasks, three dated Ministral models and a same-model direct-grounding baseline yield 3,600 assessments. The production response-act checker instantiated with 14B accepts 291 of 302 unsupported-labelled answers; the direct baseline accepts 41. Supported-answer retention is 95.2% and 66.9%, respectively. An exact promotion-correction identity tracks error through 100 chronological 3B-14B-8B-14B answer trajectories. Among 65 initially approved answers, the final stateful recheck-recovery policy increases exact-match error by 9.23 percentage points relative to the initial checkpoint (95% article-clustered interval [-1.72, 19.61]). Controlled evidence-fingerprint changes expose asymmetric freshness enforcement between publication and recovery. A separate BIPIA prompt-injection experiment records zero target insertions among 266 valid editor outputs. External Hugging Face calibration experiments transfer retrieval models from ArguAna to SciFact and NFCorpus, and diagnostic decision rules from Thunderbird to BGL, distinguishing probability calibration from ranking changes. The measurements separate semantic false approval, stale authorization and recovery-induced error at executable publication boundaries.