arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10601cs.CR

阈值选择而非样本量决定非确定性复合AI工作流的无信任验证边界

Threshold Choice, Not Sample Size, Bounds Trustless Verification of Nondeterministic Compound AI Workflows

Alper Alimoglu

首次发表
浏览论文内容

中文总结 AI 辅助

针对非确定性复合AI流水线的无信任验证,提出协议并证明瓶颈是阈值选择而非样本量,按执行推导的阈值优于固定阈值。

中文摘要 AI 辅助

复合AI流水线将大语言模型调用、检索器和工具串联起来,并且是非确定性的:采样、模型更新和易变的工具响应使得同一输入在不同运行中产生不同输出。此类流水线日益在边缘、云和轨道节点上运行,这些节点不属于任何单一主体,其优化会在检查前丢弃中间结果。在这些环境中验证复现意味着要容忍非确定性输出、可能不诚实报告的节点,以及对任何共享记录的间歇性访问;现有工作最多同时处理其中两个问题。我们提出一个覆盖全部三个问题的协议:它在策略摘要下提交每个阶段输入、输出和上下文的摘要,该策略摘要固定了度量和阈值;它不信任执行节点而进行锚定;在分区情况下延迟处理;并基于$k$次重新执行的中位数(无需法定人数)裁决挑战。该过程失效之处正是主要结果。在合成的HotpotQA流水线上,校准的固定阈值接受44/45次诚实复现,拒绝104/105个发散对,但在k=5时让27/29次同输入伪造通过,更多样本无济于事,因为采样只是锐化估计而不移动其位置。在固定该度量和此流水线的重新执行离散度的情况下,约束因素是阈值而非样本量:一个按执行推导的阈值检测到29次中的19次,而匹配到相同零诚实拒绝的最佳常数阈值达到9次;在k=5时不拒绝任何诚实承诺,尽管在k=3时15次中有3次;并捕获了针对它构建的攻击者的15次中的11次,该攻击者必须瞄准在其承诺存在后才绘制的目标。该规则是测量性的而非部署性的。

英文摘要

Compound AI pipelines chain LLM calls, retrievers, and tools and are nondeterministic: sampling, model updates, and volatile tool responses make one input yield different outputs across runs. Such pipelines increasingly run across edge, cloud, and orbital nodes owned by no single party, whose optimizations discard intermediate results before inspection. Verifying reproduction there means tolerating nondeterministic outputs, a node that may not report honestly, and intermittent access to any shared record; existing work addresses at most two at once. We give a protocol covering all three: it commits digests of each stage's inputs, outputs, and context under a policy digest pinning the metric and threshold, anchors them without trusting the executing node, defers under partition, and decides a challenge on the median of $k$ re-executions, with no quorum. Where that procedure breaks is the main result. On a synthetic HotpotQA pipeline a calibrated fixed threshold accepts 44 of 45 honest reproductions and rejects 104 of 105 divergent pairs, yet lets same-input fabrication through in 27 of 29 trials at k=5, more samples being no help since sampling sharpens an estimate without moving it. Holding that metric and this pipeline's re-execution spread fixed, the binding constraint is the threshold rather than the sample size: one derived per execution detects 19 of 29 where the best constant matched to the same zero honest rejections reaches 9, rejects no honest commitment at k=5 though 3 of 15 at k=3, and catches 11 of 15 of an attacker built against it, which has to aim at a target drawn only after its commitment exists. That rule is measured rather than deployed.

↑