人机协商作为交互式证明:无透明性下的可验证性条件
Human-LLM Deliberation as Interactive Proof: Conditions for Verifiability Without Transparency
浏览论文内容
中文总结 AI 辅助
本研究将人机协商建模为交互式证明,提出在无内部状态访问下,通过局部检查累积证据实现可验证性,并证明其可靠性条件与认证局限。
中文摘要 AI 辅助
当大型语言模型(LLM)提供一个用户无法轻易构建的论证时,用户如何决定是否接受其主张?受交互式证明的启发,我们将人机协商建模为具有无限制内部搜索的证明者与资源受限的人类验证者之间的交互。验证者请求并检查支持细节,而无法访问LLM的内部状态。通过的检查会累积证据,直至达到接受阈值。我们证明了针对自适应证明者的随时有效可靠性:只要任务提供在每次相关历史之后仍然有效的错误通过率和人类检查错误率的界限,那么最终接受错误主张的概率至多为所选错误水平。有限时域完备性界限额外要求诚实响应的充分性和足够的诊断进展的界限。进一步的检查可以加强接受的证据,但每次都需要另一个充分的响应和可靠的人类努力。这种权衡是否允许认证取决于验证者的努力预算、认知负荷、专业知识和疲劳程度。我们确定了在相同资源预算下,所提供的界限能够认证指定序列的局部检查但不能认证指定全局检查的条件。
英文摘要
When an LLM supplies an argument that a user could not readily construct, how can the user decide whether to accept its claim? Inspired by interactive proofs, we model human-LLM deliberation as an interaction between a prover with unrestricted internal search and a resource-bounded human verifier. The verifier requests and checks supporting details without access to the LLM's internal state. Passed checks accumulate evidence toward an acceptance threshold. We prove anytime-valid soundness against adaptive provers: the probability of ever accepting a false claim is at most a chosen error level, provided the task supplies bounds on false passes and human checking errors that remain valid after every relevant history. A finite-horizon completeness bound additionally requires bounds on the adequacy of honest responses and sufficient diagnostic progress. Further checks can strengthen the evidence for acceptance, but each requires another adequate response and reliable human effort. Whether this tradeoff permits certification depends on the verifier's effort budget, cognitive load, expertise, and fatigue. We identify conditions under which the supplied bounds certify a specified sequence of local checks but not a specified global check under the same resource budgets.
发表机构
- Stern School of Business(斯特恩商学院)
- New York University(纽约大学)
- Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。