arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.08961cs.LGcs.AImath.STstat.TH

NL-PAC:大语言模型介导监督中的规范模糊性与认证极小极大风险下限

NL-PAC: Specification Ambiguity and Certified Minimax Risk Floors in LLM-Mediated Supervision

发表机构奥泽金大学
查看机构详情
  • Özyeğin University(奥泽金大学)

机构由 AI 辅助整理,请以论文原文为准。

Berkay Anahtarci

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型介导监督中规范模糊性问题,引入NL-PAC框架,利用固定模型阈值解码法则定义相关内容,通过有限样本置信界认证风险,在Qwen模型审计中得出特定提示有正证书等结果,该保证有特定局限性。

中文摘要 AI 辅助

大语言模型越来越多地为自然语言指定的任务提供标签、评估和反馈。当规范存在多种解读但监督渠道未揭示实际采用哪种解读时,额外标签可减少采样误差但无法解决识别问题。我们引入自然语言PAC(NL-PAC)框架,它使用固定模型的阈值解码法则定义可接受标签和候选目标。多个标签可接受的概率等于逐点可接受目标类的直径,在目标盲监督下,每个学习者在每个样本量下都会面临至少此直径一半的最坏情况风险;通过与数据无关的策略可实现该类上的确切随机极小极大风险。有限样本置信界使这些量可从留出的未标记输入中得到认证。在对Qwen~2.5 - 3B的审计中,一个预先指定的提示产生了正的模型相对证书,而释义和精确规则控制产生零。留出的桥梁审计发现,提供的候选解读子句未通过将证书转移到连贯解读所需的可接受性条件。该保证特定于被审计的模型、提示、阈值和输入分布;将其扩展到人类解释需要外部验证。

英文摘要

Large language models increasingly provide labels, evaluations, and feedback for tasks specified in natural language. When a specification admits multiple readings but the supervision channel does not reveal which is operative, additional labels reduce sampling error without resolving the resulting identification problem. We introduce Natural Language PAC (NL-PAC), a framework that uses a fixed model's thresholded decoding law to define admissible labels and candidate targets. The probability that multiple labels are admissible equals the diameter of the pointwise-admissible target class, and under target-blind supervision every learner incurs worst-case risk of at least half this diameter, at every sample size; the exact randomized minimax risk over this class is attained by a data-independent strategy. Finite-sample confidence bounds make these quantities certifiable from held-out unlabeled inputs. In a frozen Qwen~2.5--3B audit, one prespecified prompt yields a positive model-relative certificate, whereas a paraphrase and exact-rule controls yield zero. A held-out bridge audit finds that supplied candidate reading clauses fail the admissibility condition needed to transfer the certificate to coherent readings. The guarantee is specific to the audited model, prompt, threshold, and input distribution; extending it to human interpretations requires external validation.

↑