arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Tacet:用于自动统计有效性核算的语言与类型系统

Tacet: A Language and Type System for Automatic Statistical Validity Accounting

Chiké Abuah

arXiv 2608.27451首次发表:更新:

AI 中文总结

本文提出用于自动统计有效性核算的语言与类型系统Tacet,其核心演算结合自由估计与代价主张子语言,通过Lean 4验证元理论,并在SWE-bench Verified排行榜等案例中验证了方法。

AI 中文摘要

系统间的实证比较是计算机科学研究中标准的证据形式,但很少有研究对其统计有效性进行检查:大多数比较根本未被构建为统计检验。现有的多重比较程序可控制由此产生的误差,但其所需的输入(分析所考察的内容以及观测值的排列方式)无法从p值列表中恢复。我们引入Tacet,这是一种语言,分析过程可在其中声明自身生成的内容、陈述预期发现的结果,且任何无法承担或无法正确检验的主张都会被拒绝。其核心演算T结合了自由估计子语言与代价主张子语言:自由估计子语言携带报告的 footprint(足迹)和纯度位,纯度位记录在构建值时是否参考了任何结果;代价主张子语言携带财富变换器,二者仅通过对比较进行定价的机制相连接。通过读取结果选定的样本会设置纯度位,并被永久记录为已考察所读取的所有内容,因此永远无法被赋予单侧或确证性代价,而无需系统询问分析者是否存在刻意 cherry-pick( cherry-pick 指选择性选取有利结果)的行为。比较是否配对或聚类是在读取任何数据之前,仅从工件模式、键字段之间声明的函数依赖关系静态计算得出的,且假设该结构被忽略的机制会被拒绝而非定价。由于财富变换器在已实现的p值中是反单调的,可在分析运行前检查可负担性,从而将预注册转化为类型规则。我们在Lean 4中对元理论进行了机器验证,无任何承认的 gaps( gaps 指漏洞),并在参考实现以及两个针对已发表工件的案例研究(SWE-bench Verified排行榜和BIG-Bench Hard)上验证了该方法。

英文摘要

Empirical comparisons between systems are a standard form of evidence in computer science research, but few are checked for statistical validity: most are never framed as statistical tests at all. Existing multiple-comparison procedures could control the resulting error, but need inputs (what an analysis examined, and how its observations are arranged) that are not recoverable from a list of p-values. We introduce Tacet, a language in which an analysis declares what it generated, states what it expects to find, and is refused any claim it cannot afford or cannot properly test. Its core calculus T pairs a free estimation sublanguage, carrying a reported footprint and a purity bit that records whether any outcome was consulted in building a value, with a priced claim sublanguage, carrying a wealth transformer, connected only by a mechanism that prices a comparison. A sample selected by reading outcomes sets the purity bit and is recorded as having examined everything it read, permanently, so it can never be granted a one-sided or confirmatory price, without the system ever asking whether the analyst intended to cherry-pick. Whether a comparison is paired or clustered is computed statically from the artifact schema, from declared functional dependencies between key fields alone and before any data is read, and a mechanism that assumes that structure away is refused rather than priced. Because the wealth transformer is antitone in the realized p-value, affordability can be checked before the analysis runs too, turning pre-registration into a typing rule. We prove the metatheory machine-checked in Lean 4 with no admitted gaps, and demonstrate the approach on a reference implementation and two case studies on published artifacts, the SWE-bench Verified leaderboard and BIG-Bench Hard.

Comments67 pages, 2 figures, 10 tables, including 10 appendices. Lean 4 mechanization: https://github.com/abuach/tacet-mech ; reference implementation and case-study replication code: https://github.com/abuach/tacet-python

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑