arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19575cs.PL

面向自适应数据分析的程序分析

Program Analysis for Adaptive Data Analysis

Jiawen Liu, Weihao Qu, Marco Gaboardi, Deepak Garg, Jonathan Ullman

AI总结:

本研究将自适应数据分析实现为类while程序,设计程序分析近似其适应性量化属性,实现后可分析多种不同适应性结构的具体数据分析。

AI中文摘要:

数据分析通常旨在从抽取数据的总体中识别出某种属性,进而泛化到特定数据样本之外。出于这个原因,数据分析常被设计为产生低泛化误差,使得对样本数据的分析结果与对整个总体分析所得结果不会出现过大差异。自适应数据分析可视为由多个查询构成的过程,这些查询对某些数据进行查询,且下一个要运行的查询的选择可能依赖于之前查询的结果。每个单独查询的泛化误差可通过成熟的统计技术加以控制,但当查询被任意组合时,误差会在查询链中传播,进而导致高泛化误差。为解决这一问题,已有若干技术不仅保证单个查询的边界,也保证组合分析的边界,而技术的选择往往取决于自适应数据分析可生成的查询链。在本研究中,我们将自适应数据分析实现为类while程序,并设计一种程序分析,以帮助确定应使用何种技术来控制其泛化误差。更具体地说,我们将适应性的直观概念形式化为程序的量化属性,基于该定义设计一种程序分析,以合理近似该量化属性。该分析将数据分析表示为带权依赖图,其中权重对变量可达的频率进行上界约束,并采用路径搜索策略对适应性进行上界约束。我们实现了该程序分析,结果表明其可分析多种具有不同适应性结构的具体数据分析。

英文摘要:

Data analyses are usually designed to identify some property of the population from which the data are drawn, generalizing beyond the specific data sample. For this reason, data analyses are often designed to produce a low generalization error, so that the result of an analysis on sample data does not differ too much from the result one would achieve over the entire population. An adaptive data analysis can be seen as a process composed of multiple queries interrogating some data, where the choice of which query to run next may rely on the results of previous queries. The generalization error of each individual query can be controlled using well-established statistical techniques. However, when queries are arbitrarily composed, errors can propagate through the chain of queries and lead to high generalization error. To address this issue, several techniques guarantee bounds not only on single queries but also on composed analyses. The choice of technique often depends on the chain of queries that an adaptive data analysis can generate. In this work, we consider adaptive data analyses implemented as while-like programs and design a program analysis to help identify which technique to use to control their generalization errors. More specifically, we formalize the intuitive notion of adaptivity as a quantitative property of programs. Based on this definition, we design a program analysis for soundly approximating this quantity. The analysis represents the data analysis as a weighted dependency graph, where weights upper-bound how often variables can be reached, and uses a path-search strategy to upper-bound adaptivity. We implement our program analysis and show that it can analyze several concrete data analyses with different adaptivity structures.

补充信息

↑