arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自适应数据分析是否需要随机性?

Is Randomness Necessary for Adaptive Data Analysis?

Edith Cohen, Haim Kaplan, Yishay Mansour, Shay Sapir, Uri Stemmer

arXiv 2607.07085首次发表:更新:

发表机构

Edith Cohen; Haim Kaplan; Yishay Mansour; Shay Sapir; Uri Stemmer(Edith Cohen; Haim Kaplan; Yishay Mansour; Shay Sapir; Uri Stemmer)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究自适应数据分析中查询数量与样本数量关系,针对随机机制已有成果。探讨ADA是否需随机性这一基本问题,此前未解决。在信息论随机预言模型中证明,无界分析师情况下,大量自适应查询时随机性严格必要,确定性机制易失败。

AI 中文摘要

自适应数据分析(ADA)问题将重复使用数据集时防止错误发现和过拟合的挑战形式化。输入是包含来自未知分布\(\mathcal{P}\)的\(n\)个独立同分布样本的数据集,目标是针对\(\mathcal{P}\)回答一系列\(k\)个自适应选择的统计查询。主要问题是能支持多少查询(即\(k\)能多大),这对随机机制已较了解。本文探讨ADA是否需要随机性这一基本问题,此前虽有研究但仍未解决。我们在信息论随机预言模型中解决了这一差距,表明对于大量自适应查询,随机性是严格必要的:无界分析师时,任何确定性机制在\(k = \tilde{O} (n)\)次查询后就会失败。

英文摘要

The Adaptive Data Analysis (ADA) problem formalizes the challenge of preventing false discovery and overfitting when a dataset is repeatedly reused. Formally, our input is a dataset containing $n$ i.i.d.\ samples from an unknown distribution $P$ over a domain $X$, and our goal is to answer a sequence of $k$ adaptively chosen statistical queries with respect to $P$. The main question is how many queries we can support (i.e., how large $k$ can be), primarily as a function of the number of samples $n$. This question has been intensively studied and is relatively well-understood for randomized mechanisms: there are computationally efficient mechanisms that support $k \approx n^2$ queries, and no computationally efficient mechanism can answer $k \gg n^2$ queries. In this paper, we address a fundamental question: is randomness necessary for ADA? Despite a decade of work on ADA, this question remains open. A folklore observation dating back to the initial works on ADA is that randomness is {\em not} necessary when the analyst is computationally bounded. Yet, the necessity of randomness against computationally unbounded analysts has remained elusive. Our main contribution resolves this gap in the information-theoretic setting. Perhaps surprisingly, we show that randomness is strictly necessary to answer a non-trivial number of adaptive queries: when the analyst is unbounded, any deterministic mechanism can be forced to fail after just $k = \tilde{O}(n)$ queries.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑