从多个数据提供者学习分布
Learning Distributions from Multiple Data Providers
查看机构详情
- Cornell University(康奈尔大学)
- Stanford University(斯坦福大学)
- Google Research(谷歌研究院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究从受限条件样本学习有限域[n]上未知分布p的模型,基于共现图探讨可学习性。分析PAC学习样本复杂度,指出完全共现图时为O(n^2/ε^2),[n]可查询时为Θ(n/ε^2),还确定层次可比性条件下最优复杂度近线性,且线性到二次间多项式速率范围可实现。
中文摘要 AI 辅助
受从异构和重叠数据提供者学习的启发,我们研究了从受限条件样本进行分布学习的程式化模型。目标是学习有限域[n]上的未知分布p。学习者被给予一个固定的可查询集族S⊆2^[n],对S∈S的每个查询返回来自条件分布p(·|S)的独立样本。可学习性由与S相关的共现图决定。当该图在目标支持上连通时可实现逐点一致性,PAC学习则要求更多,当共现图完全时才可能。PAC学习的最优样本复杂度从近线性到二次。每个具有完全共现图的查询族的样本复杂度为O(n^2/ε^2),在最坏情况下此界是紧的。另一方面,如果[n]是可查询的,普通采样将界提高到Θ(n/ε^2),即使每个集都是可查询的也不能进一步改进。更一般地,我们确定层次可比性是S上的一个充分结构条件,在此条件下最优复杂度近线性为Θ(n/ε^2),成对查询族是典型例子。最后,线性和二次之间的整个多项式速率范围都是可达到的:对于每个α∈(1,2),存在一个具有最优PAC速率Θ(n^α/ε^2)的查询族。
英文摘要
Motivated by learning from heterogeneous and overlapping data providers, we study a stylized model of distribution learning from restricted conditional samples. The goal is to learn an unknown distribution $p$ on a finite domain $[n]$. The learner is given a fixed family of queryable sets $\mathscr{S} \subseteq 2^{[n]}$, and each query to $S \in \mathscr{S}$ returns an independent sample from the conditional distribution $p(\cdot \mid S)$. Learnability is governed by the co-occurrence graph associated with $\mathscr{S}$: two domain elements are adjacent if they appear together in some queryable set. Pointwise consistency is achievable when this graph is connected on the target support. PAC learning requires more: it is possible when the co-occurrence graph is complete. The optimal sample complexity of PAC learning ranges from nearly linear to quadratic. Every query family with complete co-occurrence graph admits sample complexity $\widetilde O(n^2/ε^2)$, and this bound is tight in the worst case. On the other hand, if $[n]$ is queryable then ordinary sampling improves the bound to $Θ(n/ε^2)$, and this cannot be improved further even if every set is queryable. More generally, we identify hierarchical comparabilityas a sufficient structural condition on $\mathscr S$ under which the optimal complexity is nearly linear, $\widetilde Θ(n/ε^2)$, with pairwise query families as a canonical example. Finally, the full range of polynomial rates between linear and quadratic is attainable: for every $α\in (1,2)$, there exists a query family with optimal PAC rate $\widetilde Θ(n^α/ε^2)$.