发表机构
Zhejiang Normal University; Georgia State University(浙江师范大学; 佐治亚州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对公平范围摘要问题,提出目标分层采样方法,在人口统计均等和自定义比例目标下分别达到最优样本量,并改进公平几何命中集近似比,实验验证了其高效性。
AI 中文摘要
紧凑摘要是大型数据集上近似查询处理的关键工具。对于范围查询工作负载,$\varepsilon$-网提供了一个小摘要,能够命中所有足够大的范围。然而,经典的$\varepsilon$-网仅保证范围有效性,并不控制所选元组的群体构成。因此,摘要可能范围有效但代表性差,这可能将不平衡传播到下游查询结果中。受近期关于公平$\varepsilon$-网和公平几何命中集的研究启发,我们在规定的目标群体比例下研究公平感知的范围摘要。不同于先前的采样-修复方法,我们提出了一种目标分层采样方法。对于人口统计均等(其中公平比例由群体比例决定),我们的样本量为$O(A_{\varepsilon})$,与标准$\varepsilon$-网界限一致,改进了先前$O\\!\left(A_\varepsilon\log\frac{k}{\varphi}\right)$的界限。对于自定义比例目标(其中公平比例由手动定义的比例决定),我们的样本量为$O(A_{\Gamma})$,其中$\Gamma$是衡量自定义比例与人口统计均等之间差距的参数;我们证明了这种对$\Gamma$的依赖是不可避免的,最坏情况下的下界为$\Omega(\Gamma/\varepsilon)$。利用我们的目标分层采样方法,我们可以将公平几何命中集问题的先前近似比改进一个对数因子,并利用这一结果,反过来改进自定义比例公平$\varepsilon$-网的大小。在真实和合成数据集上的实验表明,我们的方法比现有方法构建更小的公平摘要,可扩展到大型数据集和细粒度群体约束,并改善下游范围查询处理。
英文摘要
Compact summaries are a key tool for approximate query processing over large datasets. For range-query workloads, an $\varepsilon$-net provides a small summary that hits every sufficiently large range. However, classical $\varepsilon$-nets only guarantee range validity and do not control the group composition of the selected tuples. As a result, the summary may be range-valid but poorly representative, which can propagate imbalance to downstream query results. Motivated by recent work on fair $\varepsilon$-nets and fair geometric hitting sets \cite{dehghankar2025fair}, we study fairness-aware range summaries under prescribed target group ratios. Different from previous sample-and-repair approach, we propose a target-stratified sampling method. For demographic parity (in which the ratio of fairness is determined by group proportion), our sample size is $O(A_{\varepsilon})$, coinciding with the standard $\varepsilon$-net bound, improving previous bound of $O\!\left(A_\varepsilon\log\frac{k}φ\right)$. For custom-ratio targets (in which the ratio of fairness is determined by manually defined proportion), our sample size is $O(A_Γ)$, where $Γ$ is a parameter measuring the gap between the customized ratio and the demographic parity; we prove that this dependence on $Γ$ is unavoidable, with a worst-case lower bound of $Ω(Γ/\varepsilon)$. Using our target-stratified sampling method, we could improve the previous approximation ratio for the fair geometric hitting set problem by a logarithmic factor, and making use of this result, we could in turn improve the size of custom-ratio fair $\varepsilon$-net. Experiments on real and synthetic datasets demonstrate that our method constructs smaller fair summaries than existing approaches, scales to large datasets and fine-grained group constraints, and improves downstream range query processing.