arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

网络结构的事后选择推断

Post-Selection Inference for Network Structure

Eric Auerbach, Jonathan Auerbach, Sidonia McKenzie

arXiv 2607.00312首次发表:更新:

AI 中文总结

针对网络数据中基于连接密度选择组群导致的推断无效问题,提出两种事后选择置信区间,其中基于Talagrand型浓度不等式的区间渐近最优,并在实证中验证了校正选择偏差的必要性。

AI 中文摘要

研究者常使用代理群体(如社区、集团或市场)之间的连接密度来刻画社会或经济网络的结构。在许多情况下,这些群体是利用网络数据选择的,这使得传统的固定群体推断程序可能无效。为解决这一问题,我们开发了两种新的置信区间,它们在事后选择意义下普遍有效,即保证渐近地覆盖所有相对规模不消失的群体对。我们的第一个区间建立在\cite{berk2013valid}的策略之上。第二个区间基于经验过程的Talagrand型浓度不等式。两个区间都易于计算且可扩展至大型网络,但我们论文的一个关键技术贡献是证明只有第二个区间在渐近意义上达到了最佳可能宽度(至多常数因子)。三个实证示例表明,考虑选择在实践中可能很重要。社交网络中的同质性证据和贸易网络中的枢纽-轮辐结构证据在修正后仍然存在,而工人转换网络中的分离市场细分证据则不再成立。

英文摘要

Researchers often use the density of connections between groups of agents, such as communities, blocs, or markets, to characterize the structure of a social or economic network. In many cases, these groups are selected using the network data, making conventional fixed-group inference procedures potentially invalid. To address this issue, we develop two new confidence intervals that are universally valid post-selection in the sense that they guarantee simultaneous coverage asymptotically over all pairs of groups whose relative sizes do not vanish. Our first interval builds on a strategy of Berk et al. (2013). Our second interval is based on a Talagrand-type concentration inequality for empirical processes. Both intervals are simple to compute and scalable to large networks, but a key technical contribution of our paper is to show that the second interval is rate-optimal over a broader class of intervals. Three empirical illustrations show that accounting for selection can matter in practice. Some evidence for homophily in a social network and a hub-and-spoke structure in a trade network survives our correction, while evidence for a segmented market structure in a worker transition network does not.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑