arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

每个边际一个提案:二元矩阵的零样本摊销序贯重要性采样

One Proposal for Every Margin: Zero-Shot Amortized Sequential Importance Sampling for Binary Matrices

Ruishuo Chen, Weijia Li, Xun Wang, Yu Chen, Leheng Cai, Longbo Huang

arXiv 2609.35514首次发表:更新:

发表机构

Institute for Interdisciplinary Information Sciences, Tsinghua University; Department of Statistics and Data Science, Tsinghua University(清华大学交叉信息研究院; 清华大学统计与数据科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对二元矩阵的计数与均匀采样问题,提出MarginFlow框架,将SIS提案设计转化为学习问题,利用边际自相似性实现零样本摊销,在1190个测试边际上优于31种解析配置,中位有效样本分数达99.8%。

AI 中文摘要

在生态学、心理测量学以及社会与金融网络分析中,二元矩阵通常在其观测到的行和列总和(即边际)条件下进行分析,这将该问题限制在具有相同边际的矩阵的有限样本空间内。两个基本问题是计算该空间的大小并从中均匀采样。序贯重要性采样(SIS)通过独立的加权样本和无偏计数估计器解决这两个问题,但其效率关键取决于提案分布。现有提案是解析设计的,其准确性可能随边际变化而显著波动。我们证明,理想的SIS提案(其中每个权重等于计数且方差为零)恰好是生成流网络(GFlowNet)的策略,该网络对具有给定边际的每个矩阵赋予单位奖励。因此,我们提出MarginFlow,一个将提案设计转化为学习问题并利用边际的自相似性跨边际进行摊销的框架。每个部分矩阵本身就是一个具有缩减边际的实例,因此一个读取剩余边际的集合变换器即可服务于每个边际。我们在1904个边际的池上训练MarginFlow,并在1190个保留边际(合成和真实,从3×3到870×6)上零样本评估。在1190个边际中的1187个上,它匹配或超越了31个解析设计配置(为每个边际事后选择)中的最佳者,其中位有效样本分数为99.8%。在最佳配置损失超过1纳特有效样本量的56个边际上,MarginFlow全部胜出,并将中位有效样本分数从10.3%提升至94.1%。

英文摘要

In ecology, psychometrics, and the analysis of social and financial networks, binary matrices are often analyzed conditional on their observed row and column sums, which restricts the problem to a finite sample space of matrices with the same margins. Two fundamental problems are to count this space and to sample uniformly from it. Sequential importance sampling (SIS) addresses both with independent weighted samples and an unbiased count estimator, but its efficiency depends critically on the proposal distribution. Existing proposals are analytically designed, and their accuracy can vary substantially with the margins. We show that the ideal SIS proposal, under which every weight equals the count and the variance vanishes, is exactly the policy of a generative flow network (GFlowNet) with unit reward on every matrix that has the given margins. We therefore propose MarginFlow, a framework that turns the design of the proposal into a learning problem and amortizes it across margins by exploiting their self-similarity. Every partial matrix is itself an instance with reduced margins, so one set transformer that reads the remaining margins serves every margin. We train MarginFlow on a pool of 1904 margins and evaluate it zero-shot on 1190 held-out margins, synthetic and real, from $3\times3$ to $870\times6$. On 1187 of the 1190 margins it matches or beats the best of 31 analytically designed configurations, chosen post hoc for each margin, and its median effective sample fraction is 99.8%. On the 56 margins where that best loses more than one nat of effective sample size, MarginFlow wins every one and raises the median effective sample fraction from 10.3% to 94.1%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑