arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Hold-Out Scoring 用于高效高斯 DAG 学习

Hold-Out Scoring for Efficient Gaussian DAG Learning

Donguk Shin, Byeongguk Kang, Inseol Lee, Gunwoong Park

arXiv 2610.02785首次发表:更新:

发表机构

Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出 HOST 算法,通过留出评分与凸回归替代子集搜索,实现无需入度上界的高效高斯 DAG 学习,以 $d\log p$ 样本复杂度在多项式时间内精确恢复图结构。

AI 中文摘要

高维高斯 DAG 学习面临统计与计算之间的差距:具有尖锐样本复杂度的方法依赖于计算代价高昂的子集搜索和预设的入度上界,而多项式时间替代方法的样本复杂度则较差。我们提出 HOST,一种高效的 DAG 学习算法,它用逐节点的留出评分和凸回归取代子集搜索,且无需预设入度上界。我们的关键洞察是,恢复正确的排序并不需要排序评分具有均匀小的估计误差,而只需对这些误差进行单侧控制。在排序步骤中,HOST 利用留出样本进行评分估计会使排序评分在期望上膨胀,这对于那些尚不应被选中的候选者是有利的方向。给定排序后,HOST 通过递归地从两节点间的总效应中移除间接效应来恢复父节点。在适当条件下,HOST 能以多项式时间精确恢复最大入度为 $d$ 的 $p$ 节点 DAG,样本复杂度为 $d\log p$ 量级。实验表明,HOST 在实现有竞争力的图恢复性能的同时,展现出良好的运行时间扩展性。

英文摘要

High-dimensional Gaussian DAG learning faces a statistical-computational gap: methods with sharp sample complexity rely on computationally expensive subset search and a supplied indegree bound, whereas polynomial-time alternatives have less favorable sample complexity. We introduce HOST, an efficient DAG learning algorithm that replaces subset search with nodewise hold-out scoring and convex regression, without requiring a supplied indegree bound. Our key insight is that recovering a correct ordering does not require uniformly small estimation errors in ordering scores but only one-sided control of those errors. In the ordering step, HOST exploits the fact that score estimation using hold-out samples inflates ordering scores in expectation, which is the favorable direction for candidates that should not yet be selected. Given the ordering, HOST recovers parents by recursively removing indirect effects from total effects between two nodes. Under suitable conditions, HOST exactly recovers a $p$-node DAG of maximum indegree $d$ with sample complexity of order $d\log p$ in polynomial time. Experiments show that HOST achieves competitive graph recovery while exhibiting favorable runtime scaling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑