arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18817stat.MLcs.LG

概率张量结构学习中的代数签名

Algebraic Signatures for Structural Learning in Probability Tensors

发表机构东京大学艺术与科学研究生院 · 日本先进科学研究院 · Shiga大学数据科学学系
查看机构详情
  • Graduate School of Arts and Sciences, The University of Tokyo, Tokyo, Japan(东京大学艺术与科学研究生院)
  • Japan Advanced Institute of Science and Technology, Nomi, Ishikawa, Japan(日本先进科学研究院)
  • Faculty of Data Science, Shiga University, Hikone, Shiga, Japan(Shiga大学数据科学学系)

机构由 AI 辅助整理,请以论文原文为准。

Akihiro Maeda, Shohei Hidaka, Satoshi Aoki

首次发表
浏览论文内容

中文总结 AI 辅助

研究从经验概率张量中观察到的消失二项式识别概率结构的逆问题,核心方法是将环面模型的消失二项式视为代数签名,通过签名匹配识别模型,在合成及真实语言数据上测试,结果为代数统计应用于计算语言学开辟新途径。

中文摘要 AI 辅助

代数统计通过多项式约束来刻画统计模型,但主要用于解析指定的模型类。本文研究逆问题:从经验概率张量中观察到的消失二项式识别概率结构。我们将环面模型的消失二项式视为其代数签名,把代数统计的理想-簇对应转化为结构学习的操作程序,通过签名匹配识别模型而无需参数估计。通过关注计算上易处理的配置矩阵类(即克罗内克堆栈类),使这些签名可明确枚举。在该类中定义最小不变约束(MIC)作为表征每个签名并推广独立性概念的原子单位。我们在合成数据以及语料库规模的真实语言数据上使用MIC测试了此方法。结果表明该方法有用,所识别的秩一结构对应于可解释的单词集。这些结果为将代数统计应用于计算语言学开辟了新途径。

英文摘要

Algebraic statistics characterizes statistical models through polynomial constraints, but it has mainly been used for analytically specified model classes. This paper studies the inverse problem: identifying probabilistic structure from vanishing binomials observed in empirical probability tensors. We treat the vanishing binomials of a toric model as its algebraic signature, and turn the ideal-variety correspondence of algebraic statistics into an operational procedure for structural learning that identifies a model by signature matching without parameter estimation. By restricting attention to a computationally tractable class of configuration matrices, which we call {\it the Kronecker-stack class}, we make these signatures explicitly enumerable. Within this class we define minimum invariant constraint (MIC) as the atomic unit characterizing each signature and generalizing the notion of independence. We tested this approach employing MICs on synthetic data as well as on corpus-scale real language data. The results suggested the utility of the method, revealing that the identified rank-one structures correspond to interpretable sets of words. These results open up a new avenue for applying algebraic statistics to computational linguistics.

补充信息

↑