arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于改进随机索引嵌入的稀疏互信息图平均方法

Sparse PPMI Graph Averaging for Random Indexing Embeddings

Sriram Loganathan, Gokul Anand, Aung Bo Bo, Yourui Shao, William B. Andreopoulos

arXiv 2608.05724首次发表:更新:

AI 中文总结

本文提出带前K剪枝的PPMI图平均方法,可修复弱RI嵌入,在童话语料库的Google家族类比任务上使RI准确率提升至30.7±2.9%,是针对弱RI嵌入的有效非梯度修复手段。

AI 中文摘要

稀疏词嵌入流水线可避免稠密共现矩阵的实例化、稠密分解及梯度训练,同时仍依赖稀疏全局语料库统计信息。本文研究通过在稀疏正点互信息(PPMI)图上加权平均优化的随机索引(RI)向量。在童话语料库上,覆盖的语义类比集包含272个Google家族类问题。在该家族子集上,PPMI前K图平均可修复较弱的RI初始化,在5个随机种子下,准确率从19.4±0.7%提升至30.7±2.9%。在单次测试运行中,相同的邻域平均会降低PPMI+奇异值分解(SVD)、二值化+SVD、连续词袋模型(CBOW)及跳元模型(Skip-gram)的家族子集类比准确率。因此,该方法在text8数据集上无法与神经基线竞争,且在SimLex-999数据集上的严格相似度相关性接近零。尽管在测试配置下布隆过滤器草图的表现不如RI,但我们发现带前K剪枝的PPMI图平均是针对弱RI嵌入的有用非梯度修复方法。在童话数据集上,PPMI前K=50图平均可提升RI,准确率从19.4±0.7%提升至30.7±2.9%,在种子42时达到最佳的34.6%。

英文摘要

We study a specific sparse post-processing pipeline for Random Indexing (RI) on kinship analogies in a small fairytales corpus. The published artifacts use uniform RI context accumulation with 200 dimensions and eight nonzeros, followed by one residual graph average, $\mathbf{E}=(1-α)\mathbf{E}_0+α\mathbf{P}\mathbf{E}_0$, where $\mathbf{P}$ is a row-normalized PPMI graph and $α=0.3$. Terminal row normalization and per-dimension median/IQR scaling are then applied. On the Google analogy benchmark's family section, 272 of 506 questions are valid for every seed. Across five paired seeds, the complete pipeline raises accuracy from 19.41\% to 30.74\%, a gain of 11.32 percentage points with a nested-bootstrap 95\% confidence interval of [6.93, 15.89]. Robust scaling alone contributes 3.24 points [1.25, 5.38], while graph averaging without robust scaling contributes 6.18 points [2.63, 9.92]. A separate 40-question general grid does not support a general improvement: the full pipeline changes accuracy by -6.00 points [-13.50, -0.50], and averaging without robust scaling changes it by -6.50 points [-14.50, -0.50]. The supported positive claim is therefore limited to the covered fairytales kinship analogy set; the results do not establish a generally effective embedding method.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑