压缩向量表示中的精确语义读出
Exact semantic readout from compressed vector representations
浏览论文内容
中文总结 AI 辅助
本研究刻画了压缩向量表示支持精确线性或仿射语义读出的条件,提出行空间秩条件,并通过实验表明预训练嵌入仅可分而不可精确仿射恢复,但监督转导训练可达到精确恢复。
中文摘要 AI 辅助
我们刻画了压缩向量表示何时允许对有限词典的真值条件进行精确的线性或仿射读出:每个谓词对应一个固定的映射,将每个实体向量发送到相应的真值向量。一个必要且充分的行空间条件决定了存在性;增广真值矩阵的秩为r,在线性情形下给出最小维度r,在仿射情形下为r-1。精确读出在共享的真值基上返回值,布尔连接词在该基上不变地作用;仅可分性需要介入阈值。对于二元关系,恒等或严格全序的精确双线性读出要求实体向量线性无关。使用GloVe和word2vec的实验区分了精确仿射恢复、线性可分性和留出预测:大多数谓词是严格可分的,但没有任何谓词允许从预训练嵌入中进行精确仿射读出。有监督的转导训练在满足界限的每个测试维度上达到数值精度的精确仿射恢复。在嵌入的原始维度上,受限于精确线性恢复的几何结构在特征范数上保留了预训练方差的98-99%,在WordNet词典上保留了80-83%。
英文摘要
We characterize when compressed vector representations admit exact linear or affine readouts of a finite lexicon's truth conditions: one fixed map per predicate, sending each entity vector to the corresponding truth vector. A necessary and sufficient row-space condition determines existence; the augmented truth matrix has rank r, giving minimum dimension r in the linear case, and r-1 in the affine. Exact readouts return values in a shared truth basis on which Boolean connectives act unchanged; separability alone requires an intervening threshold. For binary relations, exact bilinear readout of identity or strict total order requires linearly independent entity vectors. Experiments with GloVe and word2vec distinguish exact affine recovery, linear separability, and held-out prediction: most predicates are strictly separable, but none admits an exact affine readout from the pretrained embeddings. Supervised transductive training attains exact affine recovery to numerical precision at every tested dimension meeting the bound. At the embeddings' original dimension, geometries constrained to exact linear recovery retain 98-99 percent of the pretrained variance on the feature norms, and 80-83 percent on the WordNet lexicon.
发表机构
- Center for Possible Minds(可能心智中心)
- Indiana University Bloomington(印第安纳大学布卢明顿分校)
机构由 AI 辅助整理,请以论文原文为准。