发表机构
Indiana University(印第安纳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对生成数据驱动的相似性感知统计,本文证明在高斯核下,利用向量几何结构可实现单遍次线性空间流式近似算法,并给出下界,揭示几何结构能改变流式复杂性。
AI 中文摘要
受生成系统产生的数据启发,\cite{LZ26b} 通过加权相似图形式化了相似性感知统计,在经典基于频率的统计中用相似性取代相等性。尽管该框架能够捕捉非相同项之间的语义关系,但在一般相似函数下,即使粗略的单遍近似也可能需要线性空间。因此,我们探究自然向量相似性中存在的几何结构能否克服这一障碍。对于高斯核,我们对此问题给出肯定回答。对于固定维欧几里得向量流,我们通过多样性指数和高斯密度矩,研究经典频率统计(包括不同元素数量和频率矩)的相似性感知类比。我们给出利用高斯核的几何与分析性质的单遍次线性空间近似算法,并辅以相应下界。我们的结果表明,几何结构能从根本上改变相似性感知统计分析的流式复杂性。
英文摘要
Motivated by data produced by generative systems, \cite{LZ26b} formulates similarity-aware statistics via a weighted similarity graph, replacing equality with similarity in classical frequency-based statistics. Although this framework captures semantic relationships between nonidentical items, under general similarity functions even coarse one-pass approximation can require linear space. We therefore ask whether the geometric structure present in natural vector similarities can overcome this barrier. We answer this question affirmatively for the Gaussian kernel. For fixed-dimensional Euclidean vector streams, we study similarity-aware analogues of classical frequency statistics, including the number of distinct elements and frequency moments, through the diversity index and Gaussian density moments. We give one-pass sublinear-space approximation algorithms that exploit the geometric and analytic properties of the Gaussian kernel, and complement them with lower bounds. Our results show that geometric structure can fundamentally change the streaming complexity of similarity-aware statistical analysis.
Comments37 pages