AI 中文总结
研究边缘设备跨模态检索在CiM硬件部署问题,提出PolySim框架,将概率检索转为确定性管道,用低阶多项式基近似高斯嵌入维度,经实验其提升R@1,减少推理计算,是首个在CiM硬件实现概率跨模态检索的方法。
AI 中文摘要
边缘设备上的跨模态检索受益于捕获语义不确定性的概率嵌入,但在内存计算(CiM)硬件上部署它们仍然是一个开放问题。核心困难是采样差距:诸如PCME之类的概率方法在推理时依赖蒙特卡罗采样和非线性距离评估,这与仅支持确定性单步矩阵向量乘法的CiM交叉开关阵列根本不兼容。现有的概率检索方法很少能在传统交叉开关上执行。为了弥合这一差距,我们提出了PolySim,一个将概率检索重新制定为完全确定性管道的框架。PolySim使用低阶多项式基近似每个高斯嵌入维度,并通过可学习的阶双线性核计算相似度,在保留分布信息的同时消除随机采样。在跨越视频、图像和音频检索的六个基准测试中,PolySim将R@1比确定性基线提高了高达10.3%,并匹配或超过了PCME,同时将推理减少到单个与交叉开关兼容的矩阵向量乘法。在实际设备非理想情况下的CrossSim评估证实了在传统交叉开关阵列上的稳健部署。据我们所知,PolySim是第一种在CiM硬件上实现概率跨模态检索的方法。
英文摘要
Cross-modal retrieval on edge devices benefits from probabilistic embeddings that capture semantic uncertainty, but deploying them on compute-in-memory (CiM) hardware remains an open problem. The core difficulty is a sampling gap: probabilistic methods such as PCME rely on Monte Carlo sampling and nonlinear distance evaluation at inference, which are fundamentally incompatible with CiM crossbar arrays that support only deterministic, single-step matrix-vector multiplication. Few existing probabilistic retrieval methods can be executed on a conventional crossbar. To bridge this gap, we propose PolySim, a framework that reformulates probabilistic retrieval into a fully deterministic pipeline. PolySim approximates each Gaussian embedding dimension using low-order polynomial bases and computes similarity via a learnable order-bilinear kernel, eliminating stochastic sampling while preserving distributional information. In experiments on six benchmarks spanning video, image, and audio retrieval, PolySim improves R@1 over deterministic baselines by up to 10.3\% and matches or exceeds PCME, while reducing inference to a single crossbar-compatible matrix-vector multiplication. CrossSim evaluation under realistic device non-idealities confirms robust deployment on conventional crossbar arrays. To the best of our knowledge, PolySim is the first method to enable probabilistic cross-modal retrieval on CiM hardware.
CommentsAccepted by ICCAD 2026