arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种保持对称性的张量$\boldsymbol{\rm \bigstar_M}$-SVD

A Symmetry-Preserving Tensor $\star_{\mathbf{M}}$-SVD

Victor Arsenescu, Misha E. Kilmer

arXiv 2608.24985首次发表:更新:

AI 中文总结

该研究提出保持对称性的$\boldsymbol{\rm \bigstar_M}$-SVD,将其应用于对称图像识别,在保持识别率的同时大幅降低基存储量,人脸光照场景下性能优于普通张量SVD。

AI 中文摘要

图像集合和视频等多路数据无处不在,但将其展平为矩阵的常规方法会丢弃通常承载信号的跨模式结构。t-乘积及其推广形式$\boldsymbol{\rm \bigstar_M}$-乘积提供了一种类矩阵的张量代数,其张量SVD的截断在Frobenius范数下是最优的,与矩阵情形相同。许多真实数据还具有内在的反射对称性:正面人脸、工业零件和叶片均为双侧对称。我们定义了一种保持对称性的$\boldsymbol{\rm \bigstar_M}$-SVD,将Shah和Sorensen提出的矩阵保持对称性SVD推广到$\boldsymbol{\rm \bigstar_M}$代数中。当张量的变换域正面切片具有反射对称性时,其$\boldsymbol{\rm \bigstar_M}$-SVD的左基也具有对称性,因此只需存储其一半。将双侧对称图像横向放置并作为侧面切片存储,即可得到此类张量。该构造保留了每幅图像的对称部分,且仅存储一半的基。随后,我们通过将新图像投影到该基上并在系数空间中匹配最近的训练图像来识别新图像。在包含人脸、叶片、蝴蝶翅膀及其他物体的10个数据集上,对称基的识别率与普通张量SVD相当,但基存储量减少了2至11倍,MUCT数据集除外。在不同光照下的人脸数据集上,其识别率超过了普通张量SVD在任何存储水平下所能达到的最佳识别率。

英文摘要

Multiway data such as image collections and video is ubiquitous, but the usual approach of flattening them into matrices discards the cross-mode structure that often carries the signal. The t-product and its generalization, the $\star_{\mathbf{M}}$-product, give a matrix-mimetic tensor algebra with a tensor SVD whose truncation is optimal in the Frobenius norm, just as in the matrix case. Much real data also has internal reflective symmetry: frontal faces, manufactured parts, and leaves are all bilaterally symmetric. We define a symmetry-preserving $\star_{\mathbf{M}}$-SVD that extends the matrix symmetry-preserving SVD of Shah and Sorensen to the $\star_{\mathbf{M}}$-algebra. When a tensor's transform-domain frontal slices are reflectively symmetric, the left basis of its $\star_{\mathbf{M}}$-SVD is symmetric too, so only its top half must be stored. Bilaterally symmetric images, turned on their side and stored as lateral slices, give such a tensor. This construction keeps the symmetric part of each image and stores only half the basis. We then recognize new images by projecting them onto this basis and matching to the nearest training image in coefficient space. Across ten datasets of faces, leaves, butterfly wings, and other objects, the symmetric basis matches the recognition rate of ordinary tensor SVD at 2-11x less basis storage, with the exception of MUCT. On faces under varying illumination it exceeds the best rate the ordinary tensor SVD attains at any storage level.

Comments16 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑