发表机构
Institute of Science Tokyo; Center for Advanced Intelligence Project (AIP), RIKEN; The Hong Kong University of Science and Technology (Guangzhou)(东京科学大学; 理化学研究所先进智能项目中心; 香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出受量子多体理论启发的等距多重线性基,通过黎曼优化训练,在自然照片和线条图压缩上优于传统固定基,在Quick Draw任务中同质量下字节数比JPEG少约20%。
AI 中文摘要
离散傅里叶变换、离散余弦变换及其分块变体是大多数已部署的图像和视频编解码器的基础,它们的有效性依赖三个特性:运行时间接近线性(线性至多对数因子)、完全可逆、参数极少。本研究将这些基推广为等距多重线性基,允许少量额外参数(参数规模为图像大小的多对数级),同时保留上述全部三个特性。针对图像数据集,我们开发了一套系统框架,在该基族中搜索能最有效压缩数据集的基:该基被参数化为受量子多体理论启发的等距张量网络,通过酉矩阵流形上的黎曼优化进行训练。在自然照片和线条图上,训练后的基始终优于固定的非参数化对应基;在Quick Draw线条图压缩任务中,相同重建质量下,它们存储图像所需的字节数比JPEG的8×8分块余弦变换少约20%。
英文摘要
The Discrete Fourier Transform (DFT), the Discrete Cosine Transform (DCT), and their block-wise variants underpin most deployed image and video codecs. Their effectiveness rests on three properties: their runtime is near-linear (up to a polylogarithmic factor) in the image size, they are exactly invertible, and they carry few to no parameters. In this work, we generalize these bases to isometric multilinear bases, allowing a small number of extra parameters (polylogarithmic in the image size), while preserving all three properties. We develop a scheme to train a better transformation for a given image dataset: we use isometric tensor networks, inspired by quantum many-body theory, to parameterize the basis, and train it with Riemannian optimization. We show that training consistently improves performance, as our parameterized bases can represent the traditional DFT and DCT-IV (a variant of the DCT). Evidence is shown across natural photographs and line drawings. On Quick Draw line-drawing compression, for example, the best trained basis outperforms the block cosine transform used in the JPEG format by $20\%$ in terms of compressed data size.