发表机构
Innovatrics(因诺维特克斯(Innovatrics))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对亚千字节人脸压缩需求,基准测试了10种编解码器,提出定制学习型编解码器,发现512字节预算下编解码器排名会变化,为身份保留人脸压缩的实际部署提供依据。
AI 中文摘要
在身份证、智能卡生物识别和带宽受限的验证场景中,需要将人脸图像存储在严格的亚千字节预算下,这迫使编解码器丢弃大部分信号,同时保留人脸匹配器实际读取的内容——身份信息。通用编解码器优化的是像素保真度,而非驱动验证的嵌入距离,因此哪种编解码器、分辨率和设置能在1024字节或更低预算下最佳保留身份,以及在512字节时性能如何下降,目前尚不明确。我们在不同分辨率、字节预算、两个数据集(受控的Color FERET、野外的AI-Solutions-KK)及四个基准人脸匹配器上对10种通用和人脸专用编解码器进行了基准测试,14种模型(ViT和CNN)的列表证实排名与骨干网络无关。随后,我们训练了一种定制的身份保留编解码器,该编解码器通过对冻结增益表进行二分搜索,精确达到字节预算,并开展了四项研究:分辨率、人口统计学公平性、重压缩及无框对抗鲁棒性。亚千字节级的身份保留是可行的,但部署哪种编解码器完全取决于预算。在1024字节和112像素的工作分辨率下,该问题接近解决:现代编解码器在ArcFace基准上使Color FERET的等错误率低于0.35%。在512字节时,领域排名重新洗牌:AVIF、HEIF、JPEG XL及传统JPEG在FMR为1e-4时的非错误匹配率降至28%至98%,而WebP、JPEG-AI及我们的字节预算学习型编解码器则保持在该区间之外,其中WebP为24.3%,我们的精确变体在野外场景中为6.9%。这种重新洗牌,而非1024字节时的排名,是实际应用的结果:在1千字节时选择的编解码器,并非在其一半预算下部署的编解码器。
英文摘要
Storing face images under a hard sub-kilobyte budget, as required for identity documents, smart-card biometrics and bandwidth-constrained verification, forces a codec to discard most of the signal while keeping what a face matcher actually reads: identity. Generic codecs optimize pixel fidelity, not the embedding distances that drive verification, so which codec, resolution and setting best preserve identity at 1024 bytes or less, and how that degrades at 512, is unclear. We benchmark ten general and face-specific codecs across resolutions, byte budgets, two datasets (controlled Color FERET, in-the-wild AI-Solutions-KK) and four anchor face matchers, with a fourteen-model ViT and CNN roster confirming the ranking is backbone-invariant. We then train a custom identity-preserving codec that hits the byte budget exactly via binary search over a frozen gain table, and run four studies: resolution, demographic fairness, recompression, and no-box adversarial robustness. Sub-kilobyte identity preservation is feasible, but which codec to deploy depends entirely on the budget. At 1024 bytes and the 112 px working resolution the problem is close to solved: modern codecs hold Color FERET equal-error rate under 0.35 percent on the ArcFace anchor. At 512 bytes the field re-sorts: AVIF, HEIF, JPEG XL and legacy JPEG collapse to 28 to 98 percent false-non-match rate at FMR 1e-4, while WebP, JPEG-AI and our byte-budgeted learned codecs stay out of that band, with 24.3 percent for WebP against 6.9 percent for our accurate variant in the wild. That re-sort, not the 1024-byte ranking, is the operational result: a codec chosen at 1 kB is not the codec to deploy at half that.