arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IDSPACE:用于可靠评估数字身份验证系统的新型文档生成器【扩展技术报告】

IDSPACE: A Novel Document Generator for Reliable Evaluation of Digital Identity Verification Systems [Extended Technical Report]

Lulu Xie, Yancheng Wang, Kanchan Chowdhury, Rolando Garcia, Yingzhen Yang, Jia Zou

arXiv 2609.03052首次发表:更新:

发表机构

Arizona State University; Marquette University(亚利桑那州立大学; 马凯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

IDSpace是一种新型文档生成器,通过模型引导的贝叶斯优化等技术生成合成身份文档,可提升数字身份验证系统的评估一致性等指标,还发布了含10种欧洲ID类型的359240份合成文档数据集。

AI 中文摘要

随着服务向线上迁移,银行、贷款机构、政府等信任机构必须对远程用户的身份进行验证。欺诈检测工具已广泛应用,但评估和微调这些工具仍存在困难,因为身份文档具有敏感性,因此数据稀缺。合成数据生成提供了可行路径,且需求明确:我们此前在该领域的工作已被下载超过11000次(来自8个部分的汇总数据)。本文引入IDSpace,在三个方向扩展了该领域的研究:第一,我们提出模型引导的贝叶斯优化,仅使用目标域的少量样本,调整生成参数以最大化视觉相似度和与目标域模型的预测一致性;第二,我们将用户指定的元数据(人口统计信息、欺诈模式、采集设备)与自动调整的控制参数(字体样式、噪声水平、图像质量)解耦,允许用户在无需底层专业知识的情况下配置评估;第三,我们突破模板图像的限制,支持扫描件和移动设备拍摄的文档。实验表明,IDSpace在仅使用少量真实样本的情况下,相较于CycleGAN、扩散修复、无引导优化等基线方法,可将评估一致性提升15-45%,同时将训练准确率提升最高9%,与目标域的SSIM相似度提升10%。我们还发布了一个新数据集,包含359240份高质量合成文档,覆盖10种欧洲身份证件类型。

英文摘要

As services move online, trust institutions such as banks, lenders, and governments must verify the identity of remote users. Fraud detection tools are widely available, but evaluating and fine-tuning them remains difficult because identity documents are sensitive and therefore scarce. Synthetic data generation offers a path forward, and demand is clear: our prior work in this area has been downloaded over $11{,}000$ times (aggregated from eight parts). We introduce IDSpace, extending this line of research in three directions. First, we propose model-guided Bayesian optimization, which tunes generation parameters to maximize both visual similarity and prediction consistency with target-domain models given only a few samples from a target domain. Second, we decouple user-specified metadata (demographics, fraud patterns, capture device) from automatically tuned control parameters (font styles, noise levels, image quality), allowing users to configure evaluations without low-level expertise. Third, we expand beyond template images to support scanned and mobile-captured documents. Experiments show IDSpace improves evaluation consistency by $15-45\%$ over baselines including CycleGAN, diffusion inpainting, and non-guided optimization, using only a few real samples, while improving training accuracy by up to $9\%$ and SSIM similarity with the target domain by $10\%$. We also released a new dataset consisting of $359{,}240$ high-quality synthetic documents across ten European ID types.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑