发表机构
Chittagong University of Engineering and Technology(吉大港工程技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本到图像模型发展带来的合成图像源归属问题,提出融合语义深度学习与数学取证特征提取的双分支集成框架,经实验验证准确率高且计算高效,突出数学取证在实际部署中的优势。
AI 中文摘要
文本到图像(T2I)模型的快速发展需要稳健的合成图像源归属(SIA)方法。SIA的一个关键挑战是原始训练图像与经过未知后处理操作(如图像压缩和模糊)的实际部署图像之间的分布差异。在这项为ICANN 2026的DLMMDD挑战赛提出工作中,引入一种融合语义深度学习与数学取证特征提取的双分支集成框架。语义分支采用通过指数移动平均(EMA)和标签平滑正则化的EfficientNet-B0。取证分支从高通噪声残差中提取126个数学特征,通过截断奇异值分解(Truncated SVD)压缩并使用XGBoost分类。在一个有10个生成器的数据集上评估,测试集的55%被降级,该方法在私有排行榜上的准确率达到95.60%。此外,整个管道计算效率高,无需GPU加速,在标准CPU上不到6.5小时就能端到端执行,突出了数学取证在实际部署中的实用性和可扩展性。
英文摘要
The rapid advancement of text-to-image (T2I) models has necessitated robust Synthetic Image Source Attribution (SIA) methodologies. A critical challenge in SIA is the distribution shift between pristine training images and real-world deployed images, which undergo unknown post-processing operations such as JPEG compression and blurring. In this work, proposed for the DLMMDD Challenge at ICANN 2026, we introduce a dual-branch ensemble framework fusing Semantic Deep Learning with Mathematical Forensic Feature Extraction. The semantic branch employs EfficientNet-B0 regularized with Exponential Moving Averaging (EMA) and Label Smoothing. The forensic branch extracts 126 mathematical features -- including SVD spectral profiles and Local Binary Patterns -- from high-pass noise residuals, compressed via Truncated SVD and classified with XGBoost. Evaluated on a dataset of 10 generators where 55% of the test set is degraded, our approach achieves a private leaderboard accuracy of 95.60%. Furthermore, the entire pipeline is highly computationally efficient, requiring no GPU acceleration and executing end-to-end on a standard CPU in under 6.5 hours, highlighting the practicality and scalability of mathematical forensics for real-world deployment.
CommentsPeer-reviewed and accepted to the DLMMDD Challenge Workshop at the 35th International Conference on Artificial Neural Networks (ICANN 2026). 7 pages, 1 figures