多视角与混淆引导的集成框架用于鲁棒合成图像溯源
A Multi-View and Confusion-Guided Ensemble Framework for Robust Synthetic Image Attribution
- State Key Laboratory of HVDC(直流输电国家重点实验室)
- China Southern Power Grid Electric Power Research Institute(中国南方电网电力科学研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出多视角与混淆引导的集成框架,融合多种架构和增强策略,解决合成图像溯源中模型相似和后处理干扰问题,在公开和私有排行榜上分别取得99.53%和99.20%的准确率。
AI中文摘要:
随着文本到图像生成模型的快速发展,合成图像溯源(SIA)变得越来越重要。然而,由于现代基于扩散的生成器之间的相似性日益增加以及存在各种后处理操作,准确识别生成图像的来源模型仍然具有挑战性。在本报告中,我们针对ICANN 2026 DLMMDD研讨会的合成图像溯源挑战赛,提出了一种多视角与混淆引导的集成框架。我们的方法整合了多种互补架构,包括FFT-ConvNeXt、DINOv2、CLIP和Xception,从频率、语义和取证角度捕获多样化的溯源线索。为了提高对未知降级和图像操纵的鲁棒性,我们在训练过程中采用了广泛的数据增强策略,模拟压缩、调整大小、灰度转换和模糊等真实后处理操作。此外,我们分析了集成模型的混淆模式,并观察到Stable Diffusion 3和Stable Diffusion 3.5之间存在严重的模糊性。为了解决这个问题,我们引入了一个专门的二元专家分类器,在低置信度条件下选择性激活。我们还应用了类自适应置信度校准,以提高对腾讯混元等具有挑战性类别的判别能力。所提出的框架在公共排行榜上达到了99.53%的准确率,在私有排行榜上达到了99.20%的准确率。源代码和实现细节可在该https URL公开获取。
英文摘要:
Synthetic image attribution (SIA) has become increasingly important with the rapid advancement of text-to-image generation models. However, accurately identifying the source model of a generated image remains challenging due to the growing similarity among modern diffusion-based generators and the presence of diverse post-processing operations. In this report, we present a multi-view and confusion-guided ensemble framework for the Synthetic Image Attribution Challenge of the DLMMDD Workshop at ICANN 2026. Our approach integrates multiple complementary architectures, including FFT-ConvNeXt, DINOv2, CLIP, and Xception, to capture diverse attribution cues from frequency, semantic, and forensic perspectives. To improve robustness against unknown degradations and image manipulations, extensive data augmentation strategies are employed during training, simulating realistic post-processing operations such as compression, resizing, grayscale conversion, and blur. Furthermore, we analyze the confusion patterns of the ensemble model and observe severe ambiguity between Stable Diffusion 3 and Stable Diffusion 3.5. To address this issue, we introduce a dedicated binary expert classifier that is selectively activated under low-confidence conditions. We additionally apply class-adaptive confidence calibration to improve the discrimination of challenging classes such as Tencent Hunyuan. The proposed framework achieved 99.53% on the public leaderboard and 99.20% on the private leaderboard. The source code and implementation details are publicly available at https://github.com/ZOMIN28/SIA.