ControlTrace:恢复控制场以实现隐藏内容识别
ControlTrace: Recovering Control Fields for Hidden-Content Recognition
- Shanghai Jiao Tong University(上海交通大学)
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对扩散模型隐藏内容难以被VLM识别的问题,提出ControlTrace方法,通过U-Net恢复灰度控制场,显著提升轮廓识别准确率,且开销极小。
AI中文摘要:
空间条件扩散模型可以将单词和轮廓嵌入到自然图像中,但视觉语言模型(VLM)可能无法识别隐藏内容。基于变换的恢复方法依赖于参数和视图的选择。为了评估隐藏内容的恢复和识别,我们构建了FreqBlind,这是一个包含6,000张图像的基准,涵盖轮廓、真实单词和非单词,跨越三种条件强度。评估的基于变换的方法在轮廓图案和弱条件隐藏内容的识别上表现有限。为解决这一局限,我们提出ControlTrace来恢复生成过程中使用的灰度控制场。一个8.4M参数的U-Net从载体图像预测该场,然后VLM识别其内容。使用Qwen2.5-VL-7B-Instruct,ControlTrace在三种条件强度下实现了60.2%的开放式轮廓识别准确率,超过三种评估的先前方法中最佳方法26.9个百分点。在A100 GPU上,完整流程相比直接VLM推理仅增加7.4毫秒(5.3%)的延迟。恢复的场相比评估的变换视图具有更低的像素误差和更高的结构相似性。在四种评估的VLM中,ControlTrace保持了其整体轮廓识别优势。在测试的JPEG压缩、高斯噪声和下采样下,识别保持稳定。这些结果支持在评估设置中控制场恢复用于隐藏内容识别。
英文摘要:
Spatially conditioned diffusion models can embed words and contours in natural-looking images, but vision-language models (VLMs) may fail to recognize the hidden content. Transformation-based recovery depends on parameter and view selection. To evaluate hidden-content recovery and recognition, we construct FreqBlind, a 6,000-image benchmark spanning contours, real words and non-words across three conditioning strengths. The evaluated transformation-based methods show limited recognition of contour patterns and weakly conditioned hidden content. To address this limitation, we propose ControlTrace to recover the grayscale control field used during generation. An 8.4M-parameter U-Net predicts this field from the carrier image, and a VLM then identifies its content. With Qwen2.5-VL-7B-Instruct, ControlTrace achieves 60.2% open-ended contour recognition accuracy across the three conditioning strengths, exceeding the best of the three evaluated prior methods by 26.9 percentage points. On an A100 GPU, the complete pipeline adds only 7.4 ms (5.3%) to direct VLM inference. Recovered fields have lower pixel errors and higher structural similarity than the evaluated transformation views. Across four evaluated VLMs, ControlTrace retains its overall contour recognition advantage. Recognition remains stable under the tested JPEG compression, Gaussian noise and downsampling. These results support control-field recovery for hidden-content recognition in the evaluated setting.