发表机构
Guangdong Polytechnic Normal University(广东技术师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出MAD-Guard,在Qwen3-VL-8B上对比自回归生成与直接决策接口,发现直接头可大幅降低延迟和校准误差,并提升多任务归因性能。
AI 中文摘要
多模态基础模型何时应生成令牌,何时应直接输出决策?我们提出了MAD-Guard,一项针对封闭式多模态取证任务的输出决策接口的受控研究。一旦计算出多模态表示,对于输入复杂度高但输出熵低的封闭式取证决策,自回归生成是否必要?在匹配的Qwen3-VL-8B骨干网络、2,400个FakeClue训练样本以及华为昇腾910C NPU上的LoRA预算(r=16, α=32)下,我们评估了一系列决策接口(AR-SFT [生成] → Logit Slice → Binary Direct Head → +choice → +act → CLM-Head),并将延迟分解为骨干表示(53.12毫秒)、151,643路词汇投影(+85.04毫秒 → 138.16毫秒)和解码(+248.26毫秒 → 386.42毫秒)。在1对1二元监督(L_BCE)下,Binary Direct Head将延迟降低2.60倍至7.27倍(53.12毫秒),并将校准误差降低1.88倍(ECE = 0.0450 vs. 0.0845),但以-1.80%的准确率权衡(93.10% vs. 94.90%;0.9795 vs. 0.9871 ROC-AUC)为代价,源于放弃令牌先验。相对于AR-SFT的增益要么来自多任务归因和不确定性门控(+choice+act:96.44%准确率,0.9940 ROC-AUC,0.0187 ECE,53.71毫秒),要么来自分解的对比头(CLM-Head:96.55%二元和96.44%多任务准确率,0.0166 ECE,98.79% 7类归因,54.42毫秒),在不进行令牌解码的情况下保留语义先验。在来自五个基准的5,000张样本外图像上,我们的框架在合成、伪装和文档伪造方面表现出色(GenImage 96.44%,Chameleon 97.73%,Doc 91.84%),而在压缩人脸操纵上显示出明显边界(FF++ ROC-AUC = 0.5913)。
英文摘要
When should multimodal foundation models generate tokens, and when should they directly output a decision? We present MAD-Guard, a controlled study of output-decision interfaces for closed multimodal forensic tasks. Once a multimodal representation is computed, is autoregressive generation necessary for closed forensic decisions with high input complexity but low output entropy? Under a matched Qwen3-VL-8B backbone, 2,400 FakeClue training samples, and LoRA budget ($r=16, α=32$) on Huawei Ascend 910C NPUs, we evaluate a progression of decision interfaces (AR-SFT [generate] $\to$ Logit Slice $\to$ Binary Direct Head $\to$ +choice $\to$ +act $\to$ CLM-Head) and decompose latency into backbone representation (53.12 ms), 151,643-way vocabulary projection (+85.04 ms $\to$ 138.16 ms), and decoding (+248.26 ms $\to$ 386.42 ms). Under 1-to-1 binary supervision ($\mathcal{L}_{\mathrm{BCE}}$), a Binary Direct Head cuts latency by $2.60\times$-$7.27\times$ (53.12 ms) and lowers calibration error by $1.88\times$ (ECE = 0.0450 vs. 0.0845), with a -1.80% accuracy trade-off (93.10% vs. 94.90%; 0.9795 vs. 0.9871 ROC-AUC) from forfeiting token priors. Gains above AR-SFT arise either from multi-task attribution and uncertainty gating (+choice+act: 96.44% accuracy, 0.9940 ROC-AUC, 0.0187 ECE at 53.71 ms) or from a disaggregated contrastive head (CLM-Head: 96.55% binary and 96.44% multi-task accuracy, 0.0166 ECE, 98.79% 7-class attribution at 54.42 ms) retaining semantic priors without token decoding. Across 5,000 out-of-sample images from five benchmarks, our framework excels on synthetic, camouflage, and document forgeries (96.44% GenImage, 97.73% Chameleon, 91.84% Doc) while showing a clear boundary on compressed face manipulation (FF++ ROC-AUC = 0.5913).
Comments7 pages, 3 figures, 7 tables