发表机构
Sungkyunkwan University(成均馆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对社会工程欺诈的场景级分布外检测问题,提出ECoG生成式框架,结合证据跨度监督与理由-标签一致性,提升了分布外挑战性实例的检测性能。
AI 中文摘要
传统的分布内评估在训练数据和测试数据共享重复的任务特定模式或表面线索时,可能会高估模型的鲁棒性,这种风险在社会工程欺诈检测中尤为突出,攻击者可在保留恶意意图的同时改变场景、冒充实体或措辞。我们将此问题视为短信和语音钓鱼的场景级分布外(SL-OOD)检测问题,即从训练中保留整个攻击场景,同时保持标签空间固定,该设置用于测试模型是否能利用与决策相关的证据而非熟悉的场景特定线索泛化到未见过的攻击场景。通过这种SL-OOD评估,我们发现,对于基于特征、编码器和解码器的基线模型,高分布内性能无法可靠地预测保留的鲁棒性,我们将这种差距解释为场景记忆:即依赖重复的场景特定词汇或实体线索而非与决策相关的证据。我们提出ECoG,一种证据一致的生成式框架,在训练期间结合证据跨度监督和理由-标签一致性目标。在0.5B解码器上,相对于未使用一致性正则化训练的相同骨干网络,ECoG将分布外挑战性实例的Macro-F1提高了3.22个百分点,将生成的理由支持相反标签的预测占比降低了4.22个百分点,将与参考证据跨度的标记级重叠提高了8.38个百分点;在四个解码器骨干网络上,预测-理由不一致性的降低具有一致性。这些结果表明,紧凑的生成式检测器在社会工程偏移下可从证据监督和理由-标签一致性中受益。
英文摘要
Conventional in-distribution evaluation can overestimate robustness when training and test data share recurring task-specific patterns or surface cues. This risk is especially relevant in social-engineering fraud detection, where attackers can preserve malicious intent while changing the scenario, impersonated entity, or wording. We study this problem as scenario-level out-of-distribution (SL-OOD) detection for SMS and voice phishing, where entire attack scenarios are held out from training while the label space remains fixed. This setting tests whether models can generalize to unseen attack scenarios using decision-relevant evidence rather than familiar scenario-specific cues. Using this SL-OOD evaluation, we find that high in-distribution performance does not reliably predict held-out robustness across feature-, encoder-, and decoder-based baselines. We interpret this gap as scenario memorization: reliance on recurring scenario-specific lexical or entity cues rather than decision-relevant evidence. We propose ECoG, an evidence-consistent generative framework that combines evidence-span supervision with a rationale-label consistency objective during training. On the 0.5B decoder, relative to the same backbone trained without consistency regularization, ECoG raises Macro-F1 on OOD challenging instances by 3.22 points, reduces the share of predictions whose generated rationale supports the opposite label by 4.22 points, and increases token-level overlap with reference evidence spans by 8.38 points; the reduction in prediction-rationale inconsistency is consistent across four decoder backbones. These results suggest that compact generative detectors can benefit from evidence supervision and rationale-label consistency under social-engineering shift.
CommentsAccepted at CIKM 2026 (35th ACM International Conference on Information and Knowledge Management), Rome, Italy, November 2026. 12 pages, 4 figures. Code and data: https://github.com/kimsan1120/ECoG