arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18766cs.SDcs.CL

FRAUDSkill:面向音频反欺诈检测的结构化冻结权重技能优化

FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection

  • JD Technology(京东科技)
  • Meta
  • University of Science and Technology of China(中国科学技术大学)
  • Northeastern University(东北大学)
  • Peking University, Shenzhen(北京大学深圳研究生院)

机构由 AI 辅助整理,请以论文原文为准。

Chengxian Hu, Zhiming Ma, Mingjun Pan, Yifan Wang, Shun Zhang, Qifan Wang, Zhilei Zhao, Yijin Zhou, Yuxi Zhao, Huiyuan Liu, Peidong Wang, Peng Chen

AI总结:

针对音频反欺诈检测中模型难以适应演变的问题,提出FRAUDSkill框架,通过冻结权重并优化外部技能层,在TeleAntiFraud基准上实现73.50% Macro-F1,较基线提升31.96%。

AI中文摘要:

大型音频语言模型通过直接处理语音并推理与欺诈相关的证据,在反欺诈检测方面展现出潜力。然而,其部署要求预测遵循预定义的标签空间和结构化决策协议,该协议包括服务场景识别、欺诈检测和条件性欺诈类型分类。现有的微调和基于提示的方法通常将任务知识、约束和决策规则编码到模型参数或手动维护的提示中,这使得它们难以适应欺诈模式和标签策略的演变。为此,我们提出了FRAUDSkill,一种结构化的冻结权重适应框架,它保持底层音频语言模型不变,同时优化外部层的技能程序、路由特定策略和决策规则。我们进一步将结构化输出控制与验证引导的多路径推理相结合,以确保符合协议的预测。在TeleAntiFraud基准上,FRAUDSkill实现了73.50%的Macro-F1,比共享冻结模型基线高出31.96%,同时将无效输出减少到1.94%。大量实验表明,外部技能优化为结构化音频反欺诈检测提供了一种有效且可适应的解决方案,而无需修改底层模型。源代码可在该https URL获取。

英文摘要:

Large audio-language models have shown promise for anti-fraud detection by directly processing speech and reasoning over fraud-related evidence. Their deployment, however, requires predictions to follow a predefined label space and a structured decision protocol consisting of service-scenario identification, fraud detection, and conditional fraud-type classification. Existing fine-tuning and prompt-based approaches typically encode task knowledge, constraints, and decision rules into model parameters or manually maintained prompts, making them difficult to adapt as fraud patterns and labeling policies evolve. To this end, we propose FRAUDSkill, a structured frozen-weight adaptation framework that leaves the underlying audio-language model unchanged while optimizing an external layer of skill programs, route-specific policies, and decision rules. We further combine structured output control with validation-guided multi-path inference to ensure protocol-compliant predictions. On the TeleAntiFraud benchmark, FRAUDSkill achieves 73.50% Macro-F1, outperforming the shared frozen-model baseline by 31.96% while reducing invalid outputs to 1.94%. Extensive experiments demonstrate that external skill optimization provides an effective and adaptable solution for structured audio anti-fraud detection without modifying the underlying model. The source code is available at https://anonymous.4open.science/r/FRAUDSKILL-114514.

补充信息

↑