arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语义帕累托深度Q网络:一种用于金融异常检测的多目标强化学习框架

Semantic Pareto-DQN: A Multi-Objective Reinforcement Learning Framework for Financial Anomaly Detection

Cláudio Lúcio do Val Lopes, Lucca Machado da Silva

arXiv 2607.09641首次发表:更新:

发表机构

A3Data(A3数据公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对金融异常检测的类不平衡问题,提出语义帕累托深度Q网络框架,合成交易特征,优化向量奖励,映射帕累托前沿,经实验验证可打破零召回陷阱,提升少数类召回率,为金融异常发现提供新途径。

AI 中文摘要

金融异常检测面临极端类不平衡问题,导致传统单目标算法出现“欺诈崩溃”,默认选择多数类,无法平衡异常拦截与客户摩擦。为避免扭曲数据重采样来克服此问题,我们提出语义帕累托深度Q网络(Semantic Pareto-DQN)这一多目标强化学习框架。该方法将异构交易特征合成连贯的自然语言叙述,由大语言模型编码,产生稳健、尺度不变的状态表示。智能体优化明确解耦金融效能、操作摩擦和语义发现的向量奖励。通过映射连续帕累托前沿,系统动态应对错过异常与误报的不对称成本。在电子商务欺诈和UCI信用数据集上的实证评估表明,语义帕累托深度Q网络成功打破零召回陷阱,与标量化基线相比,实现了更高的少数类召回率,为金融异常发现提供了一种替代方案,以换取有限的操作摩擦。

英文摘要

Financial anomaly detection suffers from extreme class imbalance, causing traditional single-objective algorithms to exhibit ``fraud collapse'', defaulting to the majority class and failing to balance anomaly interdiction with customer friction. To overcome this without distortive data resampling, we propose the Semantic Pareto-DQN, a multi-objective reinforcement learning framework. Our approach synthesizes heterogeneous transaction features into cohesive natural-language narratives, encoded by large language models, thereby producing a robust, scale-invariant state representation. The agent optimizes a vectorial reward that explicitly decouples financial efficacy, operational friction, and semantic discovery. By mapping the continuous Pareto frontier, the system dynamically navigates the asymmetric costs of missed anomalies versus false positives. Empirical evaluations across E-Commerce fraud and UCI Credit datasets show that semantic Pareto-DQN successfully shatters the zero-recall trap. It achieves superior minority-class recall compared to scalarized baselines, providing an alternative to trade bounded operational friction for financial anomaly discovery.

CommentsBRACIS 2026 - 36th Brazilian Conference on Intelligent Systems

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑