arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于重心的对抗逆强化学习学习股票交易策略

Learning Stock Trading Policies via Barycenter-Based Adversarial Inverse Reinforcement Learning

Arishi Orra, Himanshu Choudhary, Manoj Thakur

arXiv 2608.15770首次发表:更新:

发表机构

School of Mathematical and Statistical Sciences; Indian Institue of Technology Mandi(数学与统计学院; 印度曼迪理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出基于重心的对抗逆强化学习框架BRaG,聚合多异构专家策略学习股票交易策略,在四大全球股市上表现优于经典交易规则及深度强化学习方法,风险特征更稳定。

AI 中文摘要

使用强化学习设计有效交易策略仍具挑战性,原因在于奖励存在延迟与噪声、探索效果差,且难以强制执行明确的风险约束。本研究提出BRaG——一种用于股票交易的基于重心的对抗逆强化学习框架,该框架从多种异构专家策略中学习交易行为。BRaG使用性能加权的Wasserstein重心聚合专家演示,生成稳定的伪专家表示,捕捉不同交易风格间的共享结构。该表示通过对抗式模仿学习预训练交易策略,缓解强化学习过程中不稳定的探索问题。预训练后的策略随后使用真实市场奖励进行强化学习优化。为确保风险感知决策,BRaG纳入控制障碍函数,约束动作执行并正则化策略学习以满足回撤限制。我们在四个全球主要股票市场(包括美国、英国、印度和台湾指数)上评估该方法。在所有市场中,该方法的表现优于经典交易规则和近期深度强化学习方法,同时展现出更稳定的风险特征。

英文摘要

Designing effective trading strategies using reinforcement learning remains challenging due to delayed and noisy rewards, poor exploration, and the difficulty of enforcing explicit risk constraints. In this work, we propose BRaG, a barycenter-based adversarial inverse reinforcement learning framework for stock trading that learns trading behavior from multiple heterogeneous expert strategies. BRaG aggregates expert demonstrations using a performance-weighted Wasserstein barycenter, yielding a stable pseudo-expert representation that captures shared structure across diverse trading styles. This representation is used to pretrain a trading policy via adversarial imitation learning, which alleviates unstable exploration during reinforcement learning. The pretrained policy is subsequently refined using reinforcement learning with true market rewards. To ensure risk-aware decision-making, BRaG incorporates control barrier functions that constrain action execution and regularize policy learning to satisfy drawdown limits. We evaluate the proposed approach on four major global equity markets, including the US, UK, Indian, and Taiwanese indices. Across all the markets, the proposed approach achieves stronger performance than both classical trading rules and recent deep reinforcement learning methods, while exhibiting more stable risk characteristics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑