arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15841cs.LGq-fin.CPstat.ML

用于股票交易稳定强化学习的自监督辅助任务发现

Self-Supervised Auxiliary Task Discovery for Stable Reinforcement Learning in Stock Trading

Arishi Orra, Himanshu Choudhary, Manoj Thakur

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出自监督框架自动发现辅助任务以支持股票交易的强化学习,通过广义价值函数和双网络结构优化,在四大股指上验证其较基线方法更具稳健性与交易性能。

中文摘要 AI 辅助

强化学习作为一种数据驱动的股票交易方法已受到越来越多的关注。然而,由于市场行为的非平稳性和奖励信号的噪声,学习一个既盈利又稳定的策略仍然具有挑战性。辅助任务常被用于改进表示学习并稳定训练,但它们通常是手动设计的,且高度依赖于对目标和预测范围的先验假设,这类固定设计可能无法在不断变化的市场 regime 中保持适用性。在本研究中,我们提出了一种自监督框架,该框架可自动发现辅助任务以支持股票交易的强化学习。这些辅助任务被公式化为广义价值函数(General Value Functions),使其预测结果能丰富所学的状态表示并辅助策略优化。该框架由两个网络组成:主网络与辅助预测一同学习交易策略,而次级网络则通过所学的累积量(cumulants)和折扣因子生成辅助任务的定义。这些任务使用元梯度机制进行更新,该机制考虑了它们对交易性能的长期影响并提升训练稳定性。我们在四个主要股票指数:道琼斯工业平均指数(DJI)、富时指数(FTSE)、孟买敏感指数(Sensex)和台湾加权股价指数(TAIEX)上对所提方法进行了评估。实证结果表明,与现有基线方法相比,自动发现的辅助任务可实现更稳健的学习并提升交易性能。

英文摘要

Reinforcement learning has gained increasing attention as a data-driven approach for stock trading. However, learning a policy that is both profitable and stable remains challenging due to non-stationary market behaviour and noisy reward signals. Auxiliary tasks are often used to improve representation learning and stabilize training, yet they are usually designed manually and depend heavily on prior assumptions about targets and prediction horizons. Such fixed designs may not remain suitable across changing market regimes. In this work, we propose a self-supervised framework that automatically discovers auxiliary tasks to support reinforcement learning for stock trading. The auxiliary tasks are formulated as General Value Functions so that their predictions enrich the learned state representation and assist policy optimization. The framework consists of two networks. The main network learns the trading policy along with the auxiliary predictions, while the secondary network generates the definitions of auxiliary tasks through learned cumulants and discount factors. These tasks are updated using a meta gradient mechanism that accounts for their long-term impact on trading performance and improves training stability. We evaluate the proposed approach across four major equity indices: DJI, FTSE, Sensex, and TAIEX. The empirical results demonstrate that automatically discovered auxiliary tasks lead to more robust learning and improved trading performance compared to existing baselines.

发表机构

  • School of Mathematical and Statistical Sciences, Indian Institute of Technology Mandi(印度理工大学曼迪分校数学与统计科学学院)

机构由 AI 辅助整理,请以论文原文为准。

↑