发表机构
Concordia University(康考迪亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对CLEF 2026 FinMMEval任务3,将比特币和特斯拉交易决策问题建模为离散动作马尔可夫决策过程,比较四种深度强化学习算法,引入阿尔法奖励并优化超参数,DDPG在测试集表现最佳,揭示了验证到测试的泛化差距。
AI 中文摘要
本文介绍了我们针对CLEF 2026 FinMMEval实验室任务3的系统,该任务要求使用新闻和历史市场数据对比特币(BTC)和特斯拉(TSLA)进行每日做多、平仓或做空交易决策。我们将问题表述为离散动作马尔可夫决策过程,并比较了四种深度强化学习算法:策略梯度(PG)、近端策略优化(PPO)、深度Q学习(DQL)和深度确定性策略梯度(DDPG)。智能体使用技术指标、周期性日历编码和由LLaMA 3.2 1B生成的每日新闻情感得分。为减少过拟合并使训练与超越买入并持有目标保持一致,我们引入基于超额市场回报的阿尔法奖励并随机化情节起始日期。通过Ray Tune对每个算法-资产对进行180次试验来优化超参数,基于验证夏普比率进行早期停止和模型选择。在CLEF任务3测试集上,DDPG实现了最强的整体性能。DQL因其在验证中获得最高夏普比率而被预先选作实时端点,且选择过程未使用测试期数据。对于TSLA,DDPG和DQL的累积回报率分别为54.96%和52.62%,而买入并持有为16.45%。对于BTC,DDPG实现了1.58%的正回报,而买入并持有下降了-34.27%。结果还揭示了显著的验证到测试的泛化差距,凸显了将在牛市条件下选择的策略转移到熊市状态的困难。
英文摘要
This paper presents our system for Task 3 of the CLEF 2026 FinMMEval Lab, which requires daily long, flat, or short trading decisions for Bitcoin (BTC) and Tesla (TSLA) using news and historical market data. We formulate the problem as a discrete-action Markov Decision Process and compare four deep reinforcement learning algorithms: Policy Gradient (PG), Proximal Policy Optimization (PPO), Deep Q-Learning (DQL), and Deep Deterministic Policy Gradient (DDPG). The agents use technical indicators, cyclical calendar encodings, and daily news sentiment scores produced by LLaMA 3.2 1B. To reduce overfitting and align training with the objective of outperforming buy-and-hold, we introduce an alpha reward based on excess market return and randomize episode start dates. Hyperparameters are optimized with Ray Tune over 180 trials per algorithm-asset pair, with early stopping and model selection based on validation Sharpe ratio. On the CLEF Task 3 test set, DDPG achieves the strongest overall performance. DQL was selected a priori for the live endpoint because it obtained the highest validation Sharpe ratio, with selection performed without access to the test period. For TSLA, DDPG and DQL achieve cumulative returns of 54.96% and 52.62%, respectively, compared with 16.45% for buy-and-hold. For BTC, DDPG achieves a positive return of 1.58% while buy-and-hold declines by -34.27%. The results also reveal a substantial validation-to-test generalization gap, highlighting the difficulty of transferring policies selected in bull-market conditions to a bear-market regime.
CommentsAccepted to the FinMMEval Lab at CLEF 2026. Recipient of a Merit Award recognizing promising approaches and well-documented submissions