arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08581cs.LGq-fin.CPstat.ML

AlphaRJM:用于随机回报引导的Alpha发现的奖励跳跃记忆

AlphaRJM: Reward-Jump Memory for Stochastic Return-Guided Alpha Discovery

Sayan Dhan, Selvaraju Natarajan

首次发表
浏览论文内容

中文总结 AI 辅助

AlphaRJM通过奖励跳跃记忆和随机回报评论家,解决公式化Alpha发现中延迟反馈导致的池历史缺失与中间动作价值不确定问题,实现稳定收益提升。

中文摘要 AI 辅助

公式化Alpha发现是一个依赖池的符号搜索问题,其中信息丰富的反馈主要在完整表达式被评估时观察到。这种延迟反馈产生了两个相互关联的困难:保留的Alpha池并未保存已实现评估反馈的完整历史,且中间构建动作的价值是不确定的,因为其结果取决于最终完成的公式。我们引入了AlphaRJM,通过奖励跳跃记忆(Reward-Jump Memory)解决这些困难,这是一种事件驱动的潜在状态,在令牌构建期间保持不变,仅在终端评估事件时使用已实现的池奖励和评估结果进行更新,以及一个动作条件的SDE回报评论家,用随机粒子表示未来的折扣发现回报。粒子通过其均值和不确定性指导动作选择,并使用结合能量距离匹配、均值校准和跳跃正则化的分布式贝尔曼目标进行学习。实验上,AlphaRJM在多个股票域、预测周期和随机种子中带来了强劲且稳定的收益,而消融研究证实了持久评估历史、随机回报建模和分布式监督的互补作用。

英文摘要

Formulaic alpha discovery is a pool-dependent symbolic search problem in which informative feedback is observed primarily when a complete expression is evaluated. This delayed feedback creates two coupled difficulties: the retained alpha pool does not preserve the full history of realized evaluation feedback, and the value of an intermediate construction action is uncertain because its consequence depends on the formula eventually completed. We introduce AlphaRJM, which addresses these difficulties through Reward-Jump Memory, an event-driven latent state that remains fixed during token construction and updates only at terminal evaluation events using the realized pool reward and evaluation outcome, and an action-conditioned SDE return critic that represents future discounted discovery returns with stochastic particles. The particles guide action selection through their mean and uncertainty and are learned using a distributional Bellman objective combining energy-distance matching, mean calibration, and jump regularization. Empirically, AlphaRJM delivers strong and stable gains across multiple equity universes, forecasting horizons, and random seeds, while ablations confirm the complementary roles of persistent evaluation history, stochastic return modeling, and distributional supervision.

发表机构

  • Indian Institute of Technology Guwahati(印度古瓦哈蒂理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑