arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34188stat.MLcs.AIcs.LG

AlphaPareto:基于LLM引导的多目标强化学习的公式化Alpha发现

AlphaPareto: Formulaic Alpha Discovery with LLM-Guided Multi-Objective Reinforcement Learning

Yingbo Zhao, Zeyu Yang, Zhoufan Zhu

AI总结:

提出AlphaPareto,一种基于LLM引导的多目标强化学习方法,通过状态增强和帕累托正则化解决非平稳性和多目标优化问题,在真实数据集上优于现有方法。

AI中文摘要:

公式化Alpha发现是量化交易中的核心挑战,因为识别能够良好协同工作的Alpha仍然困难。最近的强化学习(RL)方法将此任务建模为马尔可夫决策过程(MDP),但有两个重要问题仍未解决。首先,随着Alpha池的演变,奖励函数相应变化,使得MDP本质上非平稳。其次,大多数现有方法优化单一目标,通常是预测能力,而忽略了高质量Alpha池的其他重要属性。受这些挑战的启发,我们提出了AlphaPareto,一种用于公式化Alpha发现的RL方法。为解决非平稳性,AlphaPareto将状态增强为包括正在构建的Alpha和当前Alpha池,并应用大语言模型(LLM)对池进行编码。这种设计使智能体能够适应不断变化的搜索环境。为克服单目标奖励设计的局限性,AlphaPareto用多目标向量值奖励取代标量奖励,同时捕获预测能力、时间稳定性、扰动鲁棒性和多样性,并通过帕累托正则化学习过程优化这些目标。对真实数据集的实证应用表明,我们的AlphaPareto方法优于其竞争对手。

英文摘要:

Formulaic alpha discovery is a core challenge in quantitative trading, as identifying alphas that work well together remains difficult. Recent reinforcement learning (RL) methods formulate this task as a Markov decision process (MDP), but two important issues remain unresolved. First, as the alpha pool evolves, the reward function changes accordingly, making the MDP inherently non-stationary. Second, most existing methods optimize a single objective, typically predictive power, while ignoring other important properties of a high-quality alpha pool. Motivated by these challenges, we propose AlphaPareto, an RL method for formulaic alpha discovery. To address non-stationarity, AlphaPareto augments the state to include both the alpha under construction and the current alpha pool, and applies a large language model (LLM) to encode the pool. This design allows the agent to adapt to the evolving search environment. To overcome the limitation of single-objective reward design, AlphaPareto replaces the scalar reward with a multi-objective vector-valued reward that simultaneously captures predictive power, temporal stability, perturbation robustness, and diversity, and optimizes these objectives through a Pareto-regularized learning procedure. Empirical applications to real-world datasets show that our AlphaPareto method outperforms its competitors.

补充信息

↑