发表机构
X-Tech; Monash University; PandaAI; IIIS, Tsinghua University(X-Tech; 莫纳什大学; 熊猫AI; 清华大学智能产业研究院(IIIS))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AlphaSchema构建并探索结构化交易语义空间,通过解耦探索与实施、结合代理模型平衡探索与利用,在中国股市挖掘出强预测与投资组合表现的因子池,且对LLM选择具鲁棒性。
AI 中文摘要
自动阿尔法挖掘越来越多地采用大语言模型(LLM)智能体进行因子生成和迭代发现。然而,现有的基于LLM的系统通常将因子构建和搜索决策都委托给智能体本身,没有明确的探索空间或用于导航该空间的原则性机制。因此,探索在很大程度上是隐含的,难以系统地控制或优化。我们引入AlphaSchema,它构建并探索用于阿尔法挖掘的结构化交易语义空间。该空间中的每个点都是一个由事件(Event)、上下文(Context)、属性(Qualities)、方向(Direction)和输出(Output)组成的模式计划,在实施前指定候选因子的语义。AlphaSchema将探索与实施解耦:LLM将选定的模式计划转换为可执行因子,同时累积评估奖励以学习语义空间上的代理模型。迭代选择机制使用该模型平衡全局探索、代理引导的利用和局部变异。在中国股票市场上的实验表明,AlphaSchema发现的因子池具有强大的预测和投资组合表现。进一步分析显示,语义搜索过程在多样化区域导航的同时,越来越多地将评估分配给高奖励区域,并且不同LLM对相同模式计划的实施表现出相当的预测质量,这表明在我们的框架内,阿尔法挖掘质量对LLM的选择具有很大的鲁棒性。
英文摘要
Automated alpha mining has increasingly adopted large language model (LLM) agents for factor generation and iterative discovery. However, existing LLM-based systems often delegate both factor construction and search decisions to the agent itself, without an explicit exploration space or a principled mechanism for navigating that space. As a result, exploration remains largely implicit and difficult to control or optimize systematically. We introduce AlphaSchema, which constructs and explores a structured space of trading semantics for alpha mining. Each point in this space is a schema plan composed of Event, Context, Qualities, Direction, and Output, specifying the semantics of a candidate factor before implementation. AlphaSchema decouples exploration from implementation: an LLM translates selected schema plans into executable factors, while evaluated rewards are accumulated to learn a surrogate model over the semantic space. An iterative selection mechanism uses this model to balance global exploration, surrogate-guided exploitation, and local mutation. Experiments on the Chinese stock market show that AlphaSchema discovers factor pools with strong predictive and portfolio performance. Further analyses show that the semantic search process navigates diverse regions while increasingly allocating evaluations toward high-reward regions, and that implementations of the same schema plans by different LLMs exhibit comparable predictive quality, suggesting that alpha mining quality is largely robust to the choice of LLM within our framework.
CommentsInitial Version