AlphaG-OPD:符号阿尔法因子发现中用于策略内蒸馏的可靠性门控兄弟反事实方法
AlphaG-OPD: Reliability-Gated Sibling Counterfactuals for On-Policy Distillation in Symbolic Alpha Factor Discovery
浏览论文内容
中文总结 AI 辅助
针对符号阿尔法因子发现中结构决策无直接标签的问题,提出AlphaG-OPD框架,通过三个决策模块将终端因子评估转化为局部动作指导,经多市场多种子测试展现出强劲跨市场表现。
中文摘要 AI 辅助
符号阿尔法因子发现能够对已完成的表达式进行评分,但不会为生成该表达式的结构决策提供直接标签。生成流网络(GFlowNets)可在完整表达式上维持与奖励成比例的多样分布,但其轨迹级目标无法比较中间状态下未被选择的兄弟动作。我们提出AlphaG-OPD,这是一种将终端因子评估转化为局部动作指导的结构化策略内蒸馏框架,其设计分为三个决策模块:模块I确定教学位置,即暴露当前前向策略访问的部分抽象语法树(AST)状态下语法有效的兄弟节点;模块II确定足够可靠的教学内容,即评估四个共享后缀下的三个支持兄弟节点,仅当它们的匹配比较表现出足够的胜者一致性和正经验下置信界(LCB)时,才接受KL有界目标;模块III确定教学的强度和时长,即通过有界重放、分数索引过期和前向梯度平衡整合已接受的目标,无需额外因子评估。终端奖励、轨迹平衡、后向策略、语法和因子池规则保持不变。我们采用等物理分数的四臂消融测试配对教学、可靠性门控和整合模块,在中国的CSI300、CSI500、CSI1000以及美国的S&P 500市场上,完整方法在多个随机种子下均展现出强劲的跨市场表现。
英文摘要
Symbolic alpha factor discovery can score a completed expression, but it provides no direct label for the structural decisions that produced it. Generative flow networks (GFlowNets) preserve a diverse, reward-proportional distribution over complete expressions, yet their trajectory-level objective does not compare unchosen sibling actions at an intermediate state. We introduce AlphaG-OPD, a structural on-policy distillation framework that turns terminal factor evaluations into local action guidance. Its design separates three decisions. Component I determines where to teach by exposing grammar-valid siblings at partial abstract-syntax-tree (AST) states visited by the current forward policy. Component II determines what is reliable enough to teach: it evaluates three supported siblings under four shared suffixes and admits a KL-bounded target only when their matched comparisons exhibit sufficient winner agreement and a positive empirical lower confidence bound (LCB). Component III determines how strongly and for how long to teach by consolidating accepted targets through bounded replay, score-indexed expiry, and forward-gradient balancing, without additional factor evaluations. Terminal reward, Trajectory Balance, the backward policy, grammar, and factor-pool rules remain unchanged. An equal-physical-score four-arm ablation tests paired teaching, reliability gating, and consolidation. Across China's CSI300, CSI500, and CSI1000 and the U.S. S&P 500, the complete method delivers strong cross-market performance over multiple random seeds.