arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10407cs.AIcs.LGq-fin.PMq-fin.TR

SOTA:由期权隐含收益分布引导的股票期权交易智能体

SOTA: Stock Options Trading Agents Guided by Option-Implied Return Distributions

  • Carnegie Mellon University(卡内基梅隆大学)
  • Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

Yizhen Xie, Mengyang Liu

AI总结:

提出SOTA智能体交易框架,将期权策略选择抽象为策略层面决策,通过监督微调和强化学习训练Qwen模型,在美股期权上实现18.3%总回报和1.60夏普比率,并发现新闻在训练中的不对称作用。

AI中文摘要:

随着期权市场的增长和人工智能的进步,用于期权交易的智能体系统正受到越来越多的关注。基于语言模型的智能体能够对新闻等上下文信息进行推理,但期权交易提出了一个特别具有挑战性的决策问题:一只股票可能拥有数千份合约,智能体必须决定交易哪些合约以及如何组合它们。现有方法通常通过将策略限制为固定的策略结构(如跨式期权)来回避这种复杂性,从而限制了它们在市场条件变化时切换策略的能力。我们提出了SOTA(股票期权交易智能体),一个用于结构化期权策略选择的智能体交易框架。SOTA将庞大的期权空间抽象为策略层面的决策,而确定性的解析器则负责投资组合的实施。我们通过对Qwen3.8-27B进行监督微调后接强化学习的后训练来开发SOTA。SOTA在九只美国大盘股和SPY的期权上,与基于规则和机器学习的策略选择器在同一交易环境中进行评估。在六个月的样本外期间,SOTA获得了18.3%的总回报,夏普比率为1.60,最大回撤为8.96%。我们还记录了新闻的不对称作用:新闻改善了前沿教师轨迹,但在强化学习期间保留新闻会将样本外回报从18.3%降至-2.7%。

英文摘要:

As option markets grow and AI advances, agentic systems for option trading are gaining increasing attention. Language-model-based agents can reason over contextual information such as news, but option trading presents a particularly challenging decision problem: a single stock can have thousands of contracts, and the agent must decide both which contracts to trade and how to combine them. Existing approaches often sidestep this complexity by restricting the policy to a fixed strategy structure, such as a straddle, limiting their ability to switch strategies as market conditions change. We present SOTA (Stock Options Trading Agents), an agentic trading framework for structured option-strategy selection. SOTA abstracts the large option universe into strategy-level decisions while deterministic resolvers handle portfolio implementation. We develop SOTA by post-training Qwen3.8-27B with supervised fine-tuning followed by reinforcement learning. SOTA is evaluated on options on nine large-cap U.S. equities and SPY against rule-based and machine-learning strategy selectors in the same trading environment. Over a six-month out-of-sample period, SOTA earns an 18.3% total return with a Sharpe ratio of 1.60 and a maximum drawdown of 8.96%. We also document an asymmetric role of news: news improves frontier-teacher trajectories, but retaining news during reinforcement learning reduces out-of-sample return from 18.3% to -2.7%.

补充信息

↑