arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EvolveTrade:面向自进化LLM交易代理的经验驱动策略精炼

EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents

Sehee Kim, Yumin Choi, Minki Kang, Sung Ju Hwang

arXiv 2609.17632首次发表:更新:

发表机构

KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EvolveTrade提出自进化框架,通过策略代理基于决策轨迹和投资组合反馈迭代优化交易代理的系统提示策略,在多个市场机制下提升夏普比率和累计收益,增强工具使用适应性。

AI 中文摘要

大型语言模型(LLM)交易代理能够结合市场数据、新闻和可执行分析,但其行为往往受部署前固定的静态手写工具使用策略控制。这限制了它们在不断变化的市场环境中适应信息收集、工具调用、信号验证和风险管理方式的能力。我们提出EvolveTrade,一个自进化框架,将使用工具的交易代理的系统提示视为文本参数化策略。在每个更新间隔后,策略代理利用累积的决策轨迹和已实现的投资组合反馈修订该策略,同时保持骨干LLM不变。更新后的策略随后用于下一批交易决策,使代理能够随时间优化其信息获取和投资组合构建流程。跨多个市场机制和两个LLM骨干的实验表明,EvolveTrade在大多数评估设置中相较于固定策略的LLM基线,通常能提高夏普比率(Sharpe Ratio)和累计收益(Cumulative Return),在多数评估设置中实现了改进的SR和CR。行为分析进一步表明,自进化策略增加了代码介导的分析并激活了与机制相关的计算;案例级策略到收益的归因追踪了策略引发的配置变化如何促成已实现收益差异。这些结果表明,适应控制工具使用的可复用程序是构建更稳健的LLM交易代理的关键方向。

英文摘要

Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment. This limits their ability to adapt how they gather evidence, invoke tools, verify signals, and manage risk under changing market regimes. We introduce EvolveTrade, a self-evolving framework that treats the system prompt of a tool-using trading agent as a text-parameterized policy. After each update interval, a Policy Agent revises this policy using accumulated decision traces and realized portfolio feedback, while keeping the backbone LLM fixed. The updated policy is then used for the next batch of trading decisions, enabling the agent to refine its information-acquisition and portfolio-construction procedure over time. Experiments across multiple market regimes and two LLM backbones show that EvolveTrade often improves Sharpe Ratio and Cumulative Return over fixed-policy LLM baselines, achieving the improved SR and CR in most evaluated settings. Behavioral analyses further show that self-evolved policies increase code-mediated analysis and activate regime-relevant computations; case-level policy-to-return attributions trace how policy-induced allocation changes contribute to realized return differences. These results suggest that adapting the reusable procedure governing tool use is a key direction for building more robust LLM trading agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑