arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人工智能交易:评估用于技术市场分析的大语言模型

AI Trading: Evaluating Large Language Models for Technical Market Analysis

Geofrey Ntale

arXiv 2607.15414首次发表:更新:

AI 中文总结

本文系统比较评估五个大语言模型用于技术市场分析的能力,涵盖四项任务,采用多种定量指标。实验发现GPT-4 Turbo在通用模型中年化回报和夏普比率最高,FinGPT经领域微调表现出色,二者均超标准普尔500指数基准,还识别出模型存在的问题及得出相关结论。

AI 中文摘要

大语言模型已成为处理现代金融市场异构信息环境的强大工具。本文对五个著名的大语言模型:GPT-4 Turbo、Claude 3 Opus、Gemini 1.5 Pro、Llama 3 70B和领域专用的FinGPT进行了系统的比较评估,涉及它们的技术市场分析能力。评估涵盖四个结构化任务:从OHLCV数据中识别烛台模式、生成方向信号(买入/卖出/持有)、通过模拟执行管道对信号质量进行回测以及理解财务报告。实验框架采用了包括夏普比率、最大回撤、索提诺比率、信息系数、F1分数和BLEU分数等严格的定量指标。模拟回测结果表明,GPT-4 Turbo在通用模型中实现了最高的年化回报率和夏普比率,而FinGPT由于特定领域的微调而表现出具有竞争力的风险调整后性能。两个模型在测试条件下均优于被动的标准普尔500指数基准。该研究还识别了所有评估模型中持续存在的失败模式,包括数值幻觉、上下文窗口限制以及在横向市场环境中表现不一致。我们得出结论,虽然大语言模型在人工智能交易系统中具有真正的潜力,但强大的部署需要仔细的任务分解、严格的回测协议和领域感知的微调策略。

英文摘要

Large Language Models (LLMs) have emerged as powerful tools for processing the heterogeneous information environments of modern financial markets. This paper presents a systematic, comparative evaluation of five prominent LLMs: GPT-4 Turbo, Claude 3 Opus, Gemini 1.5 Pro, Llama 3 70B, and the domain-specialized FinGPT, with respect to their capacity for technical market analysis. The evaluation spans four structured tasks: candlestick pattern recognition from OHLCV data, directional signal generation (BUY/SELL/HOLD), backtesting of signal quality through a simulated execution pipeline, and financial report comprehension. Our experimental framework employs rigorous quantitative metrics, including Sharpe ratio, maximum drawdown, Sortino ratio, information coefficient, F1-score, and BLEU score. Findings from simulated backtesting indicate that GPT-4 Turbo achieves the highest annualized return and Sharpe ratio among general-purpose models, while FinGPT demonstrates competitive risk-adjusted performance due to domain-specific fine-tuning. Both models outperform a passive S&P 500 benchmark under the tested conditions. The study identifies persistent failure modes across all evaluated models, including numerical hallucination, context-window limitations, and inconsistent performance in sideways market regimes. We conclude that while LLMs hold genuine promise within AI trading systems, robust deployment requires careful task decomposition, rigorous backtesting protocols, and domain-aware fine-tuning strategies.

CommentsMaster's research paper, Georgia Institute of Technology. 31 pages, 4 figures

Journal refGeorgia Institute of Technology, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑