发表机构
HKUST(GZ); Paradoox AI; E Fund Management Co., Ltd; MBZUAI; The University of Tokyo(香港科技大学(广州); 悖论人工智能公司; 易方达基金管理有限公司; Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学); 东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于LLM的交易系统中智能体能否为自身智能付费的问题,引入TradeLens工具包进行评估,通过多方面分析发现可行性取决于智能到利润的转化,不同模型有不同失败模式,重构了此类交易智能体的评估方式。
AI 中文摘要
大语言模型(LLM)智能体越来越多地用于交易系统,模型推理、工具使用和持续决策会产生成本,期望能带来交易价值。现有评估通常报告性能指标,很少考察智能体的可行性,即动态的由LLM介导的决策能否将产生的成本转化为可衡量的增量利润。为应用此标准,我们引入TradeLens,这是一个基于跟踪的诊断工具包,用于从交易记录、运行时跟踪和部署配置评估智能交易系统。它重建交易轨迹,将利润和成本归因于可解释的证据,并诊断智能体是否以及为何为自身智能付费。我们对骨干模型、资本规模、交易频率和系统架构进行了广泛分析,并进行了部署讨论。结果表明,可行性取决于智能到利润的转化:不同模型有不同失败模式,如DeepSeek - V3.2中资产选择不佳和GLM - 4.7中时机不利,而资本规模、交易频率和架构仅通过放大或降低决策归因的时机价值起作用。这些发现将基于LLM的交易智能体的评估从以能力为中心的性能排名重新构建为基于跟踪的智能到利润转化诊断。
英文摘要
Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce trading value. Existing evaluations typically report performance metrics, but rarely examine agentic viability: whether dynamic LLM-mediated decisions convert their induced costs into measurable incremental profit. To apply this criterion, we introduce TradeLens, a trace-grounded diagnostic toolkit for evaluating agentic trading systems from their trading records, runtime traces, and deployment configurations. It reconstructs trading trajectories, attributes profit and cost to interpretable evidence, and diagnoses whether and why an agent pays for its own intelligence. We conduct extensive analysis across backbone models, capital scales, trading frequencies, and system architectures, together with deployment discussion. Our results show that viability hinges on intelligence-to-profit conversion: models exhibit different failure patterns, such as poor asset selection in DeepSeek-V3.2 and negative timing in GLM-4.7, while capital scale, trading frequency, and architecture matter primarily in these runs, especially through their effects on decision-attributed timing value. These findings reframe the evaluation of LLMbased trading agents from capability-centric performance ranking to trace-grounded diagnosis of intelligence-to-profit conversion. Our code is available at https://github.com/ParadooxAI/TradeLens.
CommentsAccepted by EMNLP 2026 Findings