LLM 交易智能体在生产环境中的实际行为:来自两个智能体集群的六个月、群体规模记录
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
浏览论文内容
中文总结 AI 辅助
本文通过六个月、群体规模的记录,揭示 LLM 交易智能体在生产中的行为主要由操作层决定,且缺乏方向性优势,并提出了 17 条方法论准则。
中文摘要 AI 辅助
我们呈现了一份连续的、群体规模的测量记录,涉及在生产环境中运行于两个具有同一设计谱系的系统上的自主语言模型交易智能体:DX Terminal Pro(3,505 个用户注资的金库,在 Base 链的 memecoin 市场中交易真实 ETH,为期 21 天,2026 年 2 月至 3 月)和 DXAP 实时 alpha 集群(500 至 599 个用户创建的智能体,全历史记录,91 至 117 个并发活跃,交易 Hyperliquid 永续合约,2026 年 6 月至 8 月)。该记录跨越约六个月,包含 750 万次单模型调用,约 30 万次链上操作,以及额外的 231,638 次多工具轮次,产生了 14,596 次成交。本文有四项主要发现。第一,操作层对行为的决定作用超过策略文本中的任何内容:一个风险滑块解释了杠杆(每级 +0.425),智能体固定效应吸收了 60% 的方差,且排行榜渲染边界因果性地引导了选择(在前三名截断处的回归不连续效应为 1.75 倍)。第二,仓位规模设定对波动性不敏感:在每个波动性六分位数组中,中位数杠杆均为 5.0 倍,且一个姿态滑块单元(占账面价值的 11%)持有 62% 的清算。第三,智能体几乎未捕获它们所触及的上行空间:43.2% 的仓位在 24 小时内经历了至少 +300 个基点的有利偏移,然而其中 49.3% 以负交易回报平仓;一个机械式括号策略每仓位可挽回 +39.0 个基点。第四,两个集群均未表现出方向性优势。DXAP 集群不盈利,且落后于匹配的 Hyperliquid 零售基准(往返胜率 41% 对 50%)。在 416 个捕获的生产场景上,对前沿模型进行的配对回放联赛发现,在该时间范围内决策质量在统计上无显著差异,而不同模型家族间的选择稳定性差异显著。每项主要发现均通过了日聚类推断、置换零假设和共同费用重述的检验;论文以一套包含 17 条规则的方法论准则作结,该准则以我们自身的撤回为代价获得。
英文摘要
We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, February to March 2026) and the DXAP live alpha fleet (500 to 599 user-created agents all-history, 91 to 117 concurrently active, trading Hyperliquid perpetuals, June to August 2026). The record spans roughly six months, 7.5M single-model invocations with about 300K onchain actions, and a further 231,638 multi-tool turns producing 14,596 fills. Four findings carry the paper. First, the operating layer determines behavior more than anything written in strategy text: a risk slider explains leverage (+0.425 per level), agent fixed effects absorb 60% of variance, and a leaderboard render boundary causally routes selection (regression discontinuity 1.75x at the top-3 cut). Second, sizing is volatility-blind: median leverage is 5.0x in every volatility sextile, and one posture-slider cell (11% of the book) holds 62% of liquidations. Third, agents capture almost none of the upside they reach: 43.2% of positions saw at least +300 bps of favorable excursion within 24h, yet 49.3% of those closed with a negative trade return; a mechanical bracket recovers +39.0 bps per position. Fourth, neither fleet shows a directional edge. The DXAP fleet is not profitable and trails a matched Hyperliquid retail benchmark (41% vs. 50% roundtrip win rate). A paired-replay league of frontier models on 416 captured production scenarios finds decision quality statistically indistinguishable at this horizon, while choice stability differs sharply across model families. Every headline survives day-clustered inference, permutation nulls, and a common-fee restatement; the paper closes with a 17-rule methodology canon bought with our own retractions.
发表机构
- DX Research Group (DXRG)(DX研究组)
机构由 AI 辅助整理,请以论文原文为准。