arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28399cs.AIq-fin.TR

RetailAgent:自条件多模态大语言模型交易智能体中的结构化不利时机

RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents

  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
  • Michigan State University(密歇根州立大学)
  • University of Rochester(罗切斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

Yupeng Zhang, Liuyuan Jiang, Hongyi Huang, Bingheng Li, Lisha Chen

AI总结:

该研究提出RetailAgent框架,发现LLM交易智能体的决策存在跨维度的持续不利时机,动作与后续收益的一致性是该效应的关键驱动因素,同时自生成记忆会增强策略持续性。

AI中文摘要:

在金融市场中,对价格变动做出系统性反应的序列策略可能会被其他市场参与者预测到。本文研究大型语言模型(LLM)智能体是否会表现出这种定向结构,研究基于实验框架RetailAgent展开:LLM观察匿名化的股票日内价格历史和允许状态,然后在后续区间收益披露前反复选择做多(持有股票)或空仓(不参与)。我们在剔除做多决策的整体占比后,比较同一股票日内路径中做多区间和空仓区间的收益。这种暴露匹配度量揭示了跨模态、跨时间范围、跨状态和跨模型家族的持续不利时机。对保存的动作序列进行打乱会显著削弱该效应,表明动作与后续收益的一致性是导致不利得分的原因。将自生成的记忆输入决策会进一步增加策略的持续性,而在智能体同时使用两种动作的股票日中,时机表现得更为不利。这些结果揭示了序列LLM金融决策中稳定、可恢复的定向结构,以及用于研究其他参与者如何响应可预测策略的行为信号。

英文摘要:

In financial markets, a sequential policy that reacts systematically to price movements may become predictable to other market participants. This paper studies whether large language model (LLM) agents exhibit such directional structure through RetailAgent, an experimental framework in which an LLM observes anonymized intraday equity price histories and permitted state, then repeatedly chooses long (hold the stock) or flat (stay out) before the subsequent interval return is revealed. We compare returns during long and flat intervals along the same stock's intraday path after removing the overall fraction of long decisions. This exposure-matched measure reveals persistent negative timing across modality, horizon, state, and model family. Shuffling saved action sequences substantially attenuates the effect, showing that alignment between actions and subsequent returns drives the negative score. Feeding self-authored memories into decisions further increases policy persistence, while timing becomes more negative among stock-days on which the agent uses both actions. These results reveal stable, recoverable directional structure in sequential LLM financial decisions and a behavioral signal for studying how another participant could respond to a predictable policy.

↑