arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MemTrial:学习何时信任LLM组合智能体中的记忆

MemTrial: Learning When to Trust Memory in LLM Portfolio Agents

Guanghao Wu, Zhuo Cai, Shoujin Wang

arXiv 2610.11732首次发表:更新:

发表机构

University of Technology Sydney(悉尼科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MemTrial通过让记忆接受“审判”、用分层贝叶斯模型池化信用,在投资组合管理中既从重要经验获益,又限制信用失效时的损失,在多个基准上优于现有方法。

AI 中文摘要

用于投资组合管理的大语言模型(LLM)智能体会从经验中学习:它们会根据使用记忆中经验所做决策的结果,为每条经验分配信用。然而在金融市场中,该结果大多反映的是当日所有决策共同面临的市场变动,因此信用追踪的是市场而非经验,这类智能体的表现往往不如直接持有等权重(1/$N$)投资组合。我们探究智能体如何为一条经验分配其带来的改变对应的信用,答案是让记忆接受“审判”:使用某条经验和不使用该经验生成的同一决策草稿,面临相同的市场环境,二者共有的结果会在差值中抵消。我们的智能体MemTrial会针对每条决策,通过部分因子设计选择8种检索经验的组合生成草稿,并以这些差值的平均值(即班扎夫值)为每条经验分配信用。由于每个日期仅出现一次,且每份草稿都是有噪声的LLM样本,这些信用存在噪声,可能无法在新日期上成立。因此MemTrial会通过分层贝叶斯模型跨日期和相似经验池化这些信用,仅在信用能预测未见过的日期时才依据其行动,否则会锚定在1/$N$这类保守基准上。在四个基准测试中,MemTrial不仅能从重要经验中获益(在半合成基准上,其为15种方法中表现最佳,且该基准的经验质量已知),还能在信用不成立时限制损失(在PortBench和InvestorBench上,其损失最多比1/$N$低2.2%,而表现最佳的经验学习智能体损失为15%-38%)。在5种设置的平均表现中,它将表现最佳的经验学习智能体的效用提升了21.2%,且在使用8种LLM时,它在InvestorBench上击败了所有基于LLM的基线方法。

英文摘要

Large language model (LLM) agents for portfolio management learn from experience: they credit each experience in their memory with the outcome of the decisions that used it. In financial markets, however, this outcome mostly reflects the market move shared by all decisions on that date, so the credit tracks the market rather than the experience, and these agents often do worse than simply holding the equal-weight (1/$N$) portfolio. We ask how an agent can credit an experience with what it changes, and answer it by putting memory on trial: drafts of the same decision with and without an experience face the same market, so the outcome they share cancels in their difference. Our agent, MemTrial, drafts each decision with eight combinations of its retrieved experiences, chosen by a fractional factorial design, and credits each experience with its Banzhaf value, the average of these differences. As each date occurs once and each draft is a noisy LLM sample, these credits are noisy and may not hold on new dates. MemTrial therefore pools them across dates and similar experiences with a hierarchical Bayesian model, acts on them only after they have predicted unseen dates, and otherwise stays anchored at a conservative reference such as 1/$N$. On four benchmarks, MemTrial not only benefits from experiences that matter (the best of 15 methods on a semi-synthetic benchmark with known experience quality) but also limits its losses when its values do not hold (at most 2.2\% below 1/$N$ on PortBench and InvestorBench, against 15--38\% for the best experience-learning agent). Averaged over five settings, it improves the utility of the best experience-learning agent by 21.2\%, and with eight LLMs it beats every LLM-based baseline on InvestorBench.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑