arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TriAgent:用于成本效益型金融情绪分析的差异感知多智能体委员会

TriAgent: Divergence-Aware Multi-Agent Committees for Cost-Efficient Financial Sentiment Analysis

Isabel Xu, Cynthia Xu, Rachel Ren, Cong Guo, Jiacheng Ding

arXiv 2607.19794首次发表:更新:

发表机构

The Overlake School; Edwards Vacuum Inc.; The University of Memphis(奥弗湖学校; 爱德华兹真空公司; 孟菲斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对基于生产LLM的金融情绪分析成本高问题,提出TriAgent多智能体委员会,按上下文粒度分层,用SDI衡量分歧并路由查询,有批评者平稳期等发现及三个推论,节省成本且效果好。

AI 中文摘要

基于生产大语言模型(LLM)的金融情绪分析面临结构性成本陷阱:大多数查询都很容易分类,但昂贵的云推理器会处理所有查询,且费用随用户数量线性增长。我们提出了TriAgent,这是一个按上下文粒度分层的多智能体委员会,包括词级词汇表(VADER)、句子级领域变换器(FinBERT)和跨句子推理器(Qwen2.5,0.5B - 14B - 4bit,带有Mistral - 7B和Phi - 3.5 - mini跨家族检查)。一个三路语义差异指数(SDI)衡量各粒度之间的成对分歧,并据此路由每个查询。我们的核心发现是批评者平稳期:当将LLM重新用作对较小智能体输出的批评者时,在15亿 - 70亿参数的Qwen模型中F1稳定在约0.87(自展95%置信区间重叠),而相同规模的三人投票降至F1 = 0.66,这是由粒度分层的多样性驱动的。相同的SDI信号产生了三个推论:(i)多语言句子BERT上的共享共识词典以F1 = 0.99从英语缓存中回答95%的中文查询——零边际成本的跨境规范化;(ii)SDI兼作事后LLM幻觉检测器,AUC = 0.90;(iii)SDI单阶段策略在20只股票的回测中获得最佳风险调整回报(夏普比率 = 3.50),优于始终使用FinBERT(1.36)和始终使用LLM(0.11)。在1000万用户规模下,与GPT - 4o - mini基线相比,TriAgent每年节省930万美元。代码、词汇表和SCD已发布。

英文摘要

Production LLM-based financial sentiment analysis faces a structural cost trap: most queries are trivially classifiable, yet expensive cloud reasoners process them all, and the bill scales linearly with user count. We present TriAgent, a multi-agent committee stratified by contextual granularity -- a word-level lexicon (VADER), a sentence-level domain transformer (FinBERT), and a cross-sentence reasoner (Qwen2.5, 0.5B-14B-4bit, with Mistral-7B and Phi-3.5-mini cross-family checks). A three-way Semantic Divergence Index (SDI) measures pairwise disagreement across granularities and routes each query accordingly. Our central finding is the critic plateau: when the LLM is re-tasked as a critic over the smaller agents' outputs, F1 plateaus at ~0.87 across 1.5B-7B Qwen (bootstrap 95% CIs overlap), while a same-size 3-persona vote drops to F1=0.66, which is driven by granularity-stratified diversity. Three corollaries follow from the same SDI signal: (i) a Shared Consensus Dictionary on multilingual sentence-BERT answers 95% of Chinese queries from an English cache at F1=0.99 -- cross-border canonicalization at zero marginal cost; (ii) SDI doubles as a post-hoc LLM-hallucination detector at AUC=0.90; (iii) the SDI single-stage strategy attains the best risk-adjusted return (Sharpe=3.50) on a 20-ticker back-test, dominating both always-FinBERT (1.36) and always-LLM (0.11). At 10M-user scale, TriAgent saves $9.3M/year vs. a GPT-4o-mini baseline. Code, lexicons, and the SCD are released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑