arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01213cs.CE

从语言模型股票排名到可检验的经济规则:一项计算审计

From language-model stock rankings to testable economic rules: A computational audit

Shuai Wu, Xue Li, Zhijun Wang, Bolun Liu, Weilin Cai, Zihao Su, Ran Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文对语言模型股票排名进行计算审计,测试其稳定性与投资结果,发现线性规则可近似排名但收益结论依赖比较设置,且个股进出场预测较弱。

中文摘要 AI 辅助

我们检验了语言模型股票排名的稳定性、可复现性及投资结果。四个模型和五个数值比较器在72个月度持有期内共享一个投资组合引擎,应用于上证50、沪深300和中证500指数。排名使用九个特征,五次重复的上证50运行衡量了相同输入下的变异性。基于开发期模型偏好拟合的线性规则在未见月份、更大池和受控干预测试前被冻结。它们与模型排名的平均Spearman一致性在上证50中为0.923-0.984,转移后为0.795-0.985。相对于零变化预测,在十二个存档特征组比较和八个匹配单特征比较中,总体排名变化误差有所下降,并在分别包含九加三和六加二家族的检验中应用了Holm调整。对个股进出场的预测仍然较弱(事件Jaccard指数为0.000-0.125)。历史平均模型复合年增长率范围为6.26%至11.17%。在每边10个基点和六个月块设置下,十二比较的模型减规则收益家族及因子控制关联未产生调整后的显著发现。三个更高成本、十二个月块的比较在其十二个测试切片内支持Terra规则,但在完整的144个测试敏感性网格中无调整后的显著发现。两个输入干预批次共包含5,184个响应,提供了配对干预收益测试。三模型批次在主要块长度下无自举调整发现;Luna行序效应在其三测试家族内出现于异方差和自相关一致(HAC)调整下,但在合并的24测试调整中不显著。紧凑规则近似总体排名;收益结论取决于比较家族、不确定性方法和并列优先级。

英文摘要

We test the stability, reproducibility and investment outcomes of language-model stock rankings. Four models and five numerical comparators share a portfolio engine over 72 monthly holding periods in the Shanghai Stock Exchange (SSE) 50, China Securities Index (CSI) 300 and CSI 500. Rankings use nine characteristics, and five repeated SSE 50 runs measure variation under identical inputs. Linear rules fitted to development-period model preferences are frozen before unseen-month, larger-pool and controlled-intervention tests. Their mean Spearman agreement with model rankings is 0.923-0.984 in the SSE 50 and 0.795-0.985 after transfer. Aggregate rank-change error falls relative to a zero-change prediction in twelve archived feature-group comparisons and eight matched single-feature comparisons, with Holm adjustments applied in separate nine-plus-three and six-plus-two families. Prediction of individual entries and exits remains weak (event Jaccard 0.000-0.125). Historical mean model compound annual growth rates range from 6.26% to 11.17%. At 10 basis points per side and six-month blocks, the twelve-comparison model-minus-rule return family and factor-controlled associations yield no adjusted finding. Three higher-cost, twelve-month-block comparisons favor a Terra rule within their twelve-test slices, with no adjusted finding across the full 144-test sensitivity grid. Two input-intervention batches totaling 5,184 responses supply paired intervention-return tests. The three-model batch has no bootstrap-adjusted finding at the primary block length; a Luna row-order effect appears under heteroskedasticity- and autocorrelation-consistent (HAC) adjustment within its three-test family but not in a pooled 24-test adjustment. Compact rules approximate aggregate rankings; return conclusions depend on comparison families, uncertainty methods and tie priorities.

发表机构

  • University of Colorado Boulder(科罗拉多大学博尔德分校)
  • Beijing Union University(北京联合大学)
  • University of North Dakota(北达科他大学)
  • City University of Macau(澳门城市大学)
  • The University of Sydney(悉尼大学)
  • Taiyuan University of Technology(太原理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑