arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从搜索到研究:自主量化因子挖掘中的搜索扩展探索

From Search to Research: Exploring Search Scaling in Autonomous Quantitative Factor Mining

Kangcheng Deng, Hui Cai, Jiacheng Lu, Chester Zhongshu Qian, Rui Sun, Beidi Luan, Jing Li, Daxin Jiang, Zuo Bai

arXiv 2609.35559首次发表:更新:

发表机构

StepFun; Shanghai Jiao Tong University; University of California, Los Angeles; FinStep(阶跃星辰; 上海交通大学; 加州大学洛杉矶分校; FinStep)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究基于50个量化因子挖掘任务,探究自主研究智能体中搜索扩展对性能的影响,发现模型能力决定初始表现、深度搜索缩小差距、并行优于顺序搜索,并指出未来需更强模型与自适应计算部署策略。

AI 中文摘要

推理扩展已被证明能够提升大语言模型(LLM)的性能,这一原理通过增加搜索预算自然延伸到自主LLM智能体,我们将其称为“搜索扩展”。尽管已有研究刻画了LLM推理扩展的机制、扩展行为及性能极限,但在自主研究领域,这些问题仍鲜为人知。因此,我们基于金融研究报告中的50个量化因子挖掘任务,探究搜索扩展如何影响研究性能及其背后的驱动机制。每个任务要求智能体执行完整的研究循环,从解读假设、用代码实现,到评估并迭代优化所得因子。在九个模型上,我们通过追踪不同预算下的性能、在模型间迁移中间研究状态以及比较不同搜索策略,考察模型能力、搜索深度和搜索组织方式如何塑造因子质量。我们发现:(1)初始性能与模型能力的关联更强,而更深的搜索可以缩小模型间的差距;(2)模型嫁接表明,早期研究状态对最终性能有实质性影响;(3)在相同迭代预算下,并行搜索优于顺序搜索,这与更广泛覆盖搜索空间带来的收益一致。进一步的轨迹分析显示,性能更高的模型能更有效地诊断失败、调整搜索方向,并在选择候选因子时保留预期的经济假设。这些发现表明,自主研究的未来进展将需要更强的模型,以及在整个研究过程中自适应部署测试时计算的政策。

英文摘要

Inference scaling has been shown to improve large language model (LLM) performance, and this principle naturally extends to autonomous LLM agents through increased search budgets, which we refer to as *search scaling*. Although prior work has characterized the mechanisms, scaling behavior, and performance limits of LLM inference scaling, much less is known about these questions in autonomous research. Therefore, we investigate how search scaling affects research performance and what mechanisms drive these gains using 50 quantitative factor-mining tasks grounded in financial research reports. Each task requires an agent to carry out an end-to-end research loop, from interpreting a hypothesis and implementing it in code to evaluating and iteratively refining the resulting factor. Across nine models, we examine how model capability, search depth, and search organization shape factor quality by tracing performance across varying budgets, transferring intermediate research states between models, and comparing different search strategies. We find that (1) initial performance is more strongly associated with model capability, while deeper search can narrow cross-model gaps; (2) model grafting shows that the early research state materially shapes final performance; and (3) parallel search outperforms sequential search under the same iteration budget, consistent with benefits from broader coverage of the search space. Further trajectory analysis shows that higher-performing models more effectively diagnose failures, revise search directions, and preserve the intended economic hypothesis when selecting candidates. These findings suggest that future progress in autonomous research will require stronger models together with adaptive policies for deploying test-time computation throughout the research process.

Comments33 pages, including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑