arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25554cs.AI

提炼时间搜索与推理:通过 harness 辅助的高效数据合成发展用于未来预测的大语言模型

Distilling Temporal Search and Reasoning: Evolving LLMs for Future Prediction via Harness-Assisted Efficient Data Synthesis

Wanxu Cai, Zhengyu Chen, Huaisheng Zhu, Wei Wang, Jingang Wang, Qiang Xu

首次发表
浏览论文内容

中文总结 AI 辅助

研究未来事件预测难题,提出时间截断 harness 方法,通过强制时间截断减少时间泄漏与对拒绝采样等的依赖,构建相关语料库和度量标准,经蒸馏实验验证该方法能提升模型表现,将高质量数据转化为模型参数提升。

中文摘要 AI 辅助

未来事件预测具有广泛的社会影响,但仍具有挑战性。当前最优方法通过外部代理框架增强大语言模型,但移除框架后预测能力消失。虽然近期工具集成推理(TIR)将深度搜索内化用于多跳事实检索,但预测还需要对历史趋势和动态变化进行时间搜索和推理。关键障碍是数据:历史查询会导致时间泄漏,使预测退化为检索。先前工作要么用静态观测冻结信息收集,要么依赖拒绝采样或未解决的新查询,丢弃大量数据,降低合成效率。我们提出一种时间截断 harness,每次都强制进行时间截断,实现从历史事件中进行 TIR 式采样,减少时间泄漏以及对拒绝采样或未解决查询的依赖,提高采样效率。我们还构建了大规模语料库和基于过程的度量标准,表明我们的 harness 自然地诱导出更广泛的时间搜索广度,提高高质量数据比例,进一步提高效率并减少对复杂规则的依赖。蒸馏实验表明,在 harness 干预数据上训练的学生模型表现最佳,证明 harness 辅助的模型发展将更高质量的时间搜索和推理数据转化为学生模型的参数提升。

英文摘要

Future event prediction carries broad social impact yet remains challenging. SOTA approaches augment LLMs with external agent frameworks whose predictive capability vanishes once the harness is removed. While recent Tool-Integrated Reasoning (TIR) internalizes deep search for multi-hop retrieval of facts, forecasting further demands temporal search and reasoning over historical trends and dynamic shifts. The key obstacle is data: historical queries induce temporal leakage that degrades forecasting into retrieval. Prior works either freeze information gathering with static observations, or rely on rejection sampling or unresolved fresh queries that discard vast amounts of data, degrading synthesis efficiency. We propose a time-truncation harness that enforces a temporal cut-off at every turn, enabling TIR-style sampling from historical events, reducing temporal leakage and reliance of rejection sampling or unsolved queries, increasing the sampling efficiency. We further build a large-scale corpus and a process-based metric and show that our harness naturally induces a broader temporal breadth of search and raises the proportion of high-quality data, further increasing the efficiency and reducing the reliance on complex rubrics. Distillation experiments show that students trained on harness-intervened data achieve the best performance, demonstrating harness-assisted model evolving that turns higher quality temporal search and reasoning data into a parametric advancement of the students.

发表机构

  • The Chinese University of Hong Kong(香港中文大学)
  • Meituan LongCat Team(美团龙猫团队)

机构由 AI 辅助整理,请以论文原文为准。

↑