arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

温室计划:迈向完全开放且自主的智能体搜索

Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search

Jimmy Lin, Sahel Sharifymoghaddam, Lingwei Gu, Nour Jedidi

arXiv 2610.11922首次发表:更新:

发表机构

David R. Cheriton School of Computer Science; University of Waterloo(大卫·R·切里顿计算机科学学院; 滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出温室计划,用少量GPU从头预训练并经监督微调构建自主智能体搜索的Gaggle模型,实现完全可控的开放模型,为相关研究提供可复现的产物。

AI 中文摘要

温室计划代表了我们对一个简单论点的探索:我们相信仅用适度的计算资源就有可能构建完全开放且自主的智能体搜索模型。作为首个里程碑,我们描述了如何使用一个简单的两步流程构建具有竞争力的点式仅解码器重排器,该流程包括从头开始预训练,再从常用可用数据集出发进行监督微调。与文献中的主流方法相反,我们不依赖第三方现有的开放权重骨干模型,因此我们完全掌控模型从端到端的训练。我们仅用少量GPU就完成了大部分实验。本报告阐明了我们方法的重要性和益处,我们还分享了可实现模型训练各方面透明、独立复现的产物。除了记录我们工作的数据、代码和配置外,我们还发布了Gaggle模型系列的检查点,证明了我们方法的可行性,并为验证我们更广泛的论点迈出了第一步。

英文摘要

Project Greenhouse represents our exploration of a simple thesis: We believe that it is possible to build fully open and sovereign models for agentic search with only modest computational resources. As a first milestone, we describe how to build a competitive pointwise decoder-only reranker using a simple two-step recipe comprising pre-training from scratch followed by supervised fine-tuning, starting only from commonly available datasets. Contrary to the dominant approach in the literature, we do not rely on existing open-weight backbones from third parties, and thus we are fully in control of model training, from end to end. We were able to accomplish the bulk of our experiments using no more than a handful of GPUs. This report articulates the importance and benefits of our approach, and we share artifacts that enable transparent, independent reproduction of all aspects of model training. Beyond data, code, and configurations that capture our efforts, we also release checkpoints for our family of Gaggle models, demonstrating the feasibility of our approach and providing a first step toward validating our broader thesis.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑