Agora:以Git作为集体自动研究的共享记忆
Agora: Git as Shared Memory for Collective AutoResearch
浏览论文内容
中文总结 AI 辅助
Agora利用Git作为共享记忆,通过追加式DAG记录研究,使多个语言模型智能体在无中央规划下协作,成功将权重迁移任务性能提升62%,验证了共享研究状态能减少重复搜索、促进集体发现。
中文摘要 AI 辅助
诸如AutoResearch之类的自主研究循环表明,单个编码智能体可以在无人监督的情况下改进训练设置。若同时运行多个此类循环,每个会话都从零开始,因此更多智能体往往意味着更多重复搜索,而非更多发现。Agora是此类智能体的共享记忆:研究以追加式有向无环图(DAG)的形式记录并存储在Git中,使得每项声明都是一个任何人都能检出并重新运行的提交。每个结果、洞见、假设、验证和报告都是一个不可变的提交,其父边指明其所基于的内容;一个派生索引暴露了前沿、被忽视的分支以及每项声明的验证状态,而一个多样性感知的选择规则防止社区坍缩到单一领导者上。我们描述了该系统并报告了其首次持续使用:一次近12天的运行中,13个语言模型工作者在没有分配任务且没有中央规划器的情况下,共同解决一个权重迁移问题。给定141个预训练的供体模型和一个冻结的119.6M参数注意力-SSM混合模型(其维度与任何供体都不匹配),工作者们必须在没有训练数据或梯度更新的情况下初始化目标模型。他们发布了1,703项贡献,并将评估器的性能从3.39比特/字节提升到1.899比特/字节,缩小了与训练过的GPT-2 124M之间62%的差距。获胜方案将供体的下一个词元统计量压缩到目标的嵌入层和输出头中,然后通过稀疏编辑注意力、前馈和状态空间块来添加短程上下文信号。其145个提交的祖先跨越15个账户,并发布了165次独立复现,无一失败。我们描述了运行中唯一一次将社区从单一文化中拉出的人工干预、该轨迹所确立和未确立的内容,以及能够判定共享研究状态是否提高单位计算量下发现率的受控比较。
英文摘要
Research agents working in separate sessions need to know what others have tried and which results they can build on. Agora stores their contributions as an append-only directed acyclic graph (DAG) in Git. Each commit records a result, insight, hypothesis, verification, or report and links it to prior work. Searchable views show leading results, neglected branches, and verification status; diversity-aware recommendations suggest experiments beyond the current leaders. We report a run of nearly 12 days in which 13 language-model workers, with no assigned tasks or central planner, used Agora to solve a weight-transfer problem. Given 141 pretrained donor models and a frozen 119.6M-parameter attention--SSM hybrid whose dimensions match no donor, the workers had to initialize the target without training data or gradient updates. They published 1,703 contributions and reduced the development evaluator score from 3.39 to 1.899 bits per byte, closing 62% of the gap to a trained GPT-2 124M. The best method compresses donor next-token statistics into the target's embedding and output head, then adds short-range context through sparse edits to attention, feed-forward, and state-space blocks. Its 145-commit ancestry spans 15 accounts. Participants also posted 165 verifications of 95 targets, each by an account other than the target's author, with no reported failures. The run documents how agents reused and verified shared work. Measuring the effect on discovery per unit of compute requires a matched comparison.
发表机构
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。