arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39111cs.CL

Bongard:训练机器直觉

Bongard: Training Machine Intuition

Li Ding, Haidi Jin, Chen Ji

首次发表
浏览论文内容

中文总结 AI 辅助

Bongard是一个开放权重的System One模型,通过三阶段训练和联合嵌入后训练,将机器直觉作为独立能力进行训练,在DecisionBench上达到78.05%准确率,并实现低延迟高吞吐决策。

中文摘要 AI 辅助

人类智能在很大程度上依赖于习得的直觉:无需显式展开每一个中间步骤即可识别模式并判断情境。我们推出Bongard,一个开放权重的System One模型,将机器直觉视为一种可独立设计和训练的能力。一个T5Gemma 2 4B-4B编码器-解码器将阅读证据与做出判断分开。编码器结合问题指令双向读取状态,而独立的解码器分支共享此编码,因此对同一情境的多个判断只需读取一次状态。一个训练好的头部返回对给定候选的概率,而不生成文本。训练分三个阶段进行,从监督判断到语义关系再到行动结果,每个阶段在一个Blackwell GPU上更新全部70.9亿可训练参数。联合嵌入后训练将保留改写集上的准确率从75.7%提升至85.9%。随后一个沙盒阶段通过行动推演和精确预言机学习结果分布,将冻结沙盒面板上的准确率从50.6%提升至64.8%。在DecisionBench上,最终模型在23,900个决策上达到78.05%的准确率,在公开比较的61个系统中排名第四。在一张RTX PRO 6000上,短请求的中位延迟为36毫秒,对一个状态的32个问题耗时221毫秒。Bongard证明了机器直觉可以通过表示学习和结果反馈进行系统训练,为高吞吐决策工作负载提供了一种开放、高效的替代方案。

英文摘要

Human intelligence relies heavily on learned intuition: recognising patterns and judging situations without explicitly unfolding every intermediate step. We introduce Bongard, an open-weight System One model that treats machine intuition as an independent capability to design and train. A T5Gemma 2 4B-4B encoder-decoder separates reading the evidence from making judgments. The encoder reads the state bidirectionally together with the question instructions, and separate decoder branches share this encoding, so many judgments about the same situation require only one reading of the state. A trained head returns probabilities over the supplied candidates without generating text. Training proceeds in three stages, from supervised judgments to semantic relationships to action outcomes, and each stage updates all 7.09 billion trainable parameters on one Blackwell GPU. Joint-embedding post-training raises accuracy on held-out rephrasings from 75.7% to 85.9%. A sandbox stage then learns outcome distributions from action rollouts and exact oracles, raising accuracy on a frozen sandbox panel from 50.6% to 64.8%. On DecisionBench, the final model reaches 78.05% accuracy over 23,900 decisions and ranks fourth of 61 systems in the public comparison. On one RTX PRO 6000, its median latency is 36 ms for short requests, and 32 questions about one state take 221 ms. Bongard demonstrates that machine intuition can be systematically trained via representation learning and outcome feedback, providing an open, efficient alternative for high-throughput decision workloads.

发表机构

  • AgentBull Pte Ltd(AgentBull私人有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑