发表机构
Indian Institute of Technology Madras(印度理工学院马德拉斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对持续预训练评估中的知识污染问题,提出虚构世界基准ORDER,通过合成物理规则和组合任务验证领域自适应,证明小模型经适应后能超越GPT-4.1,且任务性能比知识测试更能预测规划质量。
AI 中文摘要
通过持续预训练将语言模型适应到新领域,引发了一个基本的评估问题:如果训练语料与模型已知内容重叠,性能提升就不能被清晰地归因于新学习而非已有知识。这个问题对于知识密集型、任务轻量(KHTL)的机器人部署最为关键——例如药品配发、危险品处理、设施特定规程,这些场景中物理任务简单,但管理规则是专有的且关乎安全,并且大量现场测试成本高昂或不安全。我们提出了ORDER(面向具身推理的本体驱动决策),一个基于虚构世界的基准:一个342,069词元的合成语料库定义了一个自洽的物理规则,该规则不可能出现在任何模型的预训练数据中。ORDER将一个500题的知识测试(ORDER-BENCH)与一个更难的组合任务ORDER-SPATIAL配对:在熟悉和全新场景中对物体进行安全操作排序。未经适应的GPT-4.1在ORDER-SPATIAL上得分低于随机水平(Kendall's tau = 0.441),表明其先验知识与虚构物理规则存在冲突。经过持续预训练后,小模型在熟悉和全新场景上均有显著提升,这证明了真正的世界模型归纳而非记忆。我们进一步将其应用于机器人流程:在知识测试中表现良好的模型,往往需要额外的技能适应阶段才能生成有效、可执行的计划;经过该阶段后,完全离线的小模型在完整的感知到执行循环中(在模拟iiwa7机械臂上演示,并有人类在环修正)优于GPT-4.1,即使GPT-4.1被赋予对相同规则的检索访问权(Kendall's tau = 0.848 对比 0.606)。在整个过程中,ORDER-SPATIAL的性能而非知识测试准确率才是预测真实计划质量的关键。
英文摘要
Adapting language models to new domains via continual pre-training raises a basic evaluation problem: if the training corpus overlaps with what the model already knows, performance gains cannot be cleanly attributed to new learning rather than pre-existing knowledge. This matters most for knowledge-intensive, task-light (KHTL) robot deployments - pharmaceutical dispensing, hazardous-material handling, facility-specific protocols, where the physical task is simple but the governing rules are proprietary and safety-critical, and where extensive live testing is costly or unsafe. We introduce ORDER (Ontology-driven Decision-making for Embodied Reasoning), a benchmark built on a fictitious world: a 342,069-token synthetic corpus defining a self-consistent physics that cannot appear in any model's pre-training data. ORDER pairs a 500-question knowledge test (ORDER-BENCH) with a harder compositional task, ORDER-SPATIAL: ordering objects for safe manipulation across both familiar and entirely novel scenes. GPT-4.1 without adaptation scores below chance on ORDER-SPATIAL (Kendall's tau = 0.441), showing its priors actively conflict with the invented physics. After continual pre-training, small models improve substantially on both familiar and novel scenes alike evidence of genuine world-model induction rather than memorization. We then carry this through to a robot pipeline: models that answer the knowledge test well often cannot produce valid, executable plans without a further skill-adaptation stage, after which small, fully offline models outperform GPT-4.1 even when GPT-4.1 is given retrieval access to the same rules (Kendall's tau = 0.848 vs. 0.606), on a full perception-to-execution loop demonstrated on a simulated iiwa7 arm with human-in-the-loop correction. Throughout, ORDER-SPATIAL performance, not knowledge-test accuracy is what predicts real plan quality.