发表机构
Imperial College London; Tsinghua University; University of Illinois at Urbana-Champaign(伦敦帝国理工学院; 清华大学; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究考察数学大语言模型在无外部标准时(子领域间翻译)的表征,引入盲队列工具度量改写对隐含假设的陈述,发现普遍性方向导致量化域不对称扩大,且能力不主导此现象。
AI 中文摘要
大型语言模型(LLMs)在竞赛数学上已达到专家级水平,这主要归功于围绕它们的大量搜索:候选解被大量采样,并且仅当外部标准接受时才保留。这样的过程改善了幸存下来的结果,但未触及模型所表征的内容。我们在不存在外部标准的情况下考察这一问题:在相邻子领域的方言之间翻译陈述,其中保真度取决于内容被断言的普遍性水平。源文本在其词汇中隐含了这一水平,因此忠实的翻译必须从理论之间的关系中恢复它。我们引入了一种工具,将真值、内容和范围编码到独立的盲队列中,并采用一种无评判者的度量来判断改写是否陈述了其源文本中隐含的假设,并通过一个植入阳性的对照确立了其敏感性。在来自四个家族的七个模型中,向一般框架翻译时,60.6%的改写扩大了量化的域,而没有任何改写缩小它;向具体框架翻译时,28.3%的改写缩小了域,0.3%的改写扩大了它。能够防止这种情况的假设在21.6%的模型改写和4.2%的人类陈述中被陈述。能力并不支配这种不对称性:它出现在每个测试的模型中,且最有能力的模型扩大得最少。该结果在基准测试中由预注册规则保留的一半数据上以及由数学家撰写的陈述上得到复现。指示模型陈述其所需的每个假设会提高该比率,但不会提高其对方向的敏感性。我们认为,这些系统已经获得了子领域词汇之间的对象级对应,而没有翻译在理论之间携带假设到假设的约束。
英文摘要
Large language models (LLMs) have reached expert-level performance on competition mathematics largely through the volume of search placed around them: candidate solutions are sampled in quantity and retained only when an external criterion accepts them. Such a procedure improves the outcome that survives it while leaving untouched what the model represents. We examine that question where no external criterion exists: translating statements between the dialects of neighbouring subfields, where fidelity turns on the level of generality at which content is asserted. The source leaves that level implicit in its vocabulary, so a faithful translation must recover it from the relation between the theories. We introduce an instrument that codes truth, content and scope in separate blind queues, with a judge-free measure of whether a rewrite states the hypothesis implicit in its source, and establish its sensitivity with a planted-positive control. Across seven models from four families, translating towards the general framing widens the domain of quantification in 60.6% of rewrites and narrows it in none; translating towards the concrete framing narrows it in 28.3% and widens it in 0.3%. The hypothesis that would prevent it is stated in 21.6% of model rewrites and 4.2% of human statements. Capability does not govern the asymmetry: it appears in every model tested, and the most capable widens least. It replicates on the half of the benchmark held out by a pre-registered rule, and on statements written by mathematicians. Instructing a model to state every hypothesis it requires raises that rate but not its sensitivity to direction. We argue that these systems have acquired an object-level correspondence between subfield vocabularies without the constraint under which a translation between theories carries hypotheses to hypotheses.