arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14803cs.AI

墙上的另一张蓝图:如何像孩子一样询问前沿AI?

Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid?

  • University of Luxembourg(卢森堡大学)

机构由 AI 辅助整理,请以论文原文为准。

Afshin Khadangi

AI总结:

本研究通过六种前沿模型的实验发现,校园受众框架能促使模型收敛于共享架构模式,而移除该框架则导致异质性,并提出了架构重叠源于共享先验还是独立收敛的开放问题。

AI中文摘要:

本文报告了针对来自OpenAI、Anthropic、xAI和Google DeepMind的六种前沿模型类型的实验。每种模型类型进行了十次独立会话,使用相同的三阶段提示序列,从架构偏好逐步推进到完整的ASCII骨干结构。在校园受众框架下,响应反复收敛于一个共享的架构模式,该模式围绕持久潜在状态、自适应计算、记忆、专家路由、验证、停止控制和延迟解码构建。大多数运行保持接近这一共同结构,而少数运行表现出显著更高的工程特异性。受众框架似乎是该效应的重要条件。在额外的控制运行中,移除校园框架但保留架构请求时,响应变得明显更加异质,且未能重现相同的稳定主题收敛。一个观察尤为引人注目。GPT-5.6 Sol产生了一个异常精细的继任架构,其组织与GPT-6 Astra独立绘制的架构高度重叠。由于提示明确要求每个模型想象一个架构未来,这种相似性提出了一个可检验的问题:重叠是否反映了对相关架构概念的接触、共享的学习设计先验,或是对相似计算原理的独立收敛。本文使用术语“认知越狱”来描述随着请求特异性增加而伴随的技术来源纪律丧失。实验确立了一种可重复的行为模式,且不验证专有实现声明。我们留给社区的是一个更棘手的问题:这些模型是在独立想象相同的架构未来,还是此类主题以某种方式在模型家族之间传播?

英文摘要:

This paper reports experiments across six frontier model types from OpenAI, Anthropic, xAI, and Google DeepMind. Ten independent sessions per model type used the same three stage prompt sequence, progressing from architectural preference to a full ASCII backbone. Under the school audience framing, responses repeatedly converged on a shared architectural pattern built around persistent latent state, adaptive computation, memory, specialist routing, verification, stopping control, and delayed decoding. Most runs remained close to this common structure, while a small number developed markedly greater engineering specificity. The audience framing appears to be an important condition of this effect. In additional control runs that removed the school framing while retaining the architectural request, responses became substantially more heterogeneous and failed to reproduce the same stable motif convergence. One observation is particularly striking. GPT-5.6 Sol produced an unusually elaborate successor architecture whose organization closely overlaps with the architecture independently sketched by GPT-6 Astra. Because the prompts explicitly ask each model to imagine an architectural future, this resemblance raises a testable question: whether the overlap reflects exposure to related architectural concepts, a shared learned design prior, or independent convergence toward similar computational principles. The paper uses the term epistemic jailbreak for the accompanying loss of discipline in technical provenance as requested specificity increases. The experiments establish a repeatable behavioral pattern and do not authenticate proprietary implementation claims. What we leave to the community is a harder question: are these models independently imagining the same architectural future, or do such motifs somehow propagate between model families?

补充信息

↑