发表机构
Nanyang Technological University; Southern University of Science and Technology; Tianjin University; Sun Yat-sen University; Carnegie Mellon University; Georgia Institute of Technology; The Hong Kong Polytechnic University; Shenzhen University; Tsinghua University; The Hong Kong University of Science and Technology(南洋理工大学; 南方科技大学; 天津大学; 中山大学; 卡内基梅隆大学; 佐治亚理工学院; 香港理工大学; 深圳大学; 清华大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究从分叉未来采样后续操作以比较语言模型隐藏状态,发现共享接口在多模型评估中表现最优,支持测试操作库中存在可复用因果接口。
AI 中文摘要
相同的语言模型答案可能源于支持不同后续计算的隐藏状态,因此当前答案探针无法建立可复用的内部接口。我们引入分叉未来:仅在形成前缀状态后才采样后续操作,并通过这些操作诱导的响应分布比较状态。这在无需研究人员指定潜在标签的情况下,得到了隐藏状态的经验因果商。随后,共享(Shared)、局部(Local)、混合(Mixture)和分布式(Distributed)接口在预序因果描述长度下竞争,需满足后续签名保真度和匹配容量约束。在两项详细的模型评估中,共享接口具有最低的保留描述长度,在Qwen2.5-1.5B上提升0.216纳特,在Llama-3-8B上提升0.294纳特,同时保持紧密聚类的平均后续签名失真;五骨干扫描保留了共享增益的正向趋势。与图对齐的移植分析显示,共享接口在目标正确性、局部性、复制保留和复合特征上具有最强的联合表现,且API对齐路径介导了0.749的目标效应,而匹配的空路径仅为0.150。在盲四分类模型-机体测试中,16个架构中有14个被正确识别,12个非共享机体中观察到1个非共享到共享的错误。这些结果支持在测试的操作库中存在经济的可复用因果接口,但该声明明确依赖于候选架构、干预措施和保留的未来。
英文摘要
Identical language-model answers can arise from hidden states that support different future computations, so current-answer probes do not establish a reusable internal interface. We introduce forked futures: future operations are sampled only after a prefix state has formed, and states are compared through the response distributions induced by those operations. This yields an empirical causal quotient over hidden states without requiring researcher-specified latent labels. Shared, Local, Mixture, and Distributed interfaces then compete under prequential causal description length subject to future-signature fidelity and matched capacity constraints. In the two detailed model evaluations, Shared has the lowest held-out description length, with gains of 0.216 nats on Qwen2.5-1.5B and 0.294 nats on Llama-3-8B, while maintaining tightly clustered mean future-signature distortion; a five-backbone sweep preserves the positive direction of Sharedness Gain. The figure-aligned transplantation analysis gives Shared the strongest joint target-correctness, locality, copy-preservation, and composite profile, and API-aligned paths mediate 0.749 of the target effect versus 0.150 for matched null paths. In the blind four-class model-organism test, 14/16 architectures are recovered, with one observed non-Shared to Shared error among 12 non-Shared organisms. These results support an economical reusable causal interface within the tested operation banks, while keeping the claim explicitly conditional on the candidate architectures, interventions, and held-out futures.
CommentsAuthor list corrected to remove a researcher who was mistakenly included in the previous version and had no involvement whatsoever in this project