arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22478cs.AIcs.CLcs.LG

托管大语言模型中的无持久性复制:行动时间信念评估中的测量敏感性

Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation

Bhushan Kashinath Joshi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究在Regent Chess环境中区分复制、测量敏感性与持久性,发现托管模型行为评估结果因配置和标识符而异,需显式索引声明条件。

中文摘要 AI 辅助

对托管语言模型的行为评估可能因评估服务、测量工具或两者在不同运行中的差异而产生变化。我们区分了三个验证问题:先前发现是否在其历史配置下于新数据上重现(复制);当评估与推理配置在同一标识符下重建时,端点是否改变(测量敏感性);以及在一种共同工具下,该发现是否在后续测试的标识符中持续存在(持久性)。我们在Regent Chess中研究这些问题,这是一个顺序环境,其中隐藏的、可变的状态被精确记录,使得陈述的信念可以在行动时间与真实情况进行评分;正端点值意味着性能比匹配均匀比较器更差。先前报告的Gemini 3.1 Flash-Lite缺陷在其历史配置下于新游戏中重现(+0.0530,95%置信区间[+0.0329,+0.0714])。在同一公共标识符下的背靠背同日H/R比较中,重建配置下的模型减均匀端点低0.0429(H减R对比的95%置信区间[+0.0182,+0.0667]);所有六个配置组件共同变化,因此没有组件被隔离。在重建的R下,预先冻结的、交错同窗口4K比较在Gemini 3.1和Gemini 3.7之间逆转了符号,这些标识符在发布和产品层级上有所不同;额外的描述性和探索性单元显示相同的方向模式。任何额外的服务期间贡献仍未解决(-0.0166,[-0.0483,+0.0157])。因此,复制、测量敏感性和持久性在一次评估中可能产生不同的结论,这促使对托管模型行为声明按测试标识符、服务期间、测量工具和推理配置进行显式索引。

英文摘要

Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement instrument, or both differ across runs. We separate three validation questions: whether a prior finding recurs on fresh data under its historical configuration (replication), whether the endpoint changes when the evaluation-and-inference configuration is rebuilt under the same identifier (measurement sensitivity), and whether the finding persists across subsequently tested identifiers under one common instrument (persistence). We study these questions in Regent Chess, a sequential environment in which a hidden, mutable state is recorded exactly, allowing stated beliefs to be scored against ground truth at action time; positive endpoint values mean worse performance than a matched-uniform comparator. The previously reported Gemini 3.1 Flash-Lite deficit recurs on fresh games under its historical configuration (+0.0530, 95% CI [+0.0329,+0.0714]). In a back-to-back same-day H/R comparison under the same public identifier, the model-minus-uniform endpoint is 0.0429 lower under the rebuilt configuration (95% CI for the H-minus-R contrast [+0.0182,+0.0667]); all six configuration components vary jointly, so no component is isolated. Under rebuilt R, the prospectively frozen, interleaved same-window 4K comparison reverses sign between Gemini 3.1 and Gemini 3.7, identifiers that differ in release and product tier; additional descriptive and exploratory cells show the same directional pattern. Any additional serving-period contribution remains unresolved (-0.0166, [-0.0483,+0.0157]). Replication, measurement sensitivity, and persistence can therefore yield different conclusions within one evaluation, motivating explicit indexing of hosted-model behavioural claims by tested identifier, serving period, measurement instrument, and inference configuration.

↑