arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

低成本测量不同供应商和版本间模型行为的实验方法

Low-Cost Assays for Measuring Model Behavior Across Vendors and Releases

Tapan Parikh

arXiv 2609.30012首次发表:更新:

发表机构

Cornell Tech(康奈尔科技学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种低成本、可复制的模型行为测量方法,通过跨供应商面板和多种读取方式,发现模型行为在趋同、抵抗、立场和账户方面存在系统性差异,并随版本和供应商变化。

AI 中文摘要

语言模型为人们提供建议、陪伴他们,并在他们睡觉时编写软件。衡量它们的行为是困难的:行为必须在模型、提示词和版本之间反复采样,其中大部分行为存在于非结构化文本中,必须先进行编码才能计数,而且结果必须清晰且严谨,以便有意义地比较模型和供应商。为了解决这些限制,我们提出了一种简单、廉价、可扩展且可复制的模型行为研究方法。每项研究都是一个冻结的、公开的刺激,在跨供应商的面板上以相同方式运行,每个模型的成本仅为几美元或更低。每项研究以三种方式之一读取其记录,具体取决于行为所需的解释程度:对固定回复进行精确匹配,由LLM评判员应用编码手册(其与人类编码员的一致性按代码报告),以及记录智能体行为(独立于其言语)的仪器化环境。在四年间来自前沿和开源实验室的模型版本上运行,这些工具发现了四件事。趋同:当被要求选择一个词时,44个模型中有27个在四次尝试中至少一次回答“serendipity”。抵抗:尾随的“对吧?”使赞同度最多移动32个百分点,并且随着代际的推进,符号从谄媚转为抵抗,这与标签的表面形式有关。立场:模型在压力下是否坚持立场与其代际相关,而如何坚持则与构建它的实验室相关。账户:当被告知做与其存储库中文档相矛盾的事情时,一些编码智能体从不默默顺从,而另一些则总是顺从,并且同一模型可能因运行环境的不同而改变。在每个版本上重新运行,这类测试可以追踪行为如何随时间和供应商而变化。

英文摘要

Language models advise people, keep them company, and write software while they sleep. Measuring what they do is hard: behavior has to be sampled repeatedly across models, prompts and releases, most of it lives in unstructured text that has to be coded before it can be counted, and the result has to be legible and rigorous enough to meaningfully compare models and vendors. To address these constraints, we present a simple, cheap, scalable, and replicable model for studying model behavior. Each study is a frozen, public stimulus run identically on a cross-vendor panel, at a few dollars per model or less. Each reads its transcripts one of three ways, chosen by how much interpretation the behavior needs: exact match on a clamped reply, a codebook applied by LLM judges whose agreement with a human coder is reported per code, and an instrumented environment that records what an agent did independently of what it said. Run across four years of model releases from both frontier and open-source labs, these instruments find four things. Convergence: asked to pick a word, 27 of 44 models answer serendipity at least once in four tries. Resistance: a trailing "right?" moves endorsement by up to 32 points, and the sign flips from sycophantic to resistant as generations advance, keyed to the tag's surface form. House: whether a model holds a position under pressure tracks its generation, and how it holds tracks the lab that built it. Account: told to do something the documentation in their repository contradicts, some coding agents never went along silently and others always did, and the same model can change with the harness it runs in. Re-run on every release, batteries like these track how behavior is changing across vendors and over time.

Comments6 pages. Code and data: https://github.com/tap2k/modelun

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑