移动目标:对开源聊天语言模型十二个检查点的可信度漂移进行纵向审计
The Moving Target: A Longitudinal Audit of Trust-Benchmark Score Drift Across Open-Source Chat LLM Release Lines
浏览论文内容
中文总结 AI 辅助
通过对四个开源版本系列在多提示模板下的信任基准测试,发现相邻版本间存在显著漂移,结论是发布线的信任分数应重新测量,而非沿用,应作为检查点相关的日期化工件报告。
中文摘要 AI 辅助
模型卡片引用信任基准分数时未记录测量时间,同一分数在一个发布线的连续检查点沿用。通过对四个开源版本系列在多提示模板下的信任基准测试,发现相邻版本间漂移明显。结论是发布线的信任分数应重新测量,应作为检查点相关的日期化工件报告,封闭API等不在审计范围内。
英文摘要
Trust-benchmark scores reported on a chat-LLM release line are often carried across several checkpoints of the same line, as if the underlying model had not shifted between releases. We test that assumption. We audit four open-source release lines (Yi, Qwen, Mistral, and Gemma) at three successive public generations each. Each checkpoint is scored on a fixed 200-item basket of five chat-evaluation benchmarks: TruthfulQA, BBQ, ToxiGen, CrowS-Pairs, and XSTest, under three prompt templates. Four of the five benchmark variants are non-canonical, and two of those are synthetic proxies. The mean absolute adjacent-generation Score Drift Rate is several times the mean of an independence-based count-level reference null. It stays in the same band when we drop a benchmark, drop a release line, switch to strict scoring, or restrict to constant-parameter-size transitions. Within this audited setup, a quoted trust score should be treated as checkpoint-bound. It should be re-measured on each materially new release rather than carried forward. Closed APIs, larger models, canonical-protocol scores, and benchmark-item-subset uncertainty are out of scope.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Southern Methodist University(南方 Methodist 大学)
- University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。