arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当基准推理无法组合:AI评估中的可投射性

When benchmark inferences do not compose: Projectibility in AI evaluation

Brett Reynolds

arXiv 2607.26159首次发表:更新:

发表机构

Humber Polytechnic; University of Toronto(汉伯理工学院; 多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对AI评估中基准推理无法组合的问题,提出非组合原则,结合古德曼的竞争延伸问题与基于论证的有效性框架,通过案例和模拟开发可投射性审计以诊断基准到应用论证的衔接缺陷。

AI 中文摘要

AI基准结果很少能直接得出有意义的结论。评估者会将其推广到更多案例、解读为能力证据、外推到新任务、迁移到其他系统或场景,并结合对人工评审和下游结果的假设。以有效性为中心的方法要求每个主张都有证据支持。本文指出了一个更深层的认知问题:有根据的关联不会自动形成有根据的链条。一项研究的目标可能不是下一项研究的来源;系统、总体、结果或条件可能在衔接处发生变化;共享数据或模型谱系可能使看似独立的支持产生依赖。可投射性关注从观测到的案例到未观测到的案例的有限延伸是否有依据。古德曼提出了竞争延伸的问题;基于论证的有效性提供了测试这些问题的框架。本文的独特主张是一条非组合原则:只有当端点和假设一致,且依赖关系和不确定性被完整保留时,对相邻投射的支持才允许它们组合。一个法律研究案例显示,基准证据和部署研究各自合理但相互平行。重新分析和模拟表明,聚合稳定性会消除后续投射所需的差异。由此产生的可投射性审计可诊断从基准到应用的论证中缺乏支持的衔接点。

英文摘要

An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper makes explicit and operationalizes a problem those approaches leave to the analyst: warranted links don't automatically make a warranted chain. The target of one study may not be the source of the next; system, population, outcome, or conditions may change at the interface; and shared data or model lineage may make apparently independent support dependent. Projectibility concerns whether a bounded extension from observed to unobserved cases is warranted. Goodman supplies the problem of rival extensions; argument-based validity supplies an architecture for testing them. The contribution is an interface audit for distributed AI evidence: typed source and target descriptions, and a procedure separating endpoints that never meet from endpoints that meet while warrant fails to cross. A legal-research case shows how benchmark evidence and a deployment study can each be sound while remaining parallel. A known-truth demonstration shows why aggregate stability can erase distinctions a later projection requires. The resulting projectibility audit diagnoses unsupported joins in benchmark-to-use arguments.

Comments34 pages, 2 figures, 5 tables. v2 substantially revises Secs. 5-8 and the conclusion, adds a measured instance of factor-structure instability, and corrects a claim in Sec. 3.3 that endpoint alignment suffices for composition. Supersedes the withdrawn arXiv:2510.15236. Code: https://github.com/BrettRey/benchmark-inference-composition

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑