arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24032cs.AI

生成式人工智能证据的半衰期:40 条记录的审计、主张时效性框架以及前沿模型辅助研究的反思案例

The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research

  • Charles Sturt University(查尔斯·斯特尔特大学)

机构由 AI 辅助整理,请以论文原文为准。

Carlo Iacono

AI总结:

本文通过审计40条生成式人工智能相关记录,区分模型年龄与主张时效性并提出报告做法,还将自身生产过程作为前沿模型辅助研究创作的反思案例,展示如何让快速的人工智能辅助研究可检查,而不把模型输出当独立验证。

AI中文摘要:

生成式人工智能评估在发表前就可能成为历史,但日历年龄对每个结论的影响并不相同。本文有两个相关目的。首先,审计了2025年7月18日至2026年7月17日期间出现的40条实证记录的最大变异目的语料库。审计对发表途径、执行时间、模型标识、最新命名代或不可变快照的年龄、同系列替代和更新行为进行了编码。出现时,最新命名模型的中位数年龄为281天(中间50%:75 - 478天;范围:11 - 939天)。25篇期刊文章的中位数年龄为395天,14篇预印本为56天,1篇实验室报告为49天。35条记录包含被取代的系列,7条提供了精确日期标识符,3条明确更新了模型证据,1条增加了后期敏感性测试。所有40条都包含OpenAI系统,这是语料库的一个特征而非流行度估计。本文区分了模型年龄和主张时效性,并提出了六种报告做法。其次,将自身两天的生产过程视为前沿模型辅助研究创作的反思案例。ChatGPT中的GPT - 5.6 Sol Pro支持候选发现、来源核对、计算、起草和批判;作者检查来源、做出所有实质性决定并承担责任。这是一个实践证明,而非对生产力或质量的控制估计。通过应用自身的模型事实和模型时效性声明,本文展示了如何在不将模型输出视为独立验证的情况下使快速的人工智能辅助研究可被检查。标题隐喻性地使用了半衰期;未估计通用衰减率。

英文摘要:

Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked purposes. First, it audits a maximum-variation purposive corpus of 40 empirical records appearing between 18 July 2025 and 17 July 2026. The audit coded publication route, execution timing, model identity, age of the newest named generation or immutable snapshot, same-family supersession and refresh behaviour. At appearance, the newest named model was a median 281 days old (middle 50%: 75-478; range: 11-939). Median age was 395 days for 25 journal articles, 56 days for 14 preprints and 49 days for one laboratory report. Thirty-five records included a superseded family, seven supplied a precise dated identifier, three clearly refreshed model evidence, and one added a late sensitivity test. All 40 included an OpenAI system, a feature of this corpus rather than a prevalence estimate. The paper distinguishes model age from claim currency and proposes six reporting practices. Second, it treats its own two-day production process as a reflexive case of frontier-model-assisted research creation. GPT-5.6 Sol Pro in ChatGPT supported candidate discovery, source reconciliation, calculations, drafting and critique; the author checked sources, made all substantive decisions and accepts responsibility. This is a proof-of-practice, not a controlled estimate of productivity or quality. By applying its own Model Facts and model-currency statement, the paper shows how rapid AI-assisted research can be made inspectable without treating model output as independent validation. The title uses half-lives metaphorically; no universal decay rate is estimated.

补充信息

↑