发表机构
Carnegie Mellon University; Tsinghua University(卡内基梅隆大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出预测性信用协议,通过配对预测和五项检查衡量科学解释对实验预测的增量贡献,在多个基准上验证,发现自然解释信用未获确认。
AI 中文摘要
研究智能体解释计划中的实验。我们通过配对预测来衡量预测性信用,这些预测共享干预、预测者和结果,同时变化描述、匹配解释和捐赠者上下文。五项检查跟踪承诺、交付、对齐、已知信号采纳和预测增益。在受控学习中的336个前瞻性状态、12个Tox21端点和24个OpenML任务中,v5的冻结信用决策是不确定的。Tox21的预注册ROC AUC区间分数危害检验未通过($D-M=-.0026$,95%区间[$-.0174$,.0104]);OpenML的联合形成、点等价和可重复性规则未满足。相对于描述,匹配的点精度增益仍未得到确认,且Tox21/OpenML种子捐赠者区间跨越零。在请求的DeepSeek V4 Pro下,匹配和捐赠者卡片将次要Tox21漂移减少了64.5%和59.1%。DeepSeek V4 Flash重放将匹配点MAE从.01823提高到.02020,并错过了匹配捐赠者区间分数等价性。OpenML全卡片分配将名义80%区间扩大了21%,覆盖率为49.3%,而描述和内容在66/144张卡片中为51.4%。直接文本Flash交付了所有144条笔记,但没有检测到匹配点精度增益。研究者撰写的机制阳性对照将点MAE相对于描述降低了2.60个百分点。该协议衡量研究智能体基准和科学预测的预测性信用;在测试的捐赠者分辨率下,自然解释信用仍未得到确认。
英文摘要
Research agents explain planned experiments. We measure predictive credit with paired forecasts sharing an intervention, forecaster, and outcome while varying description, matched explanation, and donor context. Five checks track commitment, delivery, predictive gain, alignment, and known-signal uptake. Across 336 prospective states in controlled learning, 12 Tox21 endpoints, and 24 OpenML tasks, v5's frozen credit decision was inconclusive. Tox21's preregistered ROC AUC interval-score harm test was unmet ($D-M=-.0026$, 95 percent interval [$-.0174$, .0104]); OpenML's joint formation, point-equivalence, and repeatability rule was unmet. Matched point-accuracy gains over description remained unconfirmed, and Tox21/OpenML seed-donor intervals spanned zero. Under requested DeepSeek V4 Pro, matched and donor cards reduced secondary Tox21 drift by 64.5 and 59.1 percent. A DeepSeek V4 Flash replay raised matched point MAE from .01823 to .02020 and missed matched-donor interval-score equivalence. OpenML full-card assignment widened nominal 80 percent intervals by 21 percent, with 49.3 percent coverage versus 51.4 percent for description and content in 66/144 cards. Direct-text Flash delivered all 144 notes without detectable matched point-accuracy gain. A researcher-authored mechanism positive control lowered point MAE by 2.60 percentage points versus description. The protocol measures predictive credit for research-agent benchmarks and scientific forecasting; natural-explanation credit remained unconfirmed at the tested donor resolutions.