发表机构
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI); University of Pittsburgh(穆罕默德·本·扎耶德人工智能大学; 匹兹堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建96项配对回测基准,发现LLM审计器在召回率饱和时误报率高,引入干净感知警告可将误报率从20.8%降至0.0%,并指出报告干净对照率能显著区分模型性能。
AI 中文摘要
回测审计是一个校准问题:当模型错误地标记匹配的干净策略时,高缺陷召回率并无用处。我们构建了一个包含96个配对样本的基准,其中每个有缺陷的回测都有一个干净对照,该对照在保持策略、日期、代码风格、标签和报告框架固定的同时,仅改变一个方法论细节。一个确定性评分器区分缺陷召回率、干净对照误报率、证据定位和修复相关性。在来自四个文本端点的1440次缓存审计中,主要的DeepSeek审计器在封闭式和干净感知的代码召回上达到100.0%,但开放式提示过度标记了93.8%的干净代码对照,即使在召回率饱和的情况下,干净感知的三项特异性也仅为87.5%。一个干净感知的警告将DeepSeek代码误报率从20.8%(95%置信区间11.7–34.3)降至0.0%(0.0–7.4),同时召回率不变,而预算锚定在同一提示下仍标记了48个干净对照中的38个。仅报告召回率会将这四个模型中的三个排名相同;而报告干净对照率则使它们相差79个百分点。
英文摘要
Backtest auditing is a calibration problem: high flaw recall is not useful when the model falsely flags matched clean strategies. We build a 96-item paired benchmark in which every flawed backtest has a clean control that holds strategy, dates, code style, labels, and reporting scaffold fixed while changing one methodology detail. A deterministic scorer separates flaw recall, clean-control false positives, evidence localization, and fix relevance. Over 1440 cached audits from four text endpoints, the primary DeepSeek auditor reaches 100.0\% closed and clean-aware code recall, but open prompts over-flag 93.8\% of clean code controls, and clean-aware all-three specificity is 87.5\% even where recall saturates. A clean-aware warning drops DeepSeek code false positives from 20.8\% (95\% CI 11.7--34.3) to 0.0\% (0.0--7.4) at unchanged recall, while the budget anchor still flags 38/48 clean controls under the same prompt. Reporting recall alone would rank three of these four models identically; reporting the clean-control rate separates them by 79 points.