arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18867cs.LGcs.CL

HindsightBench:用于时间索引语言模型决策任务中参数后见之明的黑盒行为审计协议

HindsightBench: A Black-Box Behavioral Audit Protocol for Parametric Hindsight in Time-Indexed LLM Decision Tasks

  • University College Dublin(都柏林大学学院)

机构由 AI 辅助整理,请以论文原文为准。

Haozhe Jia

中文总结 AI 辅助

研究针对大语言模型在金融决策任务中泄露参数知识的问题,提出HindsightBench黑盒行为审计协议,通过特定矩阵、探测及指标分析参数后见之明,应用于多模型获三个主要模式,揭示模型特性及审计结果与服务的关系,还发布相关数据。

中文摘要 AI 辅助

大型语言模型会将已实现结果的参数知识泄露到历史金融决策任务中。虽然其存在已确定,但用户缺乏一种低成本的方法来审计给定模型。我们提出了HindsightBench,这是一种黑盒行为审计协议,能以探测级成本(无需回测、无需对数概率、无需语料库访问)在任何时间索引的语言模型决策任务中分析参数后见之明。该协议通过一个四臂日期操纵矩阵(揭示/仅日期/掩码/移植)、双记忆探测(日期恢复;结果召回)以及六个模型指标(触发强度、移植效果、截止后安慰剂、可恢复性、行为有效知识截止以及召回准确率解离系数),带有在可识别性依赖数据时的显式门。将其应用于来自七个供应商的15个模型在258节点的复古校正宏观面板上,产生了三个主要模式:(i)日期触发反射跟踪训练生成,而非规模——在2024年从1B到70B的开放权重生成中不存在,在每个测试的2026年生成模型中存在,并且在固定的混合专家架构和3B活动参数下在一个供应商谱系(Qwen3 -> Qwen3.6)中开启;(ii)有效截止在供应商之间跨度为22个月,比供应商报告的日期提前多达八个月,使日历窗口安慰剂设计无效;(iii)审计结果对服务不具有不变性——FP8参考模型的BF16服务会破坏触发估计的稳定性,而AWQ-INT4能保持稳定性,并且供应商锁定的推理机制会使一个探测不收敛——因此该协议附带操作要求(固定量化和思维机制;披露解析器和采样策略)。我们发布了面板、冻结的预注册、带有测量美元成本的每个模型审计行、成绩单以及一键再生。

英文摘要

Large language models leak parametric knowledge of what followed a historical date into decision tasks indexed by that date -- not necessarily a lookup of the realized outcome, but knowledge of the period all the same. Existence is settled; what users lack is a cheap way to audit a given model. We present HindsightBench, a black-box audit protocol that profiles parametric hindsight in any time-indexed LLM decision task at probe-level cost (no backtests, no logprobs, no corpus access). It chains a four-arm date-manipulation matrix (revealed/date-only/masked/transplanted), dual memory probes (date recovery; outcome recall), and six metrics -- trigger strength, transplant effect, post-cutoff placebo, recoverability, behaviorally effective cutoff, and recall-accuracy dissociation -- with explicit gates where identifiability is data-dependent. Applied to 15 models from seven vendors on a 258-node vintage-correct macro panel, it yields three patterns: (i) the date-trigger reflex is not a scale phenomenon -- it tracks training recency, though what installs it is not identified here: absent across every 2024 open-weight row where it is measurable, including a 70B tier with cutoff-aligned recall propensity, present in every tested 2026-generation model, and switching on within one vendor lineage (Qwen3 -> Qwen3.6) in the same MoE family at ~3B active; (ii) effective cutoffs span 22 months across vendors and precede vendor-reported dates by up to eight months, invalidating calendar-window placebos; (iii) results are not invariant to serving -- BF16 serving of an FP8-referenced model breaks the trigger estimate's stability while AWQ-INT4 preserves it, and a provider-locked reasoning regime makes one probe non-convergent -- so the protocol pins quantization and thinking regime as part of its contract. We release the panel, preregistrations, audit rows, transcripts, and one-command regeneration.

补充信息

↑