AI 中文总结
本研究提出ObserverBench基准框架,用于评估指导干预、控制等任务的内部估计器的充分性,通过理论与实验揭示估计准确性与动作损失的差异,为可解释性方法评估提供支撑。
AI 中文摘要
机制可解释性正越来越多地被用于指导激活调控、电路移除、安全监测等干预操作。然而,平均而言准确的内部估计器仍可能选择糟糕的动作。我们提出ObserverBench,这是一个基准框架,用于测试内部估计器(即观察者)对于其所指导的干预、控制或安全任务是否足够。每个任务均固定模型、信息边界、允许的动作、决策规则、保留的案例和损失。该基准将估计准确性与所选动作造成的损失分开报告。理论与实验表明为何两者都必不可少:在闭环控制中,观察者误差在起点以及允许干预可到达的方向上很重要;在GPT-2-small和Qwen2.5-7B的电路干预任务中,成对观察者对未见效应的预测更准确,但并不总能选择更好的动作,而在动作损失上训练的观察者会选择更低损失的动作;在安全分诊中,当违规具有不同成本时,能完美区分违规的分数可能会分配固定的干预预算不佳。在Qwen2.5-7B、Gemma-2-9B-it以及前瞻性冻结的Qwen3.5-9B APPS任务中,AUROC对监控器的排名与部署损失不同,且最佳信息源因模型而异;在公开的激活密度或检查点不匹配情况下,稀疏SAE读出在报告的Qwen面板上也落后于其层匹配的密集对照。ObserverBench提供固定的任务契约、可运行的基线以及基于表格的提交方式,用于通过可实现的动作评估可解释性方法。
英文摘要
Mechanistic interpretability is increasingly used to guide interventions such as activation steering, circuit removal, and safety monitoring. Yet an internal estimate that is accurate on average can still choose a poor action. We present ObserverBench, a benchmark framework for testing whether an internal estimator---an observer---is adequate for the intervention, control, or safety task it directs. Each task fixes the model, information boundary, allowed actions, decision rule, held-out cases, and loss. The benchmark reports estimation accuracy separately from the loss caused by the chosen action. Theory and experiments show why both are needed. In closed-loop control, observer errors matter at the starting point and along directions the allowed intervention can reach. On circuit-intervention tasks in GPT-2-small and Qwen2.5-7B, pairwise observers predict unseen effects more accurately without always choosing better actions; observers trained on action loss choose lower-loss actions. In safety triage, a score that perfectly separates violations can allocate a fixed intervention budget poorly when violations have different costs. Across Qwen2.5-7B, Gemma-2-9B-it, and prospectively frozen Qwen3.5-9B APPS tasks, AUROC can rank monitors differently from deployment loss, and the best information source changes across models. Sparse SAE readouts also trail their layer-matched dense controls on the reported Qwen panels, under disclosed activation-density or checkpoint mismatches. ObserverBench provides fixed task contracts, runnable baselines, and table-based submissions for evaluating interpretability methods through the actions they enable.
Comments28 pages, 4 figures. Code, benchmark, leaderboards, and submission interface: https://kwisatzh.github.io/observerbench/. Frozen artifact release: https://doi.org/10.5281/zenodo.22136091