推理模型在识别问题上准确但不健全
Reasoning Models Are Accurate but Unsound on Identification
- Illinois Institute of Technology(伊利诺伊理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究构建CERTID管道,评估推理模型在因果效应识别中的健全性,发现准确率不能代表健全性,模型间虚假声称率差异达十七倍。
AI中文摘要:
当推理模型被问及因果效应是否可从观测数据中恢复时,其失败方式有两种:拒绝一个可识别的查询,或回答一个不可识别的查询。后者更为严重,因为没有任何观测数据能验证所声称的公式。衡量这种失败需要可证明不可识别的查询,而先前的评估缺乏这一点;同时,评分需要接受任何等价形式下的正确公式,而字符串匹配无法提供这一点。我们构建了CERTID,一个形式化识别管道,解决了这两个局限。CERTID使用健全且完备的因果识别算法ID来认证给定图和查询下的效应是否可识别,并针对干预分布已知的结构因果模型验证返回的公式。CERTID进一步开发了理论结果以减轻结构泄漏、修复不可识别查询并建立评分保证。我们在涵盖4到50个顶点的1,200个认证实例上评估了三个前沿推理模型(Gemini Flash、Gemini Pro和GPT5.5)。准确率被证明是健全性的一个糟糕代理:在相同实例上,不可识别查询的虚假声称率在模型间变化达十七倍。我们还发现,在最强模型训练快照之后生成的图上,模型以97-100%的准确率决定可识别性。实例、认证程序、验证器以及逐实例记录可在该https URL获取。
英文摘要:
A reasoning model asked whether a causal effect is recoverable from observational data can fail in two ways: it refuses an identifiable query or answers a nonidentifiable one. The latter is more consequential, as no observational data can validate the claimed formula. Measuring this failure requires queries that are provably non-identifiable, which prior evaluations lack, and grading that accepts correct formulas in any equivalent form, which string matching cannot provide. We build CERTID, a formal identification pipeline that addresses both limitations. CERTID uses the sound and complete causal identification algorithm ID to certify whether an effect is identifiable from a given graph and query, and verifies returned formulas against structural causal models whose interventional distributions are known exactly. CERTID further develops theoretical results to mitigate structural leakage, repair non-identifiable queries, and establish grading guarantees. We evaluate three frontier reasoning models (Gemini Flash, Gemini Pro, and GPT5.5) on 1,200 certified instances spanning 4 to 50 vertices. Accuracy proves a poor proxy for soundness: on identical instances, the false-claim rate on non-identifiable queries varies by seventeen-fold across models. We also find that models decide identifiability with 97-100% accuracy on graphs generated after the strongest model's training snapshot. Instances, the certification procedure, the verifier, and per-instance records are available at https://anonymous.4open.science/r/certid-D718.