arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越答案置信度:黑盒决策模型中自我知识的受控审计

Beyond Answer Confidence: A Controlled Audit of Self-Knowledge in a Black-Box Decision Model

Sharath M Shankaranarayana, Davor Runje, Jan Jannink

arXiv 2610.01006首次发表:更新:

AI 中文总结

本研究通过受控干预审计黑盒决策模型Jev的自我知识,发现其置信度在超出知识边界时失效,而针对性的是/否问题能更准确反映知识状态,并强调黑盒审计需控制表面线索。

AI 中文摘要

决策模型返回的概率旨在用于路由、弃权(不执行)和自动操作。校准使这些概率在平均意义上可用,但并不能确定低置信度反映的是偶然性还是知识缺失,也不能确定当模型超出其已知范围时置信度是否会下降。我们在决策模型Jev中审计了这一区别,使用了超过15个公共数据集和6个生成的任务族,并通过配对干预来改变固定项目所提供的信息。Jev的置信度在熟悉的封闭选择任务上是校准的,但作为知识缺失的指标却失效:在没有答案相关信息时,它对一个显著选项分配高达0.80的概率,并且在超出观察到的知识边界的新闻上,其置信度超过准确率0.21至0.33,这一差距无法通过基于早期月份的重新校准来弥合。有针对性的是/否问题能更清晰地反映案例情况:结果是否已确定(AUROC 1.00)以及证据是否充分(0.95,而同一项目上置信度为0.85)。询问Jev是否知道答案似乎能标记出捏造的实体和边界后的新闻(0.91),但在使用真实姓名或移除日期的情况下,它相比答案不确定性并无优势。因此,黑盒知识审计需要明确控制表面线索。代码:此https URL。

英文摘要

Decision models return probabilities intended for routing, abstention and automated action. Calibration makes those probabilities useful on average, but does not establish whether low confidence reflects chance or missing knowledge, nor whether confidence falls when a model moves beyond what it knows. We audit this distinction in Jev, a decision model, with over 15 public datasets and 6 generated task families, with paired interventions that vary the information supplied for a fixed item. Jev's confidence is calibrated on familiar closed-choice tasks but fails as an indicator of missing knowledge: with no answer-relevant information it assigns up to 0.80 to a salient option, and on news beyond an observed knowledge boundary it exceeds accuracy by 0.21--0.33, a gap that recalibration on earlier months does not close. Targeted yes/no questions give sharper readouts of the case: whether an outcome is settled (AUROC 1.00) and whether the evidence suffices (0.95, against 0.85 for confidence on the same items). Asking whether Jev knows the answer appears to flag fabricated entities and post-boundary news (0.91), but with realistic names or with dates removed it shows no advantage over answer uncertainty. Black-box knowledge audits therefore need explicit controls for surface cues. Code: https://github.com/Syntheme/beyond-answer-confidence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑