保真度并不足够:智能体数据表提取的调度级检测工具
Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction
查看机构详情
- Infineon Technologies AG(英飞凌科技股份公司)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对智能体数据表提取,发现仅靠保真度不足,提出构建调度级工具,通过检查工具调用轨迹检测故障,验证了其对植入故障的检测能力及工具层的可移植性价值。
中文摘要 AI 辅助
我们发现,某模型在从未打开数据表的情况下通过了保真度检查,这是在为内部提取服务筛选模型时发现的:结构化输出约束悄悄禁用了工具使用,但该模型仍给出了答案,且编造了源文本。只有每个工具的调用轨迹能暴露这一问题。保真度(即提取值是否匹配源数据)是智能体文档提取的标准衡量指标,它将此类运行判定为成功。因此,我们在包含三个组件的智能体基准测试中记录每一次工具调用,该基准包含25个人工整理的声明,第四个组件另有12个声明,总计37个。我们从该调度记录中构建了两种检测工具:一种基于规则的故障归因分类器,以及一种仅检查调用了哪些工具、从不检查提取值的静默故障检测器。该检测器对三个模型家族的207个干净的保真度通过的提取结果未发出任何警报,且恢复了全部50个植入故障,这些故障恰好是其规则所检查的工具被 withheld(此处保留原术语)。两种结果并不对称:第一种限制了假阳性率,第二种从结构上看是召回率,而对那些调用了工具却仍给出错误答案的运行的检测能力未被测量。第二个独立的神谕(oracle,保留原术语)是一个因果室,用于测试数据表的声明在物理测量下是否成立,它是故意部分的:仅确认仪器可验证的37个声明中的2个的可验证范围,我们还给出了其余声明无法进行物理分级的分类。在受控扰动下,保真度全程通过,而因果室的判定恰好在测量不确定性处翻转。在三个已部署的模型栈中(其中一个因服务栈而非任何能力差距而不稳定),工具层带来了可移植性和可观测性,而非准确性,且仅当文档超出上下文窗口时才体现其价值。
英文摘要
One model passed our fidelity check without ever opening the datasheet. We found it while qualifying models for an internal extraction service: a structured-output constraint had silently disabled tool use, and the model answered anyway, with fabricated source text. Only the per-tool trace exposed it. Fidelity -- whether an extracted value matches the source -- is the standard measure for agentic document extraction, and it scores that run a success. We therefore log every tool call in an agentic benchmark of 25 hand-curated claims over three components, with 12 more on a fourth, 37 in all. From that dispatch record we build two instruments: a rule-based failure-attribution classifier, and a silent-failure detector whose two rules check only which tools were called, never the extracted value. The detector raises no flag on 207 clean fidelity-passing extractions across three model families, and recovers all 50 planted faults that withhold exactly the tools its rules check. The two results are not symmetric: the first bounds the false-positive rate, the second is recall by construction, and detection power against runs that call their tools and still answer wrongly is unmeasured. A second, independent oracle, a causal chamber that tests whether the datasheet's claims hold under physical measurement, is intentionally partial: it confirms only what the apparatus can exercise, a verifiable envelope of 2 of those 37 claims, and we give a taxonomy of why the rest are not physically gradable. Under a controlled perturbation, fidelity passes throughout while the chamber verdict flips exactly at the measurement uncertainty. Across three deployed model stacks (one destabilised by its serving stack, not by any capability gap) the tool layer buys portability and observability rather than accuracy, and earns its premium only once a document outgrows the context window.