发表机构
IIIT Naya Raipur(印度国际信息技术学院新赖布尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
谬误检测基准的“有效”类主要包含非图式匹配反例,导致低误报率是构建假象;使用图式匹配反例时误报率大幅上升,表明分类器仅识别论证图式而非检测谬误,并发布Scheme Foils数据集以警示。
AI 中文摘要
谬误检测基准将谬误类别与一个单一的“有效”或“无”类别配对,该类别包含了数据收集过程中未被标记为谬误的所有内容。这种构建方式具有误导性:分类器可以学习到在该类别上表现良好的线索,而无需学会区分谬误与正确论证。我们表明,基准报告的低误报率是该类别构建方式的人为产物,而非检测能力的证据。对谬误而言,信息量最大的反例是使用相同论证图式的正确论证,而在我们检查的四个基准中,此类论证在有效类别中最多只占几个百分点。在构建的图式匹配反例上进行评估时,误报率在CoCoLoFa上从16.6%上升至58.9%,在Reddit上从5.7%上升至62.0%。该比率取决于反例的编写方式,因此我们还比较了同一流程中仅图式身份不同的两种条件。分类器将图式匹配的反例标记为源谬误类型的频率比错误图式反例高40.9个百分点,而错误图式反例被识别为其实际使用的图式的比例为85.9%,被识别为源类型的比例仅为0.4%。分类器学习的是论证使用了哪种图式,而非其使用是否正确,并且在基准自身的测试集上,这两者是无法区分的。同样的分离现象也出现在三个从未见过这些基准的零样本LLM检测器中,并且在刻意构建的反例类别上,测量结果要低得多。我们将这些项目发布为Scheme Foils。在有效类别经过图式匹配覆盖审计之前,不应将报告的误报率视为检测能力的度量。
英文摘要
Fallacy-detection benchmarks pair fallacy classes with a single "valid" or "none" class that takes everything data collection did not label as a fallacy. A detector has two jobs, deciding whether an argument is fallacious and naming which fallacy it commits, and the false-positive rate is meant to measure the first. We show that what these benchmarks actually score is scheme recognition, the ability behind the second job. Their own test sets already show it: when a classifier misses a fallacy, the error lands on "none" rather than on another fallacy type, so detection is failing while classification holds. The reason is what the valid class lacks. The negatives that separate the two jobs are correct arguments using the same argumentation scheme as a fallacy, and they are scarce: nearly absent from the four benchmarks we examined, and rare even under deliberate search. A detector is therefore never tested where recognizing a scheme and judging its use come apart, and can pass on recognition alone. We construct the missing arguments, together with a control condition from the same pipeline that differs only in scheme, so whatever generation contributes, it contributes to both. The classifier labels the scheme-matched negatives as the source fallacy, and labels the wrong-scheme negatives as the scheme they actually use 85.9% of the time and as the source type 0.4%. The classifier has learned which scheme an argument uses, not whether it uses it correctly. The over-flagging follows: a model that scores 16.6% on CoCoLoFa's own valid class flags 58.9% of the constructed arguments. The same dissociation appears in three zero-shot LLM detectors that never saw these benchmarks. We release the items as Scheme Foils. A reported false-positive rate should not be trusted as a measure of detection until the valid class has been audited for scheme-matched coverage.
Comments13 pages. v2: reordered results and abstract to foreground the scheme-recognition finding; no changes to data or numbers. Data: https://github.com/fine2006/the-concealment-hypothesis