除非你知道自己在说什么,否则别跟我“嗯,实际上”:弱前提验证会降低通用问答性能
Don't `Well, Actually' Me Unless You Know What You're Talking About: Weak Presupposition Verification Degrades General QA Performance
浏览论文内容
中文总结 AI 辅助
该研究指出假前提问答(FPQA)的弱事实核查模块会降低通用问答性能,发现FPQ性能更好的方法往往在正常问题(TPQ)上表现更差,呼吁开发适配现实场景的FPQA方法。
中文摘要 AI 辅助
假前提问答(FPQA)测试大型语言模型(LLM)识别问题中错误前提并弃权(不执行)或纠正错误前提而非强化错误假设的能力。常用方法将该任务简化为提示LLM提取前提并对每个前提进行事实核查。虽然专用基准的性能不断提升,但评估主要聚焦于含错误前提的问题(FPQ),忽略了“正常”问题(TPQ)的性能。由于许多基准中FPQ的占比远超自然出现的情况,其结果无法反映现实世界的问答性能。通过在不同模型家族、规模和基准上开展大量实验,我们发现FPQ性能更好的方法往往在TPQ上表现更差,分析显示这是弱事实核查模块也会拒绝真实前提导致的。我们希望研究结果能指导未来FPQA方法向现实场景良好泛化的方向发展。
英文摘要
False-presupposition QA (FPQA) tests LLMs on their ability to identify false presuppositions in questions and abstain or correct them rather than reinforcing false assumptions. The common approach reduces the task to prompting LLMs to extract presuppositions and fact checking each presupposition. While the performance on dedicated benchmarks keeps improving, evaluation largely focuses on questions with false presuppositions (FPQs) while ignoring the performance on ``normal'' questions (TPQs). Since many benchmarks over-represent FPQs compared to their natural occurrence, the result is that performance on these benchmarks doesn't reflect real-world QA performance. Through extensive experiments across various model families, sizes, and benchmarks, we show that methods that perform better on FPQs tend to perform worse on TPQs. Our analysis reveals this is the result of weak fact checking modules that reject also true presuppositions. We hope our findings will help guide future work toward FPQA methods that generalize well to realistic settings.
发表机构
- University of British Columbia(不列颠哥伦比亚大学)
- Vector Institute(矢量研究所)
- Amii(阿尔伯塔机器智能研究所)
- CIFAR(加拿大高级研究所)
机构由 AI 辅助整理,请以论文原文为准。