发表机构
University of Colorado Boulder(科罗拉多大学博尔德分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大型语言模型处理复合逻辑答案选项时的失败问题,提出了一种分解原子判断并结合整数线性规划的框架,在LOGICAL-COMMONSENSEQA和LOGICAL-SATA基准上显著提升了宏F1指标。
AI 中文摘要
大型语言模型在需根据显式逻辑算子组合原子判断的答案选项上常出现失败,即便其对单个原子的判断是正确的。我们研究由AND、OR及NEITHER/NOR连接的复合选项,引入一种将每个选项分解为原子答案并对各原子的对比假设打分的框架,使模型从未见过复合选项。随后,算子约束的整数线性规划将校准后的分数组合为单一预测。我们在LOGICAL-COMMONSENSEQA上进行评估,并引入源自SATA-Bench的阅读理解基准LOGICAL-SATA。该框架在经人工验证的LOGICAL-COMMONSENSEQA划分集上将宏F1从48.3提升至77.0,在LOGICAL-SATA上从47.0提升至75.6,在NEITHER/NOR任务上提升最大。
英文摘要
Large language models often fail when answer options require combining atomic judgments under explicit logical operators, even when they judge the individual atoms correctly. We study compound options connected by AND, OR, and NEITHER/NOR, introducing a framework that decomposes each option into atomic answers and scores contrastive hypotheses about each one, so the model never sees a compound option. An operator-constrained integer linear program then composes the calibrated scores into a single prediction. We evaluate on LOGICAL-COMMONSENSEQA and introduce LOGICAL-SATA, a reading-comprehension benchmark derived from SATA-Bench. Our framework improves Macro-F1 from 48.3 to 77.0 on the human-validated LOGICAL-COMMONSENSEQA split and from 47.0 to 75.6 on LOGICAL-SATA, with the largest gains on NEITHER/NOR.
Comments21 pages, 6 figures, 10 tables