arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ArgGYM:一个用于结构化可废止推理的程序化、引擎验证基准

ArgGYM: A Procedural, Engine-Verified Benchmark for Structured Defeasible Reasoning

İbrahim Ethem Deveci, Funda Tan Çalık, Barış Deniz Sağlam, Duygu Ataman

arXiv 2609.38409首次发表:更新:

发表机构

Graduate School of Informatics; Middle East Technical University(信息学研究生院; 中东技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ArgGYM是一个程序化、引擎验证的基准,用于结构化可废止推理,包含1,440个实例和十五种课程配置,支持RLVR训练,揭示了模型在复杂推理任务中的性能差异。

AI 中文摘要

大型语言模型推理的近期进展主要由具有自动可验证奖励的基准和强化学习环境驱动,尤其是在数学、代码和形式逻辑领域。这些设置使模型准确性的评估和优化更加容易,但尚不清楚在固定问题规范和稳定评估标准下的成功在多大程度上能迁移到这些领域之外的推理。现实世界的推理通常在不完整和可修订的信息下进行:结论可能被暂时支持、被反证击败、被进一步论证恢复,或在出现更强理由时被修订。这类推理通常被称为可废止推理。我们引入了ArgGYM,一个用于结构化可废止推理的程序化基准和RLVR兼容的训练环境。ArgGYM将这种推理分解为十二个任务,并将特定任务的评分基于一个符号论证引擎,该引擎计算用于评估模型输出的形式状态。它包含一个冻结的基准,涵盖十五种课程配置中的1,440个已验证实例,两种论证偏好排序(最弱链和最后链)以及两种集合排序(精英主义和民主),而相同的生成器和验证器可以生成新的实例用于评估,以减少对静态测试集的依赖,并用于可验证奖励的训练。在冻结基准上,前沿和开放权重模型显示出截然不同的推理特征:它们可以在不解决完整任务的情况下恢复结构化答案的实质性部分,并且在具有更长依赖关系和更多交互结构的后续课程配置中性能下降。我们发布了基准、生成器和验证器,用于可复现的评估和RLVR训练。

英文摘要

Recent progress in large language model reasoning has been driven by benchmarks and reinforcement learning environments with automatically verifiable rewards, particularly in mathematics, code, and formal logic. These settings make model accuracy easier to evaluate and optimize, but it remains unclear how far success under fixed problem specifications and stable evaluation criteria transfers to reasoning outside such domains. Real-world reasoning often proceeds under incomplete and revisable information: conclusions may be supported provisionally, defeated by counter-evidence, reinstated by further arguments, or revised when stronger reasons become available. Reasoning of this kind is generally referred to as defeasible reasoning. We introduce ArgGYM, a procedural benchmark and RLVR-compatible training environment for structured defeasible reasoning. ArgGYM decomposes this reasoning into twelve tasks and grounds task-specific scoring in a symbolic argumentation engine that computes the formal states used to evaluate model outputs. It includes a frozen benchmark of 1,440 verified instances across fifteen curriculum configurations, two argument preference orderings (weakest-link and last-link), and two set orderings (elitist and democratic), while the same generators and verifiers can produce fresh instances for evaluation that reduces dependence on static test sets and for verifiable-reward training. On the frozen benchmark, frontier and open-weight models show sharply different reasoning profiles: they can recover substantial parts of structured answers without solving the complete task, and performance declines in later curriculum configurations with longer dependencies and more interacting structures. We release the benchmark, generators, and verifiers for reproducible evaluation and RLVR training.

Comments40 Pages, 16 Tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑