发表机构
Ant International; Zhejiang University; Dingtalk, Alibaba Group; Ant Group(蚂蚁国际; 浙江大学; 阿里巴巴集团钉钉; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出ACA-RL框架,结合MPB基准,训练模型应对缺失前提的推理任务,可询问、条件化或弃权,提升欠定问题处理能力。
AI 中文摘要
仅回答式强化学习(RL)用于训练推理模型以解决完全明确的问题,但许多现实查询会省略得出唯一答案所需的前提。在此场景下,有用的响应并非总是拒绝:模型应询问缺失的前提、基于未知量对答案进行条件化,或在无信息性条件响应可用时弃权(不执行)。我们提出Ask-Condition-Abstain强化学习(ACA-RL),一种针对该场景的数据增强RL框架。其推理图引导流程将适定问题转换为带有局部缺口标注的缺失前提训练实例;ACA-RL随后基于这些实例,对五种可观测响应行为施加结构化奖励进行训练。我们还推出缺失前提基准(MPB),一个包含274个人工验证实例的基准,涵盖数学、逻辑及现实世界文字问题。在Qwen3和Llama模型上,ACA-RL在MPB上始终表现优于基线,同时保持在适定推理任务上的竞争力。结合发布的代码、MPB及训练数据,本研究为NLP评估提供新方向:衡量模型能否识别任务是否欠定并处理不确定性,而非仅衡量其能否回答完全明确的问题。
英文摘要
Answer-only reinforcement learning (RL) trains reasoning models to solve fully specified problems, but many realistic queries omit a premise needed for a unique answer. In this setting, the useful response is not always refusal: the model should ask for the missing premise, condition its answer on the unknown quantity, or abstain when no informative conditional response is available. We present \emph{Ask-Condition-Abstain Reinforcement Learning} (ACA-RL), a data-augmented RL framework for this setting. Its reasoning-graph-guided pipeline converts well-posed problems into missing-premise training instances with localized gap annotations; ACA-RL then trains on these instances with a structured reward over five observable response behaviors. We also introduce the \emph{Missing-Premise Benchmark} (MPB), a 274-instance human-verified benchmark spanning mathematical, logical, and real-world word problems. Across Qwen3 and Llama models, ACA-RL consistently improves on MPB while preserving competitive performance on well-posed reasoning tasks. Together with the released code, MPB, and training data, this work supports a new mission for NLP evaluation: measuring whether models can recognize when a task is underdetermined and handle uncertainty, not only whether they can answer fully specified questions.
CommentsAccepted to EMNLP 2026