发表机构
Xi’an Jiaotong University; Huawei Technologies Ltd(西安交通大学; 华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对思维链监督的局限性,提出依赖感知中间问答监督(DAIS)框架,将教师推理依据转换为阶段级记录,在多个数据集上提升了最终答案准确率,在政策合规基准测试中有显著增益,证明其作为辅助监督信号的有效性。
AI 中文摘要
思维链(CoT)监督揭示了中间推理依据,但单一的推理目标通常只优化单个推理序列,对局部结论如何支持后续决策的监督有限。我们引入了依赖感知中间问答监督(DAIS),这是一个训练框架,将经过筛选的教师推理依据转换为阶段级问答记录。每个中间记录根据该决策所需的先前状态预测局部答案,而最终答案记录保持原始任务格式;评估仅使用原始输入和可选上下文。在多个Qwen主干的GDPR、AIACT、MedQA和FOLIO数据集上,DAIS提高了平均最终答案准确率。在政策合规基准测试中,相对于最强的非DAIS基线,它实现了5.6%的最大增益和4.2%的平均增益。消融实验表明,有效的先前状态条件作用比更长的目标或额外的中间文本更有贡献,支持依赖条件中间问答作为标准最终答案推理的轻量级辅助监督信号。
英文摘要
Chain-of-thought (CoT) supervision exposes intermediate rationales, but flat rationale targets usually optimize a single reasoning sequence and provide limited supervision on how local conclusions should support later decisions. We introduce Dependency-Aware Intermediate QA Supervision (DAIS), a training-time framework that converts filtered teacher rationales into stage-level QA records. Each intermediate record predicts a local answer conditioned on the previous states needed for that decision, while the final-answer record keeps the original task format; evaluation therefore uses only the original input and optional context. Across GDPR, AIACT, MedQA, and FOLIO with multiple Qwen backbones, DAIS improves average final-answer accuracy over answer-only, flat chain-of-thought, and independent-QA baselines. On policy-compliance benchmarks, it achieves a largest gain of 5.6% and an average gain of 4.2% over the strongest non-DAIS baseline. Controlled ablations show that valid previous-state conditioning contributes beyond longer targets or additional intermediate text, supporting dependency-conditioned intermediate QA as a lightweight auxiliary supervision signal for standard final-answer inference.