发表机构
Xinjiang University(新疆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对混合领域推理的组合泛化难题,提出RVV框架,通过领域标签路由推理过程、逐项验证约束并投票聚合答案集,自适应采样与模型组合在SCoRE 2026上取得79.4%的准确率,排名第二。
AI 中文摘要
当语言模型必须以不熟悉的方式组合熟悉的推理操作时,组合泛化仍然具有挑战性。2026年基于场景的常识推理评估(SCoRE)在训练中未出现的三个混合领域测试了这种能力,并要求模型识别每个问题的完整正确选项集。我们引入了路由-验证-投票(RVV),一个用于过程条件自洽性的框架,该框架使用无需参数更新的语言模型。路由利用提供的领域标签选择推理过程,指导模型表示和应用相关约束。验证提示模型根据这些约束评估每个选项。投票聚合完整的答案集,并为两个最频繁集之间投票数差距较小的问题分配额外样本。每个问题的样本遵循相同的特定领域过程。在官方测试集上,对每个问题的16个采样答案集进行投票,实现了74.6%的精确集准确率。自适应RVV达到77.3%,在选定领域路由上组合模型将准确率提升至79.4%。最终系统在参与系统中排名第二。这些结果支持领域特定推理过程和答案集分歧作为在混合领域推理中分配推理时计算的有用工具。
英文摘要
Compositional generalization remains challenging when language models must combine familiar reasoning operations in unfamiliar ways. The Scenario-Based Commonsense Reasoning Evaluation (SCoRE) 2026 tests this ability on three mixed domains absent from training and requires models to identify the complete set of correct options for each question. We introduce Route-Verify-Vote (RVV), a framework for procedure-conditioned self-consistency that uses language models without parameter updates. Route uses the provided domain label to select a reasoning procedure that guides the model in representing and applying the relevant constraints. Verify prompts the model to assess each option against those constraints. Vote aggregates complete answer sets and allocates additional samples to questions with a small vote-count margin between the two most frequent sets. Samples for each question follow the same domain-specific procedure. On the official test set, voting over 16 sampled answer sets per question achieves an exact-set accuracy of 74.6%. Adaptive RVV reaches 77.3%, and combining models on selected domain routes raises accuracy to 79.4%. The final system ranked second among participating systems. These results support domain-specific reasoning procedures and answer-set disagreement as useful tools for allocating inference-time computation in mixed-domain reasoning.
Comments12 pages, 8 figures, and 8 tables. Accepted for oral presentation at CCL26-Eval