arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种带标准答案锚定QLoRA、任务感知混合专家与组相对RLVR的验证器引导可解释推理框架

A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR

Thi Kim Trang Vo, Nam Tien Le, Thi Kim Nguyet Vo, Minh Khang Tran, Duy Phuong Tran

arXiv 2609.05221首次发表:更新:

发表机构

University of Information Technology (UIT); Ho Chi Minh City University of Technology (HCMUT); Vietnam National University, Ho Chi Minh City; University of Economics Ho Chi Minh City (UEH)(信息科技大学; 胡志明市理工大学; 越南国家大学胡志明市分校; 胡志明市经济大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对教育问答的可解释推理问题,提出含QLoRA、混合专家、RLVR的验证器引导框架,在438个示例上提升了推理深度与可解释性,补充了答案可靠性。

AI 中文摘要

大语言模型(LLM)具备强大的推理能力,但其解释可能存在不一致、依据不足或难以验证的问题。针对透明教育问答任务,我们提出一种验证器引导的可解释推理框架,结合标准答案锚定QLoRA、任务感知符号路由与组相对RLVR。首先采用Qwen2.5-3B-Instruct模型,以权威答案为锚点的领域加权QLoRA进行适配;轻量路由器随后将逻辑问题分配给FOL/Z3验证器,将物理问题分配给感知公式与单位的符号求解器。验证器反馈进一步用于RLVR过程中的候选评估、自修正与奖励构建。候选响应从三个互补维度评估:P1为答案正确性,P2为证据或单位一致性,P3为推理深度与可解释性。推理阶段,无标准答案的自一致性聚合多个候选响应,之后可选的仅问题物理验证器执行保守的系统级修正。在438个保留示例上,RLVR使P3从50.68%提升至72.20%,而混合P1基本稳定在55.94%;自一致性将纯模型P1从48.86%提升至50.23%,符号验证提供了剩余的混合增益。这些结果表明,RLVR主要强化显式推理结构,而符号验证通过在系统层面提升答案可靠性补充神经策略。

英文摘要

Large language models (LLMs) show strong reasoning ability, but their explanations can remain inconsistent, weakly grounded, or difficult to verify. We propose a verifier-guided explainable reasoning framework for transparent educational question answering that combines gold-anchored QLoRA, task-aware symbolic routing, and group-relative RLVR. Qwen2.5-3B-Instruct is first adapted with field-weighted QLoRA supervision anchored to authoritative answers. A lightweight router then assigns logic problems to a FOL/Z3 verifier and physics problems to a formula- and unit aware symbolic solver. Verifier feedback is further used to support candidate evaluation, self-revision, and reward construction during RLVR. Candidate responses are evaluated along three complementary dimensions: P1 for answer correctness, P2 for evidence or unit consistency, and P3 for reasoning depth and explainability. At inference, gold-free self-consistency aggregates multiple candidate responses before an optional question-only physics verifier performs conservative system-level correction. On 438 held-out examples, RLVR increases P3 from 50.68% to 72.20%, while hybrid P1 remains approximately stable at 55.94%. Self-consistency improves model only P1 from 48.86% to 50.23%, with symbolic verification providing the remaining hybrid gain. These results indicate that RLVR primarily strengthens explicit reasoning structure, while symbolic verification complements the neural policy by improving answer reliability at the system level.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑