发表机构
Can Tho University; Tay Do University(芹苴大学; 西原大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EXACT 2026竞赛要求用最多8B参数的自托管模型解决教育问答问题。CoTu团队开发神经符号思维程序管道,结合Z3编码、数值Python等,经答案类型路由等方法,在物理任务中获满分,决赛技术得分最高,总体第三,证明小模型也能实现可验证推理。
AI 中文摘要
透明教育问答要求答案不仅正确而且可解释,而使用小模型则排除了最大专有系统的推理能力。EXACT 2026竞赛具体提出了这个问题:最多8B参数的开放权重语言模型,自托管,每个答案都有自然语言解释。它将两项任务配对:大学规章制度的逻辑推理和多步物理问题解决。我们描述了团队CoTu开发的用于解决这两个问题的系统,这是一个神经符号思维程序管道,其中一个4B主干编写程序而不是直接陈述答案:对于规章制度查询,它发出一个Z3编码,其蕴含判决为推理提供依据,对于物理问题,它发出数值Python,两者都包含在一个共享的自我纠正循环和一个统一的解释JSON输出中。答案类型路由、基于蒸馏的任务微调以及一个延迟感知服务栈——带有推测解码的SGLang——使系统保持在每个查询60秒的限制内。该系统在物理任务的两个自动选择轮中都获得了满分,并获得了所有团队的最高决赛技术得分——13.44/15,将自动答案评估与专家判断的推理深度相结合——包括同等加权的展示得分,CoTu总体排名第三。在符号求解器中为答案提供依据可在4B规模上产生正确、可验证的推理,剩余的困难在于前提选择而非推理本身。
英文摘要
Transparent educational question answering asks for answers that are not only correct but explainable, and doing so with small models rules out the reasoning power of the largest proprietary systems. The EXACT 2026 competition poses this problem concretely: open-weight language models of at most 8B parameters, self-hosted, with a natural-language explanation for every answer. It pairs two tasks: logical reasoning over university regulations, and multi-step physics problem solving. We describe the system that team \cotu{} developed to address both, a neuro-symbolic Program-of-Thought pipeline in which a 4B backbone writes a program rather than stating an answer directly: for regulation queries it emits a Z3 encoding whose entailment verdict grounds the deduction, and for physics it emits numerical Python, both wrapped in a shared self-correction loop and a unified explained-JSON output. Answer-type routing, distillation-based task fine-tuning, and a latency-aware serving stack -- SGLang with speculative decoding -- keep the system within the 60-second per-query limit. The system achieved a \textbf{perfect score} on the physics task in both automated selection rounds and obtained the \textbf{highest final-round technical score} of any team -- $13.44/15$, combining automated answer evaluation with expert-judged reasoning depth -- with the equally weighted presentation score included, \cotu{} placed 3rd overall. Grounding answers in a symbolic solver yields correct, verifiable deductions at the 4B scale, and the residual difficulty lies in premise selection rather than the deduction itself.
CommentsThe 2nd International XAI Challenge for Transparent Educational Question-Answering @ IEEE IJCNN 2026 Competition