AI 中文总结
该研究以学术监督为案例,比较无支架的GPT-5聊天机器人基线与多模块系统ASuS,ASuS将小模型GPT-4o-mini封装在LangGraph支架中,经评估发现ASuS在各维度表现更优,提取了七种模式,挑战“大模型更好”直觉。
AI 中文摘要
大语言模型通常能对单步提示给出流畅回答,但将它们部署为领域决策系统的可靠组件则困难得多。弥合这一差距是驾驭工程的工作:围绕大语言模型核心精心构建确定性支架(符号过滤器、检索、模式类型的输入/输出、大语言模型作为评判循环、人工干预门、持久状态、审计跟踪)。我们展示了一个学术监督的案例研究,该领域结合了高风险推荐、纵向问责和结构化操作流程。我们将基线(ASA),即没有支架的GPT-5聊天机器人,与一个多模块系统(ASuS)进行比较,ASuS将小得多的GPT-4o-mini封装在一个带有符号语义检索、模式验证输出、有界重试的大语言模型作为评判、人工干预门、带有大语言模型叙述的确定性加权风险评分以及每个节点的SQLite审计跟踪的LangGraph支架中。评估标准针对六个支架机制维度(基础、可解释性、一致性、过程完整性、认知负荷、约束遵守)进行了重新调整。一项由十名评分者进行的盲法混合评估,辅以2x2模型-支架消融实验,发现ASuS尽管使用的基础模型小得多,但在每个维度上的得分都高于ASA。在十名评分者中,ASuS的合并平均分为4.08,而ASA为1.23,在配对威尔科克森检验中,十名评分者中有八名在α=0.05时拒绝原假设;完整数据在第6.4和6.7节中。消融实验证实了支架的结构贡献在很大程度上与模型无关。我们提取了七种反复出现的驾驭工程模式,并认为在可靠性、可追溯性和机构一致性比开放式流畅性更重要的情况下,驾驭工程挑战了普遍存在的“更大的模型更好”的直觉。
英文摘要
Large language models routinely produce fluent answers to single-shot prompts, yet deploying them as reliable components of a domain decision system is substantially harder. Closing this gap is the work of harness engineering: the deliberate composition of deterministic scaffolding (symbolic filters, retrieval, schema-typed I/O, LLM-as-judge loops, HITL gates, persistent state, audit trails) around an LLM core. We present a case study in academic supervision, a domain combining high-stakes recommendation, longitudinal accountability, and structured operational workflows. We compare a baseline Academic Supervision Assistant (ASA), a GPT-5 chatbot with no scaffolding, against a multi-module system, Academic Supervision System (ASuS) that wraps the much smaller GPT-4o-mini in a LangGraph harness with symbolic semantic retrieval, schema-validated outputs, LLM-as-judge with bounded retry, HITL gates, deterministic weighted risk scoring with LLM narration, and a per-node SQLite audit trail. The evaluation rubric is retargeted at six harness-mechanism dimensions (grounding, explainability, consistency, process integrity, cognitive load, constraint adherence). A blind ten-rater hybrid evaluation, supplemented by a 2 x 2 model-harness ablation, finds that ASuS, despite using a much smaller base model, outscores ASA on every dimension. Across ten raters the pooled mean for ASuS is 4.08 versus 1.23 for ASA, and 8 of 10 raters reject the null at alpha = 0.05 on a paired Wilcoxon test; full numbers are in Sections 6.4 and 6.7. The ablation confirms that the structural contributions of the harness are largely model-invariant. We extract seven recurring harness-engineering patterns and argue that where reliability, traceability, and institutional consistency matter more than open-ended fluency, harness engineering challenges the prevailing 'bigger model is better' intuition.
Comments16 pages, 4 tables, 1 figure. Code and data available at https://github.com/AkashRajSingh/Harnessing-LLMs-for-Reliable-Academic-Supervision