arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AIriskEval-edu演示:教育解释中教学风险的审计

AIriskEval-edu Demo: Auditing of Pedagogical Risks in Educational Explanations

Javier Irigoyen, Roberto Daza, Francisco Jurado, Julian Fierrez, Ruben Tolosana, Alvaro Ortigosa, Miguel Lopez-Duran, Aythami Morales

arXiv 2607.25634首次发表:更新:

发表机构

BiometricsAI, Universidad Autónoma de Madrid (UAM); GHIA, Universidad Autónoma de Madrid (UAM); Universidad de Las Palmas de Gran Canaria (ULPGC)(生物识别人工智能,马德里自治大学(UAM); GHIA,马德里自治大学(UAM); 大加那利岛拉斯帕尔马斯大学(ULPGC))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AIriskEval-edu演示平台可审计教学解释质量,依据五个维度评分标准评估,集成GPT-5.5和自托管评估器,有两种运行模式,本地评估器在多数指标上优于GPT-5.5,为教育机构提供实用审计方法。

AI 中文摘要

我们展示了AIriskEval-edu演示平台,它可审计教学解释的教学质量并提供可解释的审计结果。该平台依据涵盖教学风险五个维度的评分标准评估解释,对每个维度返回二元决策和置信度得分,还包括自然语言理由及证据范围。它通过外部API集成GPT-5.5和在消费级GPU上运行的自托管Llama 3.1 8B评估器,在AIriskEval-edu数据集上微调。平台有两种模式,在AI模式下评估模拟教师生成的解释,在人类模式下实时审计用户编写的解释。本地评估器在多数指标上优于GPT-5.5,为教育机构提供了在自身基础设施内审计内容的实用方法。

英文摘要

We present AIriskEval-edu Demo, a platform that audits the pedagogical quality of instructional explanations and provides explainable audit results. The platform evaluates an explanation against a rubric covering five dimensions of pedagogical risk: factual accuracy, depth and completeness, focus and relevance, student-level appropriateness, and ideological bias. For each dimension, it returns a binary decision and a confidence score. Detected risks also include a natural-language rationale and, except for Depth and Completeness, a localized evidence span. The platform integrates GPT-5.5 through an external API and a self-hosted Llama 3.1 8B evaluator that runs on consumer-grade GPUs. The local evaluator is fine-tuned on AIriskEval-edu, a dataset of K-12 instructional explanations with risk and explainability annotations. The platform operates in two modes: in AI mode, both evaluators assess stored explanations generated under six simulated teacher profiles, each representing a distinct pedagogical behavior and potential risk; in human mode, the local evaluator audits user-written explanations in real time. The local evaluator outperforms GPT-5.5 on most reported metrics, offering educational institutions a practical way to keep audited content within their own infrastructure.

Comments6 pages, 2 figures. Accepted at the 17th IAPR International Workshop on Document Analysis Systems (DAS 2026), ICDAR 2026, September 3, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑