CALLIOPE:一种基于来源的口头评估系统与合成就绪性评估
CALLIOPE: A Source-Grounded Oral Assessment System and Synthetic Readiness Evaluation
浏览论文内容
中文总结 AI 辅助
本文介绍CALLIOPE,一个基于来源的口头评估系统,集成教学材料、录音转录、自适应提问和双提供者评分,通过合成验证确认其操作路径,但未证明评分有效性。
中文摘要 AI 辅助
使用生成式AI进行口头评估需要的不仅仅是对话界面:教育工作者必须将口头回答与其源材料、评分标准、模型输出以及随后的人工判断联系起来。本技术报告介绍了CALLIOPE,一个基于来源的口头评估系统,集成了版本化的教学材料、学习者轮次录音与转录、自适应提问、双提供者评分标准评分、教育工作者审查和可导出的证据。我们检查了自2026年9月25日起的实施和保留的合成验证记录。在两次发布运行中,使用了三个口头固定装置和一个静音对照。在最终发布演练中,口头固定装置获得的AI综合分数分别为100分、62分和8分(满分100分);两次引发了提供者分歧标记。两次运行中检索到的全部八个音频文件与其输入逐字节相同。这些观察结果确立了特定已执行路径的操作,而非评分有效性或学习收益。后来的零流量候选版本增加了绑定版本的同意检查、仅插入的首次通过评分收据以及单独授权的编码导出,并得到本地回归测试的支持,但没有进行新的完整实时研究工作流评估。我们区分了已部署功能、候选保障措施以及剩余的恢复、并发性和研究操作要求。其贡献在于一个可检查的响应到审查工作流,以及对其工程证据所确立和未确立内容的发布特定说明。未报告人类参与者结果。OpenAI Codex协助了技术验证、证据综合和手稿准备。
英文摘要
Oral assessment with generative AI requires more than a conversational interface: educators must connect a spoken response to its source material, scoring criteria, model outputs and subsequent human judgement. This technical report presents CALLIOPE, a source-grounded oral assessment system integrating versioned instructional material, learner-turn recording and transcription, adaptive questioning, two-provider rubric scoring, educator review and exportable evidence. We examine implementation and retained synthetic verification records from 25 September 2026. Three spoken fixtures and a silence control were exercised across two release runs. In the final release rehearsal, the spoken fixtures received aggregate AI scores of 100, 62 and 8 out of 100; two elicited provider-disagreement flags. All eight retrieved audio files across the two runs were byte-identical to their inputs. These observations establish operation of specific exercised paths, not scoring validity or learning gains. A later zero-traffic candidate added version-bound consent checks, insert-only first-pass rating receipts and separately authorised coded exports, supported by local regression tests but not a new full live research workflow evaluation. We distinguish deployed functionality, candidate safeguards and remaining recovery, concurrency and study-operation requirements. The contribution is an inspectable response-to-review workflow and a release-specific account of what its engineering evidence does, and does not, establish. No human-participant outcomes are reported. OpenAI Codex assisted with technical verification, evidence synthesis and manuscript preparation.
发表机构
- Singapore University of Technology and Design(新加坡科技设计大学)
机构由 AI 辅助整理,请以论文原文为准。