arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28875cs.CLcs.AIcs.SD

提示词层面上下文未检测到侧边级词错误率(WER)变化:对生产级口述历史语料库的预注册消融实验

No Detectable Change in Side-Level WER from Prompt-Level Context: A Preregistered Ablation on a Production Oral-History Corpus

Theodore O. Cochran, Stephanie Dodson, Keith Nore

首次发表
浏览论文内容

中文总结 AI 辅助

该预注册实验在生产级口述历史转录工具上发现,提示词层面上下文未改变侧边级WER,仅短语有小幅改善,需结合序列对齐指标评估上下文机制。

中文摘要 AI 辅助

在推理时为大型多模态模型提供上下文是一种将语音转录适配到特定领域的低成本手段,早期针对较小模型的研究报告了显著增益。本研究在实际场景中测试了该机制,即在生产级口述历史转录工具的提示条件层,基于其自身生产语料库的样本开展实验。完整的提示词层面上下文未检测到侧边级词错误率(WER)发生可测量变化,且4项预注册假设均未得到支持。该实验设计为项目内配对消融,预注册时分析代码已在确认批次评分前通过哈希值冻结;2个已公开的gpt-4o试点侧边在评分器开发阶段已提前评分。19盒磁带侧边(约10.6小时的1970-80年代退化访谈音频)通过生产代码路径在3种提示分支下重新处理,结合2种已部署的商业配置(gpt-4o-transcribe和gemini-2.5-flash),并对照操作员修正的逐字参考进行评分。对于gpt-4o-transcribe,完整上下文分支与无上下文分支的配对中位数差异为+0.6 WER点,侧边重采样区间为[-1.1, +1.0];Gemini的估计值过于不稳定,无法支持可比的否定推断。事后重运行发现,批次间流水线变异性大于确认性差异,因此无法从每个单元的一次转录中分辨出该规模的效应。实现审计验证了操作已生效,序列对齐分析发现完整上下文列出的短语有小幅改善,但幅度太小无法实质性改变侧边级WER,且Gemini同时存在未列出标记错误恶化的情况。因此,评估上下文机制需要序列对齐的术语级、插入和说话人标签测量,以及 aggregate accuracy(聚合准确率)。

英文摘要

Supplying context at inference time to a large multimodal model is an inexpensive lever for adapting speech transcription to a domain, and earlier results on smaller models reported large gains. This work tested that mechanism where it ships, in the prompt-conditioning layer of a production oral-history transcription tool, on a sample from its own production corpus. Full prompt-level context did not detectably change side-level word error rate (WER), and none of the four preregistered hypotheses was supported. The design was a within-item paired ablation, preregistered with the analysis code frozen by hash before the confirmatory batch was scored; two disclosed gpt-4o pilot sides had been scored earlier, during scorer development. Nineteen cassette sides, about 10.6 hours of degraded 1970s-80s interview audio, were reprocessed through the production code path under three prompt arms, crossed with two deployed commercial configurations, gpt-4o-transcribe and gemini-2.5-flash, and scored against operator-corrected verbatim references. For gpt-4o-transcribe the median paired difference between the full-context and no-context arms was +0.6 WER points, with a side-resampled interval of [-1.1, +1.0]; the Gemini estimates were too unstable to support a comparable negative inference. A post-hoc rerun found run-to-run pipeline variability larger than the confirmatory differences, so effects of that size cannot be resolved from one transcription per cell. An implementation audit verified the manipulation was live, and sequence-alignment analysis found a small improvement on complete context-listed phrases, too small to materially change side-level WER, and for Gemini coexisting with worsened unlisted-token error. Evaluating context mechanisms therefore requires sequence-aligned term-level, insertion, and speaker-label measures alongside aggregate accuracy.

发表机构

  • AI for Altruism(利他智能)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑