发表机构
University of Tsukuba(筑波大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长文本翻译API返回不可用结果的问题,设计实现了带恢复协议的可恢复大语言模型翻译系统,该系统通过延迟发布、验证输出等方式保障翻译可用性,经测试符合多项规则要求。
AI 中文摘要
长文本翻译请求可能在API层成功,但仍会产生不可用的结果,输出可能为空、被截断、被过滤、被源文本或提示词内容主导,或在生成有价值的文本后中断。本报告描述了为具有异构输入和提供商API的已部署翻译系统开发的恢复协议:将首次可见发布延迟在64字符窗口之后,验证组装后的输出,并使用类型化流事件区分替换与续传;仅当可从源文本重新推导段落或句子前缀时,才保留中断的工作;进一步的尝试遵循稳定的模型顺序和共享截止日期,之后进入带有来源标记的回退路径。一个经过清理的配套构件实现了该协议,并通过了38项公开测试,其修复案例重现了全部14项配置的完成标签,在235个字符可见前包含4个早期无效前缀,在4个中断流中保留了31个边界安全字符,并在两个端到端场景中满足尝试、事件和来源规则。这些结果是对已发布控制流的可执行检查,而自然生成输出的翻译质量和检测器性能则需要不同的评估。
英文摘要
A long-form translation request can succeed at the API layer and still produce an unusable result. The output may be empty, truncated, filtered, dominated by source or prompt material, or interrupted after producing text worth keeping. This report describes a recovery protocol developed for a deployed translation system with heterogeneous inputs and provider APIs. It delays the first visible release behind a 64-character window, validates the assembled output, and uses typed stream events to distinguish replacement from continuation. Interrupted work is retained only when a paragraph or sentence prefix can be re-derived from the source. Further attempts follow a stable model order and a shared deadline before entering a provenance-marked fallback path. A sanitized companion artifact implements the protocol and passes 38 public tests. Its fixed cases reproduce all 14 configured completion labels, contain four early-invalid prefixes before any of their 235 characters become visible, retain 31 boundary-safe characters across four interrupted streams, and satisfy the attempt, event, and provenance rules in two end-to-end scenarios. These results are executable checks of the published control flow. Translation quality and detector performance on naturally occurring outputs require a different evaluation.
Comments9 pages, 2 figures. A sanitized reference implementation is included as ancillary material