AI 中文总结
LoopVSR是一种循环工程框架,可让代码智能体结合端到端执行证据自动修复VSR推理流水线,在CMLR VSR系统上的表现远优于静态防护,可实现100%平均修复率。
AI 中文摘要
视觉语音识别(VSR)在音频嘈杂或不可用时,可通过唇部动作恢复语音。其多阶段推理流水线涵盖视频解码、嘴部区域提取、预处理、模型调用及解码,上游故障可能掩盖下游缺陷,因此流水线维护仍很大程度依赖预定义检查和手动调试。本文提出LoopVSR,这是一种循环工程框架,可让代码智能体利用端到端执行证据自动诊断和修复VSR推理流水线。该框架将受约束的仓库级诊断与补丁修复,和一个外部控制器相结合,该控制器会审查变更、运行实际推理,并根据故障和字符错误率(CER)接受或回滚补丁。生成的反馈循环将新观测到的异常、张量统计数据和识别错误返回给智能体,逐步揭示被上游故障掩盖的缺陷。在CMLR VSR系统上,LoopVSR修复了全部11个主要故障,平均修复率达100%,而静态防护(Static guard)仅修复了11个中的2个,平均修复率为18.13%。此外,它在7次有效迭代中解决了3个级联任务,并在包含200个视频的独立隐藏测试集上保持了修复能力。这些结果表明,LoopVSR可实现可衡量的端到端VSR推理流水线自动修复。
英文摘要
Visual speech recognition (VSR) recovers speech from lip movements when audio is noisy or unavailable. Its multi-stage inference pipeline spans video decoding, mouth-region extraction, preprocessing, model invocation, and decoding, where upstream failures can mask downstream faults. Pipeline maintenance therefore still relies largely on predefined checks and manual debugging. We propose LoopVSR, a Loop Engineering framework that enables a code agent to automatically diagnose and repair VSR inference pipelines using end-to-end execution evidence. It couples constrained repository-level diagnosis and patching with an external controller that audits changes, runs real inference, and accepts or rolls back patches using failures and character error rate (CER). The resulting feedback loop returns newly observed exceptions, tensor statistics, and recognition errors to the agent, progressively exposing faults masked by upstream failures. On the CMLR VSR system, LoopVSR repairs all 11 main faults with 100% mean recovery, whereas the Static guard repairs 2 of 11 with 18.13% mean recovery. It also resolves three cascading tasks in seven accepted iterations and preserves recovery on an independent 200-video hidden set. These results demonstrate that LoopVSR enables measurable, end-to-end automated repair of VSR inference pipelines.