arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

写作时少读:流式多模态解码器的闭式带宽旋钮

Reading Less While Writing: A Closed-Form Bandwidth Dial for Streaming Multimodal Decoders

Yasir Mehmood, Kashif Javed

arXiv 2609.20845首次发表:更新:

发表机构

University of Engineering and Technology (UET)(工程技术大学(UET))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出ZENDAYA调度,用单一参数γ闭式控制流式解码器源读取量,使离线与实时统一,实验证明少读源可提升文本质量,并在三个语料库上显著优于固定调度。

AI 中文摘要

将视频或音频转换为文本的解码器通常会在输出一个词之前消耗整个输入。离线时,这仅仅是超出任务所需;而实时时,这是不可能的,因为字幕不能等待比赛结束。流式系统会附加一个固定规则,如 wait-$k$,即在每个词之前等待相同数量的输入令牌,无论输入的长度或节奏如何。我们用 ZENDAYA 取代固定偏移量,这是一种由单个连续参数 $\gamma$ 控制的调度。它使可见源前缀成为生成进度的闭式函数,并按输入自身的预测长度缩放,因此普通离线解码器和实时流式解码器成为一个家族的两个端点,而不是单独的模型。同一标量以闭式形式固定每个输出词消耗的源的平均比例,$\bar{E}(\gamma) \approx 1/(1+\gamma)$,使其既是延迟旋钮又是可解释的预算。我们证明了一个结构依赖定理:在任何预先固定且非递减的调度下,任何输出的令牌都不能依赖于尚未到达的输入。该保证对训练和未训练的权重均成立,并扩展到任意异步到达下的无界流。实证结果违反直觉:看到更少可以产生更好的文本,因为当模型自身输出最少时,源信息的涌入恰好稀释了注意力。从零开始训练,跨越两种模态和三个公共语料库(Charades-STA、ActivityNet Captions、LibriHeavy),一个紧凑的 29M 参数解码器在读取更少源的情况下匹配或超越固定调度,在最低延迟下收益最显著,而固定偏移量在此崩溃。流式 METEOR 增益在所有三个语料库上均具有统计显著性。

英文摘要

A decoder that turns video or audio into text conventionally consumes the entire input before emitting a word. Offline this is merely more than the task requires; live it is impossible, since a caption cannot wait for a match to end. Streaming systems bolt on a fixed rule such as wait-$k$, which waits for the same number of input tokens before every word, regardless of the input's length or pace. We replace the fixed offset with ZENDAYA, a schedule governed by a single continuous parameter $γ$. It makes the visible source prefix a closed-form function of generation progress, scaled by the input's own predicted length, so an ordinary offline decoder and a real-time streaming decoder become two endpoints of one family rather than separate models. The same scalar fixes, in closed form, the mean fraction of source consumed per emitted word, $\bar{E}(γ) \approx 1/(1+γ)$, making it at once a latency dial and an interpretable budget. We prove a structural dependency theorem: under any schedule fixed in advance and non-decreasing, no emitted token can depend on input that has not yet arrived. The guarantee holds for trained and untrained weights alike, and extends to unbounded streams under arbitrary asynchronous arrival. The empirical result is counterintuitive: seeing less can produce better text, because a flood of source dilutes attention exactly when the model has the least of its own output to anchor on. Trained from scratch across two modalities and three public corpora (Charades-STA, ActivityNet Captions, LibriHeavy), a compact 29M-parameter decoder matches or beats the fixed schedule while reading less of the source, with the sharpest gains at the lowest latencies, where a fixed offset collapses. Streaming METEOR gains are statistically significant on all three corpora.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑