arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

压缩短文本生成中的质量断点在哪里:阶段式瓶颈定位

Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization

Alexey Gavrilov, Alan-Barsag Gazzaev, Sergey Muravyov

arXiv 2607.24176首次发表:更新:

AI 中文总结

研究压缩短文本生成中质量断点问题,通过阶段式验证协议,在TinyStories案例中分离编解码器与潜在生成器影响,发现编解码器保真度是质量瓶颈,提出可重复使用的阶段式诊断方法,为质量问题研究提供新思路。

AI 中文摘要

压缩短文本生成器可能在两个不同位置失败:编解码器在生成开始前可能丢弃信息,或者潜在生成器可能产生弱代码。若不区分这些失败模式,研究人员可能在错误组件上浪费计算资源。我们在由分层VQ-VAE-2编解码器和掩码离散扩散生成器(MDLM)构建的64比16的TinyStories案例研究中进行了可控研究。使用阶段式验证协议,在一个共享外部GPT-2评分器下分离编解码器重建保真度、潜在生成质量和辅助潜在诊断,同时报告用于几何研究的补充语义指标。在测试配置中,仅编解码器重建就使外部困惑度中位数从15.17提高到27.36(+80.4%),p95从25.10提高到98.91(+294.1%),表明主要质量损失出现在潜在生成开始之前。在相同评分器下,代码空间MDLM在实质上比令牌空间扩散更强,分别将均值、中位数和p95降低了32.9%、30.9%和36.6%。几何感知正则化改善了局部潜在代理,但在现有运行中未改善解码文本指标。贡献是方法性而非算法性的:本文为一个具体管道提出了可重复使用的阶段式诊断,并表明在这种设置下,编解码器保真度而非潜在去噪设定了实际质量上限。

英文摘要

Compressed short-text generators can fail in two different places: the codec may discard information before generation starts, or the latent generator may produce weak codes. Without separating these failure modes, researchers can spend compute improving the wrong component. We study this problem in a controlled 64-to-16 TinyStories case study built from a hierarchical VQ-VAE-2 codec and a masked discrete diffusion generator (MDLM). We use a staged validation protocol that separates codec reconstruction fidelity, latent generation quality, and auxiliary latent diagnostics under one shared external GPT-2 scorer, while reporting complementary semantic metrics for the geometry study. In the tested configuration, codec reconstruction alone raises median external perplexity from 15.17 to 27.36 (+80.4%) and p95 from 25.10 to 98.91 (+294.1%), showing that the dominant quality loss appears before latent generation begins. Under the same scorer, code-space MDLM remains materially stronger than token-space diffusion, reducing mean, median, and p95 by 32.9%, 30.9%, and 36.6%, respectively. Geometry-aware regularization improves local latent proxies but does not improve decoded-text metrics in the available runs. The contribution is methodological rather than algorithmic: the paper presents a reusable staged diagnosis for one concrete pipeline and shows that, in this setting, codec fidelity rather than latent denoising sets the practical quality ceiling.

Comments8 pages, 3 figures, 14 tables. Published in the Proceedings of FRUCT'39

Journal refProceedings of the 39th Conference of the Open Innovations Association FRUCT (FRUCT'39), 2026, pp. 69-76

DOI:10.23919/FRUCT70069.2026.11506553

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑