arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为何GPT风格模型无法直接迁移到符号音乐:错误坐标系中的压缩

How Far Should Tokenization Go? Predictive Effectiveness and Relational Losslessness

Yi Wang

arXiv 2608.18025首次发表:更新:

发表机构

Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对GPT风格模型无法直接迁移到符号音乐的问题,提出有效性-无损性框架,通过构建预测有效且关系无损的坐标系优化标记化,实验验证了相关边界,揭示了跨模态迁移的关键在于标记化接口而非架构。

AI 中文摘要

GPT风格模型通过用可复用离散标记的有限词汇表示语言,实现了出色性能。这一成功促使符号音乐的标记化方法将和弦、动机、乐句等重复音乐结构视为类似语言标记的可复用单元。然而,标记化的优势不仅来自可复用组合,更在于压缩:有效压缩需要具备这样的坐标系,其中重复的规律性形成稳定且可预测的条件分布。因此,关键问题并非寻找更大的音乐组合,而是发现能让音乐事实变得可预测、可压缩的坐标系。我们提出了有效性-无损性框架,将标记化定义为构建预测有效且关系无损的坐标系。预测有效性原则定义了事实-标记边界:解耦和嵌套构造了暴露预测规律性的坐标接口。关系无损性原则定义了标记-状态边界:标记化在依赖上下文的关系被确定前停止,将其计算留给模型状态。受控符号音乐实验验证了这些边界。有效的坐标系构建提升了预测压缩性,而固定的关系投影则限制了上下文建模。仅序列压缩无法保证预测压缩,而保留上下文自由度能让高阶音乐组织在无需显式结构标签的情况下显现。这些结果揭示了为何GPT风格模型无法跨模态直接迁移:架构可迁移,但标记化接口不可迁移。标记化必须在保留关系自由度的同时发现有效表示,而上下文结构正是从该自由度中显现的。

英文摘要

GPT-style models have achieved remarkable success with finite vocabularies of reusable tokens, making the token interface a central component of modern sequence modeling. Symbolic music appears naturally compatible with this paradigm: it consists of discrete note events and recurring structures such as chords, motifs, and phrases. However, when tokenization moves beyond language, the interface must be specified for each domain. Existing work offers many effective designs, but no unified criterion for deciding what tokenization should represent and how far it should go. Using predictive codelength as a common criterion, we formulate the Effectiveness--Losslessness Framework to define where tokenization should begin and where it should end. The Fact--Token Boundary marks where observation-determined structure should enter the token interface, through operations such as coordinate construction. Within this interface, the resulting carrier may be reversibly recoded without changing the represented facts. The Token--State Boundary marks where tokenization should stop: relations that depend on context should remain for model-state computation rather than being fixed in advance by the tokenizer. We validate the framework through controlled multi-seed symbolic-music experiments, with an independent-corpus replication of the temporal intervention. Making musical time explicit consistently reduces predictive code and also improves pitch and duration prediction, while tonal-frame canonicalization and pitch factorization provide further gains. Fixed circle-of-fifths pitch coordinates instead increase predictive code, suggesting that imposing a fixed pitch relation before context can burden prediction. Reversible BPE substantially shortens the carrier but increases predictive codelength in every seed, showing that carrier compaction alone does not guarantee predictive gain.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑