发表机构
University of Central Florida(中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文首次对上下文自回归秩转码隐写术(CARTS)进行严格安全分析,证明其精确正确性,并通过Llama 3 8B实验验证了载荷精确恢复和抗攻击性,为语言模型在密码学中的应用奠定基础。
AI 中文摘要
自回归语言模型可通过保留跨上下文的逐位置秩信息,将载荷文本转换为相同令牌长度的隐写文本——我们将这一方法形式化为上下文自回归秩转码隐写术(CARTS)。虽然Norelli等人的Calgacus构造在实验上证明了这一现象,但此前缺乏正式的安全性分析。本文首次对CARTS进行了严格处理。我们在确定性模型假设下证明了其精确正确性,引入了一种秩坐标表示,其中密钥在秩向量空间上充当双射,定义了相关的安全概念以及与该构造自然相关的计算问题——上下文搜索、密钥碰撞、消息歧义和编码映射的非交换性——并研究了它们之间的理论关系,包括消息歧义在上下文搜索方面的表征,以及密钥碰撞与消息歧义之间的张力。对Llama 3 8B的实证研究在所有测试案例中确认了原始载荷的精确恢复,在随机密钥生成下未发现密钥碰撞,证明手工构造的碰撞是局部而非全局的,且未发现可交换的密钥对——表明对所研究的攻击向量具有抵抗性。这项工作为语言模型在密码学和隐私保护通信中的建设性应用开辟了一个有正式基础的研究议程。
英文摘要
Autoregressive language models can be used to transform a payload text into a stegotext of identical token length by preserving per-position rank information across contexts - a methodology we formalize as Contextual Autoregressive Rank Transcoding Steganography (CARTS). While the Calgacus construction of Norelli et al. demonstrated this phenomenon experimentally, no formal security analysis existed. This paper provides the first rigorous treatment of CARTS. We show its exact correctness under deterministic model assumptions, introduce a rank-coordinate representation in which keys act as bijections on rank-vector space, define relevant security notions and the computational problems naturally associated with the construction - context search, key collisions, message equivocation, and non-commutativity of the encoding maps - and study the theoretical relationships between them, including the characterization of message equivocation in terms of context search, and the tension between key collisions and message equivocation. An empirical study on Llama 3 8B confirms exact recovery of the original payload in all tested cases, finds no key collisions under random key generation, establishes that a hand-crafted collision is local rather than global, and finds no commuting key pairs - suggesting resistance to the attack vectors studied. This work opens a formally grounded research agenda for the constructive use of language models in cryptography and privacy-preserving communication.