arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23597cs.CRcs.CL

HiTMS:一个高通量多流语言隐写框架

HiTMS: A High-Throughput Multi-Stream Linguistic Steganography Framework

  • Graduate School of Informatics, Kyoto University(京都大学情报学研究科)
  • School of Cyberspace Security, Beijing University of Posts and Telecommunications(北京邮电大学网络空间安全学院)

机构由 AI 辅助整理,请以论文原文为准。

Ruiyi Yan, Zhongliang Yang, Yugo Murawaki

AI总结:

HiTMS针对现有语言隐写术单流方案的局限,提出在多轮交互响应中分配秘密,通过批处理调用提高吞吐量,用自描述框架和密钥派生调度保证可恢复性,实验显示其速度提升且降低隐写分析器AUROC,并发增加时吞吐量持续提高。

AI中文摘要:

生成式语言隐写术在大语言模型的采样随机性中隐藏秘密比特。现有方案是单流的,通过对单个提示的单个响应传达整个秘密。这带来两个限制:缺乏对批处理多流推理的协议级支持,且简单的共同批处理无法隐藏插槽占用或有效载荷完成情况。我们提出HiTMS,它在连续几轮交互中联合生成的多个响应之间分配秘密。每轮在单个批处理调用中嵌入和提取多个流,从而分摊模型调用成本并大幅提高吞吐量。为确保可恢复性,HiTMS用自描述框架包装每个响应,并采用密钥派生调度将流绑定到插槽,并用诱饵填充未使用插槽,保证精确恢复同时隐藏活动流数量。该框架与语言模型和隐写编码器无关。在八个数据集 - 模型 - 编码器设置中,八流HiTMS的嵌入和提取速度比单流基线高4.3倍,同时将隐写分析器的平均AUROC从0.681降至0.601。4到64流的额外实验表明,随着并发增加,吞吐量持续提高。

英文摘要:

Generative linguistic steganography conceals secret bits within the sampling randomness of large language models. Existing schemes are single-stream, conveying an entire secret through a single response to a single prompt. This convention incurs limitations: it provides no protocol-level support for batched multi-stream inference, and naive co-batching does not conceal slot occupancy or payload completion. We propose the High-Throughput Multi-Stream (HiTMS) framework, which distributes a secret across multiple responses produced jointly over successive rounds of interaction. Each round embeds and extracts several streams within a single batched call, thereby amortizing the cost of model invocation and substantially improving throughput. To ensure recoverability, HiTMS wraps each response in a self-describing frame and employs a key-derived schedule that binds streams to slots and fills unused slots with decoys, guaranteeing exact recovery while concealing the number of active streams. The framework is agnostic to both the language model and the steganographic coder. Across eight dataset-model-coder settings, eight-stream HiTMS achieves up to 4.3 times higher embedding and extraction speeds than single-stream baselines, while reducing the average area under the receiver operating characteristic curve (AUROC) of steganalyzers from 0.681 to 0.601. Experiments with 4 to 64 streams demonstrate sustained throughput gains as concurrency increases. GitHub repository for this work is https://github.com/ryehr/HiTMS_steganography.

↑