发表机构
The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多源文本交错问题,提出证据保持所有权路由(EPOR),将所有权预测与重建解耦,结合约束解码实现精确词级分离,在UNMIXBENCH基准上显著降低词错误率。
AI 中文摘要
当归属元数据丢失时,来自多个来源的文本可能会交错合并成单一序列,例如重叠的语音转录、文档阅读流程或并发的智能体流。我们将这一挑战形式化为词级文本分离:给定一个交错的词汇流和来源数量K,恢复原始来源序列,同时精确保留每个词的出现及其在来源内的顺序。直接使用大语言模型生成分离文本可能会遗漏、重复或虚构词汇,从而违反这一精确重建目标。因此,我们提出证据保持所有权路由(EPOR),该方法将来源所有权预测与重建解耦。EPOR适配一个因果大语言模型,根据混合流和先前的路由决策来预测规范的所有权路由。在推理时,结合完成安全的约束解码与确定性的索引重建,生成结构有效的K来源分区,精确且仅一次地保留每个观察到的词出现。我们还引入了UNMIXBENCH基准,涵盖受控合成混合物、来自AMI和ICSI的基于时间戳的语音、来自ReadingBank的基于布局的文档流以及模拟的并发数字输出。在五个评估轨道中,一个4B参数的EPOR模型在微调基线中实现了最低的平均最小置换词错误率,相对于紧凑源数组生成,五个轨道的平均值降低了22.3%,并与零样本前沿大语言模型保持竞争力。这些结果表明,当词汇证据被完全观察时,将所有权推断与词汇再生分离为直接生成提供了一种可靠的替代方案。
英文摘要
Text from multiple sources can become interleaved into a single sequence when attribution metadata is lost, such as overlapping speech transcripts, document reading flows, or concurrent agent streams. We formalize this challenge as Word-Level Text Unmixing: given an interleaved lexical stream and source count K, recover the original source sequences while preserving every word occurrence and its within-source order exactly. Directly generating separated texts with LLMs can omit, duplicate, or hallucinate words, violating this exact-reconstruction objective. We therefore propose Evidence-Preserving Ownership Routing (EPOR), which decouples source-ownership prediction from reconstruction. EPOR adapts a causal LLM to predict canonical ownership routes conditioned on the mixed stream and prior routing decisions. At inference, completion-safe constrained decoding is combined with deterministic indexed reconstruction, yielding structurally valid K-source partitions that preserve every observed occurrence exactly once. We also introduce UNMIXBENCH, covering controlled synthetic mixtures, timestamp-derived speech from AMI and ICSI, layout-derived document streams from ReadingBank, and simulated concurrent digital outputs. Across five evaluation tracks, a 4B EPOR model achieves the lowest mean minimum-permutation word error rate among finetuned baselines, reducing the five-track mean by 22.3% relative to compact source-array generation and remaining competitive with zero-shot frontier LLMs. These results show that when lexical evidence is fully observed, separating ownership inference from lexical regeneration provides a reliable alternative to direct generation.
Comments34 pages, 5 figures