发表机构
ETH Zürich; University of Copenhagen(苏黎世联邦理工学院; 哥本哈根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对转导语言模型提出无放回重加权的递归束求和算法,实现目标前缀概率的无偏估计,在文本、DNA等任务上较基线方法有更优性能,可支撑长目标字符串的前缀概率估计。
AI 中文摘要
转导语言模型(Transduced Language Models, TLMs)将预训练的源语言模型与功能性有限状态转导器相结合,以在目标字符串上诱导出一个语言模型。计算TLM下目标前缀的概率,相当于对转导器映射到以该前缀开头的目标字符串的所有源字符串的源模型概率进行求和,而这个集合可能呈指数级庞大甚至无穷大。现有研究采用基于源前缀概率的计算捷径,再通过阈值剪枝的束搜索求和来近似该求和,这会产生误差未知的下界。相反,我们无放回地重采样源前缀,并通过其包含概率的倒数对每个选中的前缀进行重加权。我们证明,递归应用此校正可得到目标前缀概率的无偏估计量,并能估计阈值剪枝损失的质量。我们的束求和算法会扩展保留的源前缀,并采样要保留的前缀,随着运行估计中添加更多概率质量而减少前缀数量,这可节省计算量并保证运行以概率1终止。我们在百科全书文本和DNA上,针对有放回重采样的序贯蒙特卡洛基线评估了该方法:在文本上实现了更好的计算-方差权衡,在DNA上相同最大粒子数下误差更低;在DNA到氨基酸的转导任务中,相对于阈值剪枝的束求和,其运行时间减少了数个数量级,使长目标字符串的前缀概率估计成为可能。在已发表的阅读时间分析中,用无偏采样取代阈值剪枝,大幅降低了估计的语料惊讶度,但未改变已发表的结论。
英文摘要
Transduced language models (TLMs) compose a pretrained \emph{source} language model with a functional finite-state transducer to induce a language model over \emph{target} strings. Computing the probability of a target prefix under a TLM amounts to summing the source-model probabilities of all source strings that the transducer maps to target strings beginning with that prefix. This set can be exponentially large or infinite. Prior work uses a computational shortcut based on source prefix probabilities, then approximates the resulting sum with threshold-pruned beam summing. This produces a lower bound with unknown error. Instead, we resample source prefixes without replacement and reweight each selected prefix by the inverse of its inclusion probability. We show that applying this correction recursively gives an unbiased estimator of the target prefix probability and lets us estimate the mass lost by threshold pruning. Our beam-summing algorithm extends the retained source prefixes and samples which prefixes to keep, reducing their number as more probability mass is added to the running estimate. This can save computation and guarantees that the run halts with probability one. We evaluate the method on encyclopedic text and DNA against sequential Monte Carlo baselines that resample with replacement. It achieves a better compute--variance tradeoff on text and lower error at the same maximum number of particles on DNA. On a DNA-to-amino-acid transduction, it reduces runtime by several orders of magnitude relative to threshold-pruned beam summing and makes estimating prefix probabilities for long target strings feasible. Replacing threshold pruning with unbiased sampling in a published reading-time analysis substantially lowers the estimated corpus surprisal but leaves the published conclusions unchanged.