[论文] Stochastic Estimation of Transduced Language Models

## 论文概要 **研究领域**: NLP **作者**: Vésteinn Snæbjarnarson, S...

论文概要

研究领域: NLP 作者: Vésteinn Snæbjarnarson, Samuel Kiegeland, Manuel de Prada Corral 发布时间: 2026-08-28 arXiv: 2508.11365

中文摘要

转导语言模型(TLMs)将预训练的源语言模型与功能有限状态转导器组合,以在目标字符串上诱导出语言模型。计算TLM下目标前缀的概率,相当于对所有被转导器映射到以该前缀开头的目标字符串的源字符串,求其源模型概率之和。这个集合可能是指数级大或无限的。先前工作使用基于源前缀概率的计算捷径,然后用阈值剪枝束求和来近似所得和。这产生一个误差未知的下界。相反,我们进行无放回地重采样源前缀,并通过包含概率的倒数重新加权每个选定前缀。我们表明递归应用此校正给出目标前缀概率的无偏估计量,并让我们估计阈值剪枝损失的质量。我们的束求和算法扩展保留的源前缀并采样保留哪些前缀,随着更多概率质量被添加到运行估计中而减少其数量。这可以节省计算并保证运行以概率一终止。我们在百科全书文本和DNA上评估该方法,与有放回重采样的顺序蒙特卡洛基线相比。它在文本上实现了更好的计算-方差权衡,在DNA上在相同最大粒子数下误差更低。在一次DNA到氨基酸转导中,它将运行时间相对于阈值剪枝束求和减少了几个数量级,并使长目标字符串的前缀概率估计变得可行。在已发表的阅读时间分析中用无偏采样替代阈值剪枝显著降低了估计的语料库惊奇度,但已发表的结论保持不变。

原文摘要

Transduced language models (TLMs) compose a pretrained source language model with a functional finite-state transducer to induce a language model over target strings. Computing the probability of a target prefix under a TLM amounts to summing the source-model probabilities of all source strings that the transducer maps to target strings beginning with that prefix. This set can be exponentially large or infinite. Prior work uses a computational shortcut based on source prefix probabilities, then approximates the resulting sum with threshold-pruned beam summing. This produces a lower bound with unknown error. Instead, we resample source prefixes without replacement and reweight each selected prefix by the inverse of its inclusion probability. We show that applying this correction recursively…

自动采集于 2026-08-29

#论文 #arXiv #NLP #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1