纳什均衡文本:一种用于文本生成的博弈论解码框架
Nash Equilibrium Text: A Game-Theoretic Decoding Framework for Text Generation
浏览论文内容
中文总结 AI 辅助
本文提出一种基于博弈论的文本生成解码框架,将修订建模为纳什均衡,并设计纳什解码算法,在问答基准上无需微调即可显著超越更大规模的自回归模型。
中文摘要 AI 辅助
文本修订已成为大语言模型不可或缺的组成部分。本文将修订问题形式化为允许存在纳什均衡的形式:词元位置是玩家,词汇项是动作,每个玩家的效用是语言模型的对数条件概率。我们通过证明随着序列长度增长,纳什均衡的对数似然可能比自回归输出呈指数级更高来激励这一修订。我们进一步提出了纳什解码算法,该算法在给定提示条件下访问词元的联合概率时,能在O(1/ε)时间内达到ε-纳什均衡。在实践中,我们使用大语言模型的条件概率估计来运行纳什解码,并在问答基准上评估所得均衡。在CLAPNQ、PubMedQA和CoQA上,从掩码语言模型获得的纳什均衡在F1和ROUGE分数上优于大至18倍的自回归模型,且无需任何微调或重新训练,代价是额外的测试时计算。
英文摘要
Text revision has become an integral component of large language models. This paper formulates revision such that it admits a Nash equilibrium: Token positions are players, vocabulary items are actions, and each player's utility is the language model's log conditional probability. We motivate the revision by showing that Nash equilibria can have exponentially higher likelihood than autoregressive outputs as the sequence length grows. We further propose Nash decoding, an algorithm that reaches an $\varepsilon$-Nash equilibrium in $O(1/\varepsilon)$ time given access to the joint probability of tokens conditioned on a prompt. In practice, we run Nash decoding using conditional probability estimates from large language models and evaluate the resulting equilibria on question-answering benchmarks. On CLAPNQ, PubMedQA, and CoQA, Nash equilibria obtained from masked language models achieve higher F1 and ROUGE scores than autoregressive models up to $18\times$ larger, without any fine-tuning or retraining, at the cost of additional test-time computation.
发表机构
- University of Virginia(弗吉尼亚大学)
- Augusta University(奥古斯塔大学)
- Qualcomm AI Research(高通人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。