arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向共同上下文问答的提示内并行解码

Intra-Prompt Parallel Decoding for Common-Context Question Answering

Theodore Glavas, Nikhita Vedula, Dushyanta Dhyani, Antonios Valkanas, Yilun Zhu, Shervin Malmasi

arXiv 2609.05707首次发表:更新:

发表机构

Amazon.com, Inc.; McGill University; Mila; Int. Lab. Learning Systems(亚马逊公司; 麦吉尔大学; 米拉研究所; 国际学习系统实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出提示内并行解码(IPPD),在单提示内并行回答多个共享上下文问题,通过共享注意力和虚拟位置ID实现,无需微调,吞吐量提升达7倍且不损质量。

AI 中文摘要

在共同上下文问答(CCQA)任务中,多个输入问题共享一个共同上下文来作为其答案的基础。然而,大型语言模型通常使用独立的提示(prompt)来生成每个答案。虽然现有的批处理和缓存技术有助于提高并行性并减少重复计算,但问题在提示之间的分离限制了可实现的加速比,因为现代GPU在注意力阶段因内存瓶颈而未被充分利用。我们提出了提示内并行解码(IPPD),这是一种新颖的推理方法,能够在单个提示内并行回答多个共同上下文的问题。IPPD通过高效地共享注意力过程中的内存和计算来直接解决瓶颈问题,因为每个问题的下一个token都在单次推理步骤中解码。IPPD使用虚拟位置ID和注意力掩码操作来生成与标准提示相同的结果,无需微调或对LLM架构进行任何更改。由于所有并行性都发生在提示内,IPPD与批处理推理完全兼容,即使每个提示具有不同的上下文也是如此。我们的实验表明,IPPD在标准解码的基础上实现了高达7倍的有效吞吐量,且没有质量下降,并且在大多数设置中优于使用PagedAttention的前缀缓存。

英文摘要

In common-context question answering (CCQA) tasks, multiple input questions share a common context to base their answers from. However, Large Language Models typically generate each answer using an independent prompt. While existing batching and caching techniques help improve parallelism and reduce repeated computations, the separation of questions across prompts limits the achievable speedup, as modern GPUs are underutilized due to a memory bottleneck during attention. We present Intra-Prompt Parallel Decoding (IPPD), a novel inference method that answers multiple common-context questions in parallel within a single prompt. IPPD directly addresses the bottleneck by efficiently sharing both memory and computation during the attention process, as the next token for every question is decoded in a single inference step. IPPD uses virtual position IDs and attention mask manipulation to generate the same output as standard prompting without requiring fine-tuning or any changes to the LLM architecture. Since all parallelism occurs within a prompt, IPPD is fully compatible with batched inference, even when each prompt features a different context. Our experiments show that IPPD delivers up to 7X the effective throughput as standard decoding without quality degradation, and outperforms prefix caching with PagedAttention in most settings.

CommentsAccepted to EMNLP 2026 main conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑