RIS-Kernel:一种通过稀疏注意力实现长上下文语言模型推理的模型无关架构
RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention
- Federal University of Uberlândia (UFU)(乌贝兰迪亚联邦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对大语言模型全自注意力复杂度高限制长上下文分析的问题,提出RIS-Kernel模型无关架构,通过稀疏随机几何降低复杂度,经实验验证其在不同设置下的有效性及在普通CPU服务器上实现长上下文推理的可行性。
AI中文摘要:
大语言模型中的全自注意力复杂度为O(N^2),限制了长上下文文档分析且需昂贵GPU集群。Reduced Interaction Sampling(RIS)推理引擎作为模型无关架构解决此问题。它通过稀疏随机几何将自注意力复杂度降至O(N log N),且不修改权重,能适应普通内存限制。在Qwen2-1.5B-Instruct上验证,在32768个token的控制评估中,1%密度和70个集成种子的RIS-Stochastic准确率达75.00%,优于原生密集基线(71.88%);5%密度和10个种子时与之匹配。在最紧预算下,1%密度和10个种子的RIS-Structural准确率达68.75%。在65536个token时,RIS比零上下文基准提高了14.06个百分点。所有评估在普通未加速CPU服务器上运行,证明长上下文语言模型推理在无GPU加速的标准学术硬件上可行。
英文摘要:
Full self-attention in large language models scales as O(N^2), which limits long-context document analysis to 65,536 tokens and requires costly GPU clusters. The Reduced Interaction Sampling (RIS) inference engine addresses this constraint as a model-agnostic architecture. Without modifying weights, RIS reduces self-attention complexity to O(N log N) using sparse stochastic geometry that fits within commodity memory limits. We validate RIS on Qwen2-1.5B-Instruct across two regimes. In controlled evaluations at 32,768 tokens (where native dense attention serves as the upper bound), RIS-Stochastic at 1% density and 70 ensemble seeds achieves 75.00% accuracy, outperforming the native dense baseline (71.88%), while RIS-Stochastic at 5% density and 10 seeds matches it (71.88%). This demonstrates that sparse attention acts as a regularizer: low density (1%) over multiple seeds filters out sequence-level noise, whereas higher density (5%) reintroduces distractor noise. Under the tightest budget, RIS-Structural reaches 68.75% accuracy at 1% density with just 10 seeds, recovering 75% of the contextual gap relative to the zero-context floor (59.38%). At 65,536 tokens, where dense attention triggers out-of-memory faults, RIS yields retrieval gains of up to 14.06 percentage points over the zero-context floor (51.56%), which is confirmed as marginally significant under McNemar's paired test (p = 0.078 < 0.10). All evaluations run on commodity, unaccelerated CPU servers (16-128 GB of RAM), demonstrating that long-context LLM inference is feasible on standard academic hardware without GPU acceleration.