用于数据工程的神经符号学:无需微调即可实现长上下文令牌减少
Neurosymbolics for Data Engineering: Achieving Long Context Token Reduction Without Finetuning
AI总结:
本文提出一种可即插即用的神经符号层,无需微调即可增强LLM逻辑推理能力,在BIRD-CRITIC等基准上平均准确率提升85%,还能将长上下文任务有效时间复杂度从O(n²)降至约O(n),减少令牌使用超50%。
AI中文摘要:
大型语言模型(LLM)越来越多地被部署用于复杂的数据工程任务,例如从自然语言生成结构化查询(Text-to-SQL)、自动化复杂电子表格操作等。然而,要最大化其效用,既需要更高的无微调准确率,也需要解决Transformer架构固有的二次时间复杂度(O(n²))带来的计算瓶颈。本文提出了一种新型可即插即用的神经符号层,旨在无缝集成到现有LLM主干中,以增强逻辑推理能力并缓解长上下文资源消耗。在推理方面,该层可立即显著提升性能,在包括BIRD-CRITIC和LiveSQLBench在内的严格基准测试中,平均准确率提升达85%,且关键在于无需任何特定任务微调或RLHF即可实现这些提升。同时,我们将该方法用于解决长上下文推理的严重计算压力:通过利用符号处理来优先排序和压缩相关上下文信息,该层可将有效令牌使用量减少50%以上,并在某些长上下文任务上将有效时间复杂度从O(n²)降至约O(n)。这种双重影响方法不仅使LLM在数据工程中更加可靠,还大幅降低了推理芯片的计算压力,使长上下文任务更易处理且成本效益更高。
英文摘要:
Large Language Models are increasingly deployed for sophisticated data engineering tasks such as generating structured queries from natural language, Text-to-SQL, and automating complex spreadsheet operations. However, maximizing their utility demands both higher finetuning-free accuracy and solutions to the computational bottleneck imposed by the Transformer architectures inherent quadratic (On2) time complexity. This paper introduces a novel drop-in neurosymbolic layer designed to seamlessly integrate into existing LLM backbones enhancing logical reasoning and mitigating long-context resource consumption. On the reasoning front, the layer immediately and significantly improves performance yielding an average accuracy increase of 85% across rigorous benchmarks including BIRD-CRITIC and LiveSQLBench, critically achieving these gains without any task specific finetuning or RLHF. Concurrently, we repurpose this approach to address the severe computational strain of long context inference. By leveraging symbolic processing to prioritize and compress relevant contextual information the layer reduces the effective token usage by over 50% and brings the effective time complexity down from O(n2) to approximately O(n) on certain long context tasks. This dual impact approach not only makes LLMs substantially more reliable for data engineering but also drastically reduces the computational pressure on inference chips, making long context tasks more manageable and cost effective.