发表机构
Indian Institute of Technology Madras; University of Birmingham(马德拉斯印度理工学院; 伯明翰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对动态金融环境中量化大语言模型的顺序内存编辑退化问题,提出CACHE-UK框架,将知识退化降低11%-17%,测试泛化率达28%,提升6个百分点。
AI 中文摘要
部署在动态金融环境中的大语言模型(LLMs)面临关键挑战:随着市场条件、监管政策和企业事实不断变化,需保持事实准确性。4位量化虽能实现高效部署,但严重限制了顺序内存编辑的可行性:现有方法在这种“量化稳定性危机”下会出现灾难性性能下降。我们提出CACHE-UK(面向英国金融的上下文自适应持续混合编辑器),这是专为领域特定量化大语言模型设计的稳定性感知内存编辑框架。CACHE-UK集成三个组件:将编辑限制在低秩适配器子空间的秩1 LoRA扰动机制、实现内容自适应编辑强度的金融领域优先级模块、跟踪“退化债务”以防止顺序更新中灾难性遗忘的闭环稳定性控制器。在4位量化的OpenLLaMA-3B模型及含88021份文档的精选英国金融语料库上评估,CACHE-UK在相同4位约束下,相较于适配基线将知识退化降低了11%-17%(这是其最显著的效果),同时达到本实验设置中观测到的最高测试成功(泛化)率(28%,较最强适配基线提升6个百分点)。这些结果表明,稳定性感知编辑可提升资源受限金融大语言模型部署中的事实维护能力,不过绝对泛化率仍然较低。
英文摘要
Large Language Models (LLMs) deployed in dynamic financial environments face a critical challenge: maintaining factual accuracy as market conditions, regulations, and corporate facts change continuously. While 4-bit quantization enables efficient deployment, it severely limits the viability of sequential memory editing: existing methods undergo catastrophic performance degradation under this "quantization stability crisis." We introduce CACHE-UK (Contextual Adaptive Continual Hybrid Editor for UK Finance), a stability-aware memory editing framework specifically designed for domain-specific, quantized LLMs. CACHE-UK integrates three components: a rank-1 LoRA perturbation mechanism that confines edits to the low-rank adapter subspace, a financial domain prioritization module for content-adaptive edit strength, and a closed-loop Stability Controller that tracks "degradation debt" to prevent catastrophic forgetting across sequential updates. Evaluated on a 4-bit quantized OpenLLaMA-3B model with a curated UK financial corpus of 88,021 documents, CACHE-UK reduces knowledge degradation by 11-17% relative to adapted baselines under identical 4-bit constraints -- its most robust effect -- while attaining the highest test success (generalization) rate observed in our setting (28%, a 6 percentage point improvement over the strongest adapted baseline). These results indicate that stability-aware editing can improve factual maintenance in resource-constrained financial LLM deployments, though absolute generalization rates remain low.
Comments10 pages, 12 figures