发表机构
AI Virtual Assistant (AVA) Lab; Georgia Institute of Technology(人工智能虚拟助手(AVA)实验室; 佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对语言模型推理,提出一类含MaLoRA和MaRA的适配器,通过在令牌和上下文级别引入选择性状态空间递归实现适应,在多个冻结主干和推理基准上提升了推理准确性。
AI 中文摘要
低秩适应引入了对每个输入相同应用的静态学习更新。该更新提供任务级适应,但未明确表示令牌级或实例级状态变化。提出了一类适配器,在两个互补粒度上引入选择性状态空间递归。在令牌级别,MaLoRA使适配器的缩放因子成为具有跨令牌循环状态的动态输入相关函数。在上下文级别,MaRA在调制语言模型生成答案之前跟踪跨段状态并选择与查询最相关的段。在三个冻结主干和两个推理基准上,该类方法提高了推理准确性。
英文摘要
Low-rank adaptation introduces a static learned update applied identically to every input. The update provides task-level adaptation but does not explicitly represent token-level or instance-level state variation. A family of adapters is proposed that introduces selective state-space control at two complementary granularities. At the token level, MaLoRA (Mamba-modulated low-rank adaptation) makes the adapter's scaling factor a dynamic input-dependent function with recurrent state across tokens, in contrast to the stateless modulators of prior work. The token-level adapter improves over low-rank adaptation. On the other hand, it differentiates tokens by structural role but not by contextual relevance, which motivates placing evidence selection at the context level. At the context level, MaRA (Mamba Retrieval Adapter) tracks cross-segment reasoning state and selects the segments most relevant to the query. State-space controlled retrieval of approximately three million parameters exceeds an eight-billion-parameter dense retriever on supporting-paragraph recall. Although base models perform poorly on the task without adaptation (14 to 25 F1), MaRA recovers the evidence relevance latent in their representations. Across three frozen backbones and two multi-hop reasoning benchmarks, the end-to-end family improves reasoning accuracy on every cell of the 3-by-2 grid, by +6.4 F1 (+10.0% relative) on average over the LoRA baseline.
CommentsAccepted to EMNLP 2026 (Main Conference). 22 pages, 5 figures, 20 tables. Code: https://github.com/atahandokme/malora-mara