发表机构
City University of Hong Kong(香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出无梯度框架CoRA,实现设备内任务条件检索,无需检索器微调等操作,在多数据集及树莓派5部署中展现有效性。
AI 中文摘要
设备内上下文学习(ICL)依赖推理前检索,为下游模型推理选择有用上下文的演示示例。该检索需利用任务特定信息,同时在有限计算、内存和数据暴露预算下于本地内存中运行。我们提出条件检索对齐(CoRA),一种无梯度框架,利用配对的候选输入与输出将冻结编码器转换为任务条件检索器。CoRA选择互补编码器层,从候选内存构建源自输出的条件空间,并通过闭式岭回归将候选输入表示对齐到该空间。随后低秩分解生成紧凑检索基,其中候选输出仅在离线索引构建期间使用,而查询时检索仅需查询输入和预计算索引。我们证明CoRA的秩约束基是输出条件拟合表示的最优低秩压缩,并推导了精确的两趟流式构造,避免完整拟合矩阵的实体化。我们还通过将视觉表示纳入条件和检索空间,将框架扩展至多模态示例检索。在10个文本数据集、4个多模态基准上,使用Llama-3.2-1B、MobileLLM-Pro、OpenFlamingo-3B和Qwen3.5-2B,以及端到端树莓派5部署的实验表明,CoRA支持有效的任务条件检索,无需检索器微调、反向传播或目标模型调用。
英文摘要
On-device in-context learning (ICL) relies on pre-inference retrieval to select demonstrations for useful context before downstream model inference. This retrieval must exploit task-specific information while operating over local memories under limited computation, memory, and data-exposure budgets. We propose Conditional Retrieval Alignment (CoRA), a gradient-free framework that converts a frozen encoder into a task-conditioned retriever using paired candidate inputs and outputs. CoRA selects complementary encoder layers, constructs an output-derived conditioning space from candidate memory, and aligns candidate input representations to this space through closed-form ridge regression. Low-rank factorization then produces a compact retrieval basis where candidate outputs are used only during offline index construction, whereas query-time retrieval requires only the query input and precomputed index. We show that CoRA's rank-constrained basis is the optimal low-rank compression of the output-conditioned fitted representation, and derive an exact two-pass streaming construction that avoids materializing the full fitted matrix. We further extend the framework to multimodal exemplar retrieval by incorporating visual representations into the conditioning and retrieval spaces. Experiments across ten textual datasets and four multimodal benchmarks with Llama-3.2-1B, MobileLLM-Pro, OpenFlamingo-3B, and Qwen3.5-2B, as well as end-to-end Raspberry Pi~5 deployment demonstrate that CoRA supports effective task-conditioned retrieval without retriever fine-tuning, backpropagation, or target-model calls.
CommentsUnder review