arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14198cs.LGcs.CL

MINT:一种用于交易数据的通用零样本预测器

MINT: A Universal Zero-Shot Predictor for Transaction Data

Parameswaran Kamalaruban, Viktor Drobnyi, Maeve Madigan, Julia Rozanova, David Sutton, Stuart Burrell

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出通用零样本预测框架MINT,通过轻量级嵌入注入等技术连接交易序列编码器与LLM,在交易预测问答任务中性能领先且资源消耗更低,证实紧凑交易嵌入更具优势。

中文摘要 AI 辅助

银行会对序列金融交易数据进行分析,以完成诸多任务,包括欺诈预防、信用风险评估和个性化推荐。为提升这些任务的预测准确率,支付基础模型(Payments Foundation Models)会将交易序列数据编码为丰富的上下文嵌入,之后可将这些嵌入作为特征提供给特定任务模型。然而,这些基础模型并非为跨新下游预测任务进行灵活的零样本推理而设计,限制了其适应性和实用性。现有的基于大型语言模型(LLM)的零样本预测方法往往无法充分利用交易数据中的预测信号,同时依赖成本高昂的文本序列化或特定任务架构,扩展性较差。为解决这些局限,我们提出了交易多模态指令网络(Multimodal Instruction Network for Transactions,MINT),该框架通过轻量级嵌入注入、交易-语言对齐和指令调优,将预训练的交易序列编码器与仅解码器结构的LLM相连。我们发现,MINT在分布内和分布外问题中均实现了最先进的预测问答性能,同时与文本序列化基线相比,大幅减少了输入令牌数、延迟和内存消耗。通过对表示、对齐策略、训练数据和历史长度的综合分析,我们证实,对于多模态推理和零样本预测任务,紧凑的交易嵌入是比文本序列化更优的交易表示方法。

英文摘要

Banks analyse sequential financial transaction data to perform many tasks, including fraud prevention, credit risk assessment and offer personalization. To improve the predictive accuracy of these tasks, Payments Foundation Models encode transaction sequence data as rich contextual embeddings, which can then be provided to task-specific models as features. However, these Foundation Models are not designed for flexible zero-shot reasoning across novel downstream prediction tasks, limiting their adaptability and utility. Existing LLM-based approaches to zero-shot prediction often fail to fully exploit the predictive signal within transaction data, while relying on costly text serialization or task-specific architectures that scale poorly. To address these limitations, we present the Multimodal Instruction Network for Transactions (MINT), a framework that connects a pretrained transaction sequence encoder to a decoder-only LLM through lightweight embedding injection, transaction-language alignment, and instruction tuning. We find that MINT achieves state-of-the-art predictive question-answering performance in both in-distribution and out-of-distribution questions, while substantially reducing input tokens, latency, and memory consumption compared to text-serialization baselines. Through comprehensive analyses of representations, alignment strategies, training data, and history length, we establish that compact transaction embeddings are a superior approach to transaction representation than text serialization for multimodal reasoning and zero-shot prediction tasks.

发表机构

  • Visa Inc.(维萨公司)

机构由 AI 辅助整理,请以论文原文为准。

↑