发表机构
University of Bristol; System Holdings Limited(布里斯托大学; 系统控股有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究通过分析SBERT和DeBERTa的嵌入几何,发现内部微调的SBERT在发票分类中表现优于零-shot LLM,且特定客户发票数据可提升其泛化性能,同时揭示结构化输入对SLM无帮助的反直觉结论。
AI 中文摘要
将发票分类到正确的总账(GL)代码是财务报告和税务合规的基础。这是一项需要专业会计判断的任务,而非常规任务:正确的类别微妙地取决于采购业务的性质、供应商和发票文本。尽管人工智能正越来越多地被各行业采用以自动化包括发票分类在内的任务,但基于内部小型语言模型(SLM)的实施可同时降低成本并提高数据安全性、保密性和可解释性。我们通过首先分析小型句子转换器(SBERT)和经典SLM(DeBERTa)的预训练嵌入几何来研究这种方法。该金融语料库的句子嵌入空间整体呈各向异性,但由局部各向同性的聚类组成,将先前的标记级发现扩展到金融环境中的句子嵌入,且这些聚类与供应商身份密切相关。在单个GPU上微调的SBERT在发票分类上达到0.96的准确率,优于零-shot大语言模型(LLM)和供应商身份基线,提升了较小、具有挑战性的类别和新客户的性能。对于这一重要的泛化问题,SBERT在拥有约100个特定客户发票时达到0.9的F1值,表明内部SLM实施具有前景。将这些结果与几何分析结合显示,预训练嵌入几何与分类性能相关,并揭示了一个反直觉的发现:对人类读者有帮助的结构化输入并不能提升SLM的性能。
英文摘要
Categorising invoices into the correct General Ledger (GL) code underpins financial reporting and tax compliance. This is a skilled accounting judgement rather than a routine task: the correct category depends subtly on the nature of the purchasing business, the vendor and the invoice text. Whilst AI is increasingly being adopted across industries to automate tasks, including invoice categorisation, implementations built on in-house small language models (SLMs) can simultaneously reduce cost and improve data security, confidentiality, and interpretability. We investigate this approach by first analysing the pre-trained embedding geometry of a small sentence transformer (SBERT) and classic SLM (DeBERTa). The sentence-embedding space of this financial corpus is globally anisotropic but composed of locally isotropic clusters, extending prior token-level findings to sentence embeddings in a financial setting, and these clusters are strongly correlated with the vendor identity. SBERT fine-tuned on a single GPU reaches 0.96 accuracy on invoice classification, above both a zero-shot LLM and a vendor identity baseline, increasing performance for smaller, challenging categories and new clients. For this important generalisation problem, SBERT reaches 0.9 F1 with roughly 100 client-specific invoices, showing that an in-house SLM implementation is promising. Combining these results with geometric analysis shows that pre-trained embedding geometry is associated with classification performance and reveals a counterintuitive finding that a structured input that would help a human reader does not improve the SLM performance.
Comments22 pages, 10 figures