arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于大语言模型驱动数据处理的外挂式可验证来源

Bolt-on, Verifiable Provenance for LLM-Powered Data Processing

Yiming Lin, Sepanta Zeighami, Aditya G. Parameswaran

arXiv 2608.25210首次发表:更新:

AI 中文总结

针对LLM黑盒无法提供答案来源及可信度的问题,提出外挂式框架BLIP,可高效生成小尺寸可验证来源,在七个数据集上准确率较最优基线高30%以上且成本低。

AI 中文摘要

大语言模型(LLM)是处理数据的强大工具,但同时也是复杂的黑盒,在返回数据查询答案时,不会提供答案来源或是否可信的任何指示。我们提出了LLM驱动数据处理的来源概念。现有启发式方法(如嵌入相似度或直接询问LLM)虽能提供答案来源的部分线索,但无法保证答案可通过识别出的来源推导得出,且往往不准确。相反,我们提出可验证来源的概念,即识别输入文本的一个子集,该子集可复现与完整文本相同(或等效)的答案,并引入最小性概念,即可验证来源尽可能小。识别此类来源的朴素解决方案需用LLM检查源数据的所有可能子集,成本过高。我们提出BLIP,这是一种外挂式框架,可针对任何LLM驱动的数据处理任务、使用任何LLM高效推断出小尺寸的可验证来源。作为BLIP的一部分,我们引入八种策略,每种策略都保证能找到最小可验证来源,还引入一种自适应策略,结合各策略优势以进一步降低成本。我们还将BLIP扩展为生成多个最小可验证来源。在七个数据集上的实验表明,BLIP生成的来源始终能保证复现答案,与表现最佳的基线方法相比,准确率高出30%以上,且来源规模相当;此外,BLIP产生的成本较低,与原始数据上的原始查询成本相当。

英文摘要

Large Language Models (LLMs) are powerful tools for processing data. However, LLMs are also complex black-boxes, returning answers to queries on data, without any indication for where the answer came from or whether it is trustworthy. We introduce the notion of provenance for data processing with LLMs. While existing heuristics (such as embedding similarity or directly asking an LLM) could provide some hints for where the answer was derived, they provide no guarantees that the answer can be derived using the identified provenance, and indeed, are often incorrect. Instead, we propose the notion of verifiable provenance wherein we identify a subset of the input text that reproduces the same (or equivalent) answer as that on the complete text, and introduce the notion of minimality, where the verifiable provenance is as small as possible. To identify such a provenance, a naive solution would require checking all possible subsets of the source data with the LLM, which is prohibitively expensive. We present BLIP, a bolt-on framework for efficiently inferring a small-sized verifiable provenance for any LLM-powered data processing task, with any LLM. As part of BLIP, we introduce eight strategies, each guaranteed to find a minimal verifiable provenance, as well as an adaptive strategy that combines their strengths to reduce cost further. We further extend BLIP to produce multiple minimal verifiable provenances. Experiments on seven datasets show that the provenance generated by BLIP is always guaranteed to reproduce the answer, achieving over 30% higher accuracy than the best-performing baseline with a comparable provenance size. Moreover, BLIP incurs a low cost, comparable to the original query on the original data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑