使用回答集编程和大语言模型进行逻辑引导的数据提取
Logic-Guided Data Extraction with Answer Set Programming and Large Language Models
浏览论文内容
中文总结 AI 辅助
研究针对大语言模型用于语义数据提取时的不足,提出结合回答集编程与大语言模型的逻辑引导数据提取框架,通过交错调用与推导减少大语言模型调用次数,提高提取质量。
中文摘要 AI 辅助
当大语言模型(LLMs)用于从非结构化文本中进行语义数据提取时,对于需要复杂组合推理和全局一致性的任务可能不可靠。本文提出了一个将基于LLM的提取与回答集编程(ASP)相结合的逻辑引导数据提取框架。LLM生成候选事实,ASP进行验证、推理、一致性检查和控制。该方法利用ASP推理确定每个阶段逻辑上可接受的谓词并指导提取查询。通过交错LLM调用和ASP推导,框架能推断隐含事实并及早检测不一致。实验表明该框架减少了LLM调用并提高了提取质量。
英文摘要
When Large Language Models (LLMs) are used for semantic data extraction from unstructured text, producing candidate relational facts from natural language, they may remain unreliable for tasks requiring complex combinatorial reasoning and global consistency. This paper proposes a logic-guided data extraction framework combining LLM-based extraction with Answer Set Programming (ASP). The LLM produces candidate facts, whereas ASP performs validation, inference, consistency checking, and control. Unlike existing pipelines that query the LLM independently for all target predicates, the proposed approach uses ASP reasoning to identify which predicates are logically admissible at each stage and to guide extraction queries. By interleaving LLM calls with ASP derivation, the framework infers logically implied facts without further extraction and detects inconsistencies early. We formalize the pipeline and prove that, under mild assumptions, it is equivalent to the baseline approach with respect to the final extracted facts, while requiring fewer LLM calls. We also introduce a caching mechanism for logic-based control queries, exploiting monotonicity of conjunctive queries over incrementally constructed fact sets to reduce solver invocations. Experiments on ASP-derived benchmarks show that the framework reduces LLM calls and improves extraction quality by mitigating spurious outputs, demonstrating the value of non-monotonic logic programming for controlled semantic extraction.
发表机构
- University of Calabria(卡拉布里亚大学)
- University of Campania Luigi Vanvitelli(坎帕尼亚路易吉·万维泰利大学)
机构由 AI 辅助整理,请以论文原文为准。