arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PRISM-RAG:面向烟草产品与立法政策推理的多模态超图检索增强生成

PRISM-RAG: Multimodal Hypergraph Retrieval-Augmented Generation for Tobacco Product and Legislative Policy Reasoning

Manuel Serna-Aguilera, Raegan Anderes, Page Dobbs, Khoa Luu

arXiv 2609.23769首次发表:更新:

发表机构

University of Arkansas; University of Arkansas for Medical Sciences(阿肯色大学; 阿肯色大学医学科学分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对跨司法管辖区法规检索的上下文冲突问题,提出多模态超图RAG框架PRISM-RAG及数据集NicoPRISM,通过司法管辖区感知路由实现93.9%合规查询正确检索,超越现有方法。

AI 中文摘要

跨司法管辖区语义相似的成文法文本的消歧是一个现有方法未能解决的检索问题。这种跨上下文冲突可能引导生成模型基于主题相关但司法管辖区不正确的来源,自信地生成答案。烟草和尼古丁法规因美国司法管辖区而异,通常措辞相似,因此,稳健的推理需要识别哪个司法管辖区的法律适用于给定产品,而不仅仅是检索相关文本。新兴产品(如袋装尼古丁)利用模糊定义逃避监管。最先进的(SOTA)文档检索增强生成(RAG)方法难以解决这种跨上下文冲突,因此难以将图像属性(例如,丰富的属性描述)与一组相似的立法文本连接起来。我们引入了NicoPRISM(尼古丁产品与法规图像和文本监测多模态),包含161,563张图像、属性描述、涵盖13个美国司法管辖区的产品、健康和立法文档知识库,以及跨两个任务的1,495个经过验证的问答对:政策合规问答和产品知识问答。我们还提出了PRISM-RAG,一种多模态超图RAG框架,构建于图像、描述和实体之上,在索引时不使用任何LLM调用,将每个查询基于产品图像,并通过司法管辖区感知的上下文组装机制路由检索,保证所查询司法管辖区的成文法文本通过构造到达语言模型。PRISM-RAG在93.9%的政策合规查询中从正确的司法管辖区检索到段落,比标准RAG高出48.6个百分点(p<0.001),在索引时使用零次LLM调用,查询时使用一次,并在关键词、语义、司法管辖区和合规性准确性指标上与SOTA RAG框架竞争或超越。

英文摘要

The disambiguation of semantically similar statutory text across jurisdictions is a retrieval problem that existing methods do not solve. This inter-context conflict can steer generative models toward confidently produced answers grounded in topically relevant but jurisdictionally incorrect sources. Tobacco and nicotine regulations vary by US jurisdiction, often sharing similar language, thus, robust reasoning requires identifying which jurisdiction's law governs a given product, not merely retrieving relevant text. Emerging products (e.g., pouches) exploit ambiguous definitions to evade regulation. State-of-the-art (SOTA) document retrieval-augmented generation (RAG) methods struggle to address this inter-context conflict, and thus struggle to connect image attributes (e.g., rich attribute captions) to the set of similar legislation texts. We introduce NicoPRISM (Nicotine Product and Regulation Image-and-Text Surveillance Multimodal), comprising 161,563 images, attribute captions, a knowledge base of product, health, and legislative documents spanning 13 US jurisdictions, and 1,495 validated question-answer pairs across two tasks: policy compliance QA and product knowledge QA. We also propose PRISM-RAG, a multimodal hypergraph RAG framework built over images, captions, and entities without any LLM calls at index time, grounding every query in a product image and routes retrieval through a jurisdiction-aware context assembly mechanism guaranteeing that statutory text from the queried jurisdiction reaches the language model by construction. PRISM-RAG retrieves passages from the correct jurisdiction in 93.9% of policy compliance queries, a 48.6 percentage point advantage over standard RAG (p<0.001), using zero LLM calls at index time and one at query time, and is competitive with or outperforms SOTA RAG frameworks across keyword, semantic, jurisdiction-, and compliance-accuracy metrics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑