arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

验证FKG.in:大语言模型增强的印度食品知识的合理性评估

Validating FKG.in: Soundness Assessment in LLM-Augmented Indian Food Knowledge

Saransh Kumar Gupta, Armaan Shah, Lipika Dey, Partha Pratim Das, Ramesh Jain

arXiv 2608.29249首次发表:更新:

发表机构

Ashoka University; Mphasis AI and Applied Tech Lab; Koita Centre for Digital Health, Ashoka University; Institute for Future Health, UC Irvine(阿肖克大学; 美斐西斯人工智能与应用技术实验室; 阿肖克大学科伊塔数字健康中心; 加州大学欧文分校未来健康研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对LLM生成食谱存在的问题,提出半自动化评估框架,验证LLM增强的印度食品知识图谱FKG.in的合理性,该框架可推广至多语言多元文化烹饪领域,为食品知识基础设施提供支撑。

AI 中文摘要

在线烹饪生态系统中,由大语言模型(LLM)生成、修改或总结的食谱内容日益增多。这类内容虽看似合理,但可能包含虚构的食材、错误的用量或不符合文化常识的组合,限制了其在下游应用及知识图谱构建中的适用性。本文提出一种半自动化的合理性评估工作流,用于验证LLM从非正式烹饪来源提取并增强的结构化食谱数据,该工作流作为印度食品知识图谱FKG.in的一部分开发,通过结合形式语法、基于词汇的检查、统计启发式方法、基于Set Transformer的一致性建模及基于检索的验证这一多阶段流程,识别并解决常见的失败模式,包括结构不一致、语义与逻辑不连贯及与源文本的偏差。尽管在印度食谱上进行了评估,但所提方法适用于更广泛的多语言和多元文化烹饪领域。本文提供了一个实用、可审计且与应用无关的框架,用于验证LLM增强的食谱数据,从而在LLM生成内容的时代,强化机器可读食品知识基础设施的基础。

英文摘要

The online culinary ecosystem is increasingly populated by recipe content generated, modified, or summarized by Large Language Models (LLMs). While often plausible, such outputs may contain hallucinated ingredients, misrepresented quantities, or culturally implausible combinations, limiting their suitability for downstream applications and knowledge graph construction. In this paper, we present a semi-automated soundness assessment workflow for validating structured recipe data extracted and augmented by LLMs from informal culinary sources. Developed as part of FKG(.in), a knowledge graph of Indian food, the pipeline identifies and addresses common failure modes, including structural inconsistencies, semantic and logical incoherence, and deviations from the source text, through a multi-stage process combining formal grammars, vocabulary-based checks, statistical heuristics, Set Transformer-based coherence modeling, and retrieval-based verification. Although evaluated on Indian recipes, the proposed methods are applicable to broader multilingual and multicultural culinary domains. We provide a practical, auditable, and application-agnostic framework for validating LLM-augmented recipe data, thereby strengthening the foundations of machine-readable food knowledge infrastructures in the era of LLM-generated content.

Comments15 pages, 2 figures, 5 tables, 27 references

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑