arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

食谱数据结构框架及其在烹饪与营养洞察中的应用

A framework for recipe data structure with applications for culinary and nutritional insights

Mansi Goel, Sumit Bhagat, Saloni Srivastava, Malav Patel, Shlok Vinodkumar Mehroliya, Ganesh Bagler

arXiv 2609.22099首次发表:更新:

发表机构

Indraprastha Institute of Information Technology Delhi (IIIT-Delhi)(德里印度信息技术学院(IIIT-Delhi))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对食谱自由文本不可计算的问题,提出结构化食谱数据框架,构建RecipeDB2数据集,实现食材解析、营养映射与饮食分类,使烹饪遗产可计算分析。

AI 中文摘要

烹饪是一个将生食材转化为美味且营养的菜肴的复杂过程,然而编码这一过程的食谱在很大程度上仍是自由文本,人类可读但无法直接计算。现有的食谱集合捕获了这些信息的片段,但没有任何共享表示能在单个可查询的模式中,将食谱的结构化食材组成、其地理文化来源及其营养概况联系起来。我们通过形式化一个食谱数据结构框架来解决这一表示空白,该框架将每个食谱分解为类型化食材实体,将这些实体锚定在参考营养数据库中,并用地理文化和饮食背景对其进行注释。我们提出了RecipeDB2,一个包含来自32个地区和99个国家的128,942个食谱、35,474种食材的结构化汇编。食材短语通过基于Transformer的命名实体模型被解析为七个烹饪属性;食材通过BERT嵌入策略链接到USDA参考表(在200种最频繁食材的人工裁定集上F1=87.90),为每个映射食材产生148个营养参数;一个随机森林分类器将34个食材类别传播到整个词汇表;一个确定性、保守的规则集为每个食谱分配一种饮食风格。通过RecipeDB2(此HTTPS URL),我们展示了一个使食谱可计算的规模化框架,将烹饪遗产(长期被视为艺术而非定量对象)转变为数据驱动的分析。

英文摘要

Cooking is a complex process that transforms raw ingredients into delicious and nutritious dishes, yet the recipes that encode this process remain largely free text; readable by people but not directly computable. Existing recipe collections capture fragments of this information, but no shared representation links a recipe's structured ingredient composition, its geo-cultural provenance, and its nutritional profile within a single queryable schema. We address this representation gap by formalizing a framework for recipe data structure that decomposes each recipe into typed ingredient entities, grounds those entities in a reference nutritional database, and annotates them with geo-cultural and dietary context. We present RecipeDB2, a structured compilation of 128,942 recipes with 35,474 ingredients from 32 regions and 99 countries. Ingredient phrases are parsed into seven culinary attributes using a transformer-based named-entity model; ingredients are linked to the USDA reference tables through a BERT embedding strategy (F1 = 87.90 on a manually adjudicated set of the 200 most frequent ingredients), yielding 148 nutritional parameters per mapped ingredient; a Random Forest classifier propagates 34 ingredient categories across the full vocabulary; and a deterministic, conservative rule set assigns each recipe a dietary style. Through RecipeDB2 (https://cosylab.iiitd.edu.in/recipedb2/), we demonstrate a scalable framework for making recipes computable, turning culinary heritage (long treated as an artistic rather than a quantitative object) into a data-driven analysis.

CommentsMain Text (11 pages, 4 figures, 3 tables); Supplementary Information (6 pages, 5 figures, 1 table)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑