arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TasteBench:从分子到可持续食品的感官预测多模态基准

TasteBench: Multimodal Benchmark for Sensory Prediction, from Molecules to Sustainable Foods

Anna T. Thomas, Sohum Patnaik, Caroline Cotto, Benjamin Sanchez-Lengeling

arXiv 2610.02599首次发表:更新:

发表机构

Stanford Computer Science; Food Intelligence Lab; NECTAR; University of Toronto Chemical Engineering(斯坦福大学计算机科学系; 食品智能实验室; NECTAR; 多伦多大学化学工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TasteBench是一个多模态基准,通过食品级排名和分子级分类任务评估感官预测模型,为可持续蛋白质发现提供计算筛选工具。

AI 中文摘要

可持续蛋白质发现缺乏类似分子对接或密度泛函理论的快速计算代理,而这些代理加速了药物和材料的发现。评估一种新型食品是否尝起来像其动物源性目标,需要昂贵的人类感官小组,这成为设计-构建-测试循环的瓶颈。我们引入了TasteBench,一个用于感官预测的多模态基准和隐私保护竞赛,涵盖两个任务:一个基于24个产品类别中215种植物性食品的21K+人类评估构建的食品级排名任务,产生了935对类别内排名对;以及一个支持性的分子级味觉分类任务,涵盖15K种风味分子。为了能够严格解释模型性能,我们表征了真实数据:评估者之间的评分者间一致性较低(Krippendorff's α = .077),而小组聚合排名的分半信度上限为.825,确立了在此基准上应评估机器学习系统的范围。我们评估了四种输入模态的基线;在评估者评分的相同对上,最佳模型达到了.661的成对准确率,与单个评估者的中位数(.650)相当,并在所有类别内对上达到了.683。TasteBench为衡量可持续蛋白质发现的计算筛选进展提供了评估基础设施和基线。

英文摘要

Sustainable protein discovery lacks the fast computational proxies, analogous to molecular docking or density functional theory, that accelerate drug and materials discovery. Evaluating whether a novel food tastes like its animal-based target requires expensive human sensory panels, bottlenecking the design-build-test loop. We introduce TasteBench, a multimodal benchmark and privacy-preserving competition for sensory prediction, spanning two tasks: a food-level ranking task built on 21K+ human evaluations across 215 plant-based foods in 24 product categories, yielding 935 within-category ranking pairs, and a supporting molecular-level taste classification task over 15K flavor molecules. To enable rigorous interpretation of model performance, we characterize the ground truth: inter-rater agreement among panelists is low (Krippendorff's $α= .077$), and the split-half reliability ceiling of panel-aggregated rankings is .825, establishing the range within which ML systems on this benchmark should be assessed. We evaluate baselines across four input modalities; on the same pairs panelists rated, the best model achieves .661 pairwise accuracy, competitive with the median individual panelist (.650), and .683 across all within-category pairs. TasteBench provides the evaluation infrastructure and baselines for measuring progress on computational screening for sustainable protein discovery.

CommentsFirst two authors contributed equally. Accepted to NeurIPS 2026, Evaluations & Datasets track. Code available at https://github.com/Food-Intelligence-Lab/tastebench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑