发表机构
Institute of Edible Fungi, Shanghai Academy of Agricultural Sciences(上海市农业科学院食用菌研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SaltyMeta构建了含580条肽的咸味肽基准,评估描述符与ESM2嵌入融合模型,最佳模型测试ROC-AUC达0.715,提供网络筛选工具。
AI 中文摘要
过量钠摄入仍然是公共卫生面临的主要挑战,而咸味肽和增咸肽为在低钠食品中保持感官咸味提供了一条潜在途径。然而,关于咸味肽的机器学习研究受到数据集规模小、证据标准不一、阴性标签不确定以及序列相似性泄漏等因素的制约。在此,我们提出SaltyMeta,一个精选基准和网络可访问的筛选框架,用于咸味或增咸短肽。该基准包含580条肽,包括280条阳性肽和300条阴性肽,全部标准化为2-15个残基的单字母氨基酸序列。质量控制未发现非标准残基、完全重复或正负重叠。按相似性分组划分后,保留456条肽用于训练,124条用于独立测试。我们在分组交叉验证下评估了548个可解释的肽描述符、8M、35M和150M参数规模的冻结ESM2嵌入,以及描述符-嵌入融合模型。由训练集分组交叉验证选择的传统ExtraTrees基线在交叉验证中达到ROC-AUC=0.693,在独立测试集上达到ROC-AUC=0.704、PR-AUC=0.702、F1=0.626和MCC=0.304。最佳的初始蛋白质语言模型融合模型是传统描述符加ESM2-8M嵌入,分组交叉验证ROC-AUC=0.696,测试ROC-AUC=0.700。使用PCA95降维和ExtraTrees特征重要性过滤的高级优化产生了一个实用的ESM2-8M PCA95前300模型,测试ROC-AUC=0.715和PR-AUC=0.703,尽管重复交叉验证的增益仍然有限。SaltyMeta是一个透明的优先级排序工具,而非感官验证的替代品。我们提供基准、模型、脚本、GitHub和Streamlit,用于可重复地筛选食源性肽。
英文摘要
Excess sodium intake remains a major public health challenge, while salty and saltiness-enhancing peptides offer a potential route to preserve sensory saltiness in reduced-sodium foods. Machine-learning studies of salty peptides, however, are constrained by small datasets, heterogeneous evidence standards, uncertain negative labels, and sequence similarity leakage. Here we present SaltyMeta, a curated benchmark and web-accessible screening framework for salty or saltiness-enhancing short peptides. The benchmark contains 580 peptides, including 280 positive peptides and 300 negative peptides, all standardized to 2-15 residue one-letter amino-acid sequences. Quality control found no non-standard residues, exact duplicates, or positive-negative overlaps. A similarity-grouped split retained 456 peptides for training and 124 for held-out testing. We evaluated 548 interpretable peptide descriptors, frozen ESM2 embeddings at 8M, 35M, and 150M parameter scales, and descriptor-embedding fusion models under grouped cross-validation. The traditional ExtraTrees baseline selected by training-set grouped cross-validation achieved ROC-AUC=0.693 in cross-validation and ROC-AUC=0.704, PR-AUC=0.702, F1=0.626, and MCC=0.304 on the held-out test set. The best initial protein-language-model fusion was traditional descriptors plus ESM2-8M embeddings, with grouped CV ROC-AUC=0.696 and test ROC-AUC=0.700. Advanced optimization using PCA95 dimensionality reduction and ExtraTrees feature-importance filtering yielded a practical ESM2-8M PCA95 top-300 model with test ROC-AUC=0.715 and PR-AUC=0.703, although repeated-CV gains remained modest. SaltyMeta is a transparent prioritization tool, not a sensory validation substitute. We provide benchmark, models, scripts, GitHub, and Streamlit for reproducible screening of food-derived peptides.
Comments18 pages, 8 figures, 6 tables; Corresponding author: Wanchao Chen (chenwanchao@saas.sh.cn)