arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

食用油拉曼光谱的决策树与K-Means分析:一种物理信息人工智能方法

Decision Tree and K-Means Analysis of Raman Spectra for Edible Oils: A Physics-Informed AI Approach

Amrita Shaw, Chandrasekar S. N., Sai Muthukumar V., Jhinuk Gupta, Deepak L. N. Kallepalli

arXiv 2608.20440首次发表:更新:

AI 中文总结

本研究结合拉曼光谱与PI-AI框架,用t-SNE、K-Means等方法分析纯油及含基质油的光谱,实现食用油高精度分类,且得到的紧凑特征可大幅减少数据量,为食品质量监测提供基础。

AI 中文摘要

加工食品中食用油的鉴别对于食品质量、欺诈预防和法规合规至关重要。本研究建立了一个结合拉曼光谱与机器学习的集成框架,该框架关联了内在光谱组织、可解释分类以及物理信息人工智能(PI-AI)。研究以五种食用油为对象,分别分析其纯品状态和在炸薯片基质中的状态,采用了t-SNE、K-Means聚类、决策树以及基于非负最小二乘法(NNLS)的光谱分解等方法。无监督分析显示,纯油的类别组织性和可分性显著更强,而食品基质效应会引发明显的光谱重叠。决策树仅利用原始1866个特征的光谱空间中的四个拉曼变量,就实现了纯油100%的分类准确率;这些变量被预剪枝和后剪枝模型一致识别,仅占可用光谱信息的约0.21%,同时保留了完美的测试集性能。对于含基质的样品,基于NNLS的PI-AI光谱分解通过将油相关特征与纸张、马铃薯的贡献分离,大幅提升了分类效果。优化后的后剪枝模型在经纸张扣除和纸张加马铃薯扣除的数据集上,准确率分别达到86.4%和85.4%,同时将重要拉曼变量的数量减少至5个和4个;这种紧凑的四特征表示进一步将数据占用量降低了99.44%,且未损失分类准确率。总体而言,这些研究结果表明,基于拉曼的食用油准确识别可通过具有物理意义、高度紧凑且可解释的光谱表示实现,为节俭AI、边缘AI、便携式传感及嵌入式食品质量监测提供了有前景的基础。

英文摘要

Classification of edible oils in processed foods is important for food quality, fraud prevention, and regulatory compliance. This study develops a Mutually Exclusive, Collectively Exhaustive framework integrating spectral organization, interpretable classification, Physics-Informed Artificial Intelligence (PI-AI), and Frugal AI-based feature reduction. Five edible oils were analyzed in pure form and within a fried-potato-chip matrix using t-SNE, K-means clustering, Decision Trees, and Non-Negative Least Squares (NNLS)-based spectral decomposition. Unsupervised analyses showed stronger class organization and separability in pure oils, while food-matrix effects caused substantial spectral overlap. Decision Trees achieved 100% classification accuracy for pure oils using only four Raman variables from 1866 spectral features. These variables represented only 0.21% of the available spectral information while retaining perfect test-set performance. Two variables associated with lipid unsaturation (about 1650 cm-1) and hydrocarbon-chain organization (about 1127 cm-1) remained important after NNLS matrix correction. Their combined contribution increased from 50% in pure oils to about 62% and 89% in paper-subtracted and paper-plus-potato-subtracted datasets, respectively. NNLS-based PI-AI improved food-matrix classification by separating oil signatures from paper and potato contributions. Optimized post-pruned models achieved nearly 80% test accuracy using only five and four Raman variables, respectively. The four-feature representation reduced the data footprint by 99.44% without loss of accuracy. These findings demonstrate that Raman-based oil identification can use compact, physically meaningful, and interpretable spectral representations, supporting Frugal AI, Edge AI, portable sensing, and embedded food-quality monitoring.

Comments46 pages, 11 figures, 2 tables, 1 supplementary table, 9 supplementary figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑