arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型在化妆品化学与皮肤健康领域的准确性与可靠性:一项基准测试研究

Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study

Amelia Liu

arXiv 2608.14631首次发表:更新:

AI 中文总结

本研究测试14个LLMs在化妆品化学与皮肤健康领域的表现,发现其总体准确性差、技术深度不足,无法可靠提供相关信息,需微调验证数据集并改进算法推理。

AI 中文摘要

随着消费者越来越多地转向AI聊天机器人获取护肤建议,大型语言模型(LLMs)在化妆品化学领域的技术准确性在很大程度上未得到评估。我们在一组结构化的化妆品化学相关主题上对14个LLMs进行了基准测试,包括特定化妆品成分的化学性质以及消费者可能感兴趣的常见化妆品场景。整个测试过程中禁用了网络搜索,以评估每个模型的内化知识而非其互联网检索能力。总体表现较差,在定量推理和结构识别任务中缺陷最为明显。尽管模型处理一般护肤问题的表现尚可,但回应始终缺乏知情消费者决策所需的技术深度。值得注意的是,与AI对话可能存在风险:听起来权威但包含技术错误的输出,与明确承认不确定性的回应相比,不太可能引发怀疑。这些发现表明,主要基于未验证的公开数据训练的通用LLMs目前并非化妆品化学信息的可靠来源。要使这些工具被视为公共使用资源,可能需要在两个方面取得进展:对经过验证的化学和皮肤病学数据集进行微调,以及对算法推理进行实质性改进。

英文摘要

As consumers increasingly turn to AI chatbots for skincare advice, the technical accuracy of Large Language Models (LLMs) in cosmetic chemistry remains largely under-evaluated. We benchmarked 14 LLMs on a structured set of topics related to cosmetic chemistry, including the chemical properties of specific cosmetic ingredients and common cosmetic scenarios that may be of interest to consumers. Web search was disabled throughout to assess each model's internalized knowledge rather than its internet retrieval capacity. Overall performance was poor, with the most pronounced deficits in quantitative reasoning and structural identification tasks. While models handled general skincare questions with reasonability, responses consistently lacked the technical depth required for informed consumer decision-making. Notably, conversation with AI can pose a risk: outputs that sound authoritative but contain technical errors are less likely to generate skepticism compared to responses that explicitly acknowledge uncertainty. These findings suggest that general-purpose LLMs, trained predominantly on unverified public data, are currently not reliable sources of cosmetic chemistry information. Progress on two fronts, fine-tuning verified chemical and dermatological datasets, and substantial improvements to algorithmic reasoning, will likely be needed before these tools can be considered as resources for public use.

Comments14 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑