大语言模型在注册营养师考试中的准确性与一致性:提示工程与知识检索的影响
Accuracy and Consistency of LLMs in the Registered Dietitian Exam: The Impact of Prompt Engineering and Knowledge Retrieval
- iHealth Labs(iHealth实验室)
- University of California, Irvine(加州大学欧文分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文利用1050道注册营养师考试题评估GPT-4o、Claude 3.5 Sonnet和Gemini 1.5 Pro的准确性与一致性,首次系统比较ZS、CoT、CoT-SC和RAP提示策略的影响,发现不同提示与领域下性能差异显著,为营养聊天机器人选择合适模型与提示方法提供了依据。
AI中文摘要:
大语言模型(LLMs)正在从根本上改变健康与福祉领域中面向人类的应用:提升患者参与度、加速临床决策并促进医学教育。尽管最先进的LLMs已在多个对话式应用中展现出卓越性能,但在营养与饮食应用中的评估仍显不足。本文提出采用注册营养师(RD)考试,对最先进的LLMs——GPT-4o、Claude 3.5 Sonnet和Gemini 1.5 Pro——进行标准且全面的评估,考察其在营养查询中的准确性与一致性。我们的评估包含1050道RD考试题目,涵盖多个营养主题和能力水平。此外,我们首次考察了零样本(ZS)、思维链(CoT)、带自一致性的思维链(CoT-SC)以及检索增强提示(RAP)对回答准确性与一致性的影响。研究结果表明,尽管这些LLMs取得了可接受的整体表现,但其结果在不同提示和问题领域之间存在显著差异。采用CoT-SC提示的GPT-4o优于其他方法,而采用ZS的Gemini 1.5 Pro记录到最高一致性。对于GPT-4o和Claude 3.5,CoT提升了准确性,CoT-SC同时提升了准确性与一致性。RAP对GPT-4o回答专家级问题尤其有效。因此,根据能力水平和特定领域选择合适的LLM与提示技术,可以减轻饮食与营养聊天机器人中的错误和潜在风险。
英文摘要:
Large language models (LLMs) are fundamentally transforming human-facing applications in the health and well-being domains: boosting patient engagement, accelerating clinical decision-making, and facilitating medical education. Although state-of-the-art LLMs have shown superior performance in several conversational applications, evaluations within nutrition and diet applications are still insufficient. In this paper, we propose to employ the Registered Dietitian (RD) exam to conduct a standard and comprehensive evaluation of state-of-the-art LLMs, GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro, assessing both accuracy and consistency in nutrition queries. Our evaluation includes 1050 RD exam questions encompassing several nutrition topics and proficiency levels. In addition, for the first time, we examine the impact of Zero-Shot (ZS), Chain of Thought (CoT), Chain of Thought with Self Consistency (CoT-SC), and Retrieval Augmented Prompting (RAP) on both accuracy and consistency of the responses. Our findings revealed that while these LLMs obtained acceptable overall performance, their results varied considerably with different prompts and question domains. GPT-4o with CoT-SC prompting outperformed the other approaches, whereas Gemini 1.5 Pro with ZS recorded the highest consistency. For GPT-4o and Claude 3.5, CoT improved the accuracy, and CoT-SC improved both accuracy and consistency. RAP was particularly effective for GPT-4o to answer Expert level questions. Consequently, choosing the appropriate LLM and prompting technique, tailored to the proficiency level and specific domain, can mitigate errors and potential risks in diet and nutrition chatbots.