arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16590cs.CL

审计的挑战:大型语言模型在健康领域输出的变异性

Challenges of Auditing: Variability in Outputs of Large Language Models for Health

Yuan Pu, Yewon Chang, Furong Jia, Xunjian Yin, Jessica Ma, Ayman Ali, Monica Agrawal

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示大型语言模型在不同访问模式下输出存在系统性差异,影响健康建议评估的有效性,呼吁模型提供商支持忠实复现以进行严格审计。

中文摘要 AI 辅助

人们越来越多地使用前沿AI模型获取健康建议,但通过不同的访问模式(例如ChatGPT、ChatGPT Health、API)并采用不同的设置。在此,我们发现不同访问模式之间存在系统性差异。由于评估通常依赖API,而消费者通过聊天机器人界面交互,这些差异限制了评估的有效性。我们的研究结果强调,模型提供商迫切需要支持对消费者体验和设置进行忠实复现,以进行严格的审计。

英文摘要

People increasingly use frontier AI models for health advice, but via different access modes (e.g., ChatGPT, ChatGPT Health, APIs) with varying settings. Here, we find systematic differences across access modes. Because evaluations typically rely on APIs while consumers interact through chatbot interfaces, these discrepancies limit evaluation validity. Our findings underscore an urgent need for model providers to enable faithful replication of consumer experiences and settings for rigorous audits.

发表机构

  • Duke University(杜克大学)
  • Durham VA Health System(达勒姆退伍军人健康系统)

机构由 AI 辅助整理,请以论文原文为准。

↑