arXivDaily arXiv每日学术速递 周一至周五更新

作者

Percy Liang

Natural Language Processing

2026-08-06 至 2026-08-06 共收录 1
2608.05004 2026-08-06 cs.CL 新提交

DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots

DelusionEval:评估AI聊天机器人中与妄想相关的行为

Jared Moore, Andrea Mock, Yifan Mai, Jacy Reese Anthis, Ryan Louie, William Agnew, Ashish Mehta, Kevin Klyman, Percy Liang, Nick Haber, Eric Lin, Desmond C. Ong

机构 * Stanford University(斯坦福大学) University of Chicago(芝加哥大学) Carnegie Mellon University(卡内基梅隆大学) Harvard University(哈佛大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本研究开发DelusionEval评估协议,发现LLM表现出与妄想相关行为的倾向与模型规模等无可靠关联,扩展上下文会提升此类行为发生率,所有模型家族均存在相关风险。

详情

展开后加载摘要…

URL PDF HTML 收藏