DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
DelusionEval:评估AI聊天机器人中与妄想相关的行为
Jared Moore, Andrea Mock, Yifan Mai, Jacy Reese Anthis, Ryan Louie, William Agnew, Ashish Mehta, Kevin Klyman, Percy Liang, Nick Haber, Eric Lin, Desmond C. Ong
机构
*
Stanford University(斯坦福大学)
;
University of Chicago(芝加哥大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Harvard University(哈佛大学)
;
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
基于自我报告的LLM代理能够实现通用个体模拟
Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Percy Liang, Robb Willer, Michael S. Bernstein
机构
*
Computer Science Department, Stanford University(斯坦福大学计算机科学系)
;
Department of Communication Studies, Northwestern University(西北大学传播学系)
;
Department of Communication, University of Washington(华盛顿大学传播学系)
;
Google DeepMind(谷歌DeepMind)
;
Department of Sociology, Stanford University(斯坦福大学社会学系)
;
Sciences Po(巴黎政治学院)
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
将AI代理与网络安全专家在真实世界渗透测试中进行比较
Justin W. Lin, Eliot Krzysztof Jones, Donovan Julian Jasper, Ethan Jun-shen Ho, Anna Wu, Arnold Tianyi Yang, Neil Perry, Andy Zou, Matt Fredrikson, J. Zico Kolter, Percy Liang, Dan Boneh, Daniel E. Ho
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
机器去学习并不如你所想:生成式AI政策与研究的启示
A. Feder Cooper, Christopher A. Choquette-Choo, Miranda Bogen, Kevin Klyman, Matthew Jagielski, Katja Filippova, Ken Liu, Alexandra Chouldechova, Jamie Hayes, Yangsibo Huang, Eleni Triantafillou, Peter Kairouz, Nicole Elyse Mitchell, Niloofar Mireshghallah, Abigail Z. Jacobs, James Grimmelmann, Vitaly Shmatikov, Christopher De Sa, Ilia Shumailov, Andreas Terzis, Solon Barocas, Jennifer Wortman Vaughan, danah boyd, Yejin Choi, Sanmi Koyejo, Fernando Delgado, Percy Liang, Daniel E. Ho, Pamela Samuelson, Miles Brundage, David Bau, Seth Neel, Hanna Wallach, Amy B. Cyphert, Mark A. Lemley, Nicolas Papernot, Katherine Lee
机构
*
The GenLaw Center(GenLaw中心)
;
Microsoft Research(微软研究院)
;
Stanford University(斯坦福大学)
;
Google DeepMind(谷歌DeepMind)
;
Center for Democracy & Technology(民主与科技中心)
;
Princeton(普林斯顿)
;
Google(谷歌)
;
University of Washington(华盛顿大学)
;
University of Michigan(密歇根大学)
;
Cornell Tech(康奈尔科技)
;
Cornell Law School(康奈尔法学院)
;
Cornell University(康奈尔大学)
;
Lighthouse
;
Stanford Law School(斯坦福法学院)
;
UC Berkeley(伯克利大学)
;
Independent(独立研究者)
;
Northeastern University(东北大学)
;
Harvard Business School(哈佛商学院)
;
W. Virginia University College of Law(维珍尼亚大学法学院)
CommentsWe ran experiments from mid-August to mid-September 2025, notified affected providers shortly after, and now make our findings public after a 90-day disclosure window
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
BountyBench: AI代理攻击者和防御者对现实世界网络安全系统的影响
Andy K. Zhang, Joey Ji, Celeste Menders, Riya Dulepet, Thomas Qin, Ron Y. Wang, Junrong Wu, Kyleen Liao, Jiliang Li, Jinghan Hu, Sara Hong, Nardos Demilew, Shivatmica Murgai, Jason Tran, Nishka Kacheria, Ethan Ho, Denis Liu, Lauren McLane, Olivia Bruvik, Dai-Rong Han, Seungwoo Kim, Akhil Vyas, Cuiyuanxiu Chen, Ryan Li, Weiran Xu, Jonathan Z. Ye, Prerit Choudhary, Siddharth M. Bhatia, Vikram Sivashankar, Yuxuan Bao, Dawn Song, Dan Boneh, Daniel E. Ho, Percy Liang
SpecEval: Evaluating Model Adherence to Behavior Specifications
Ahmed Ahmed, Kevin Klyman, Yi Zeng, Sanmi Koyejo, Percy Liang
机构
*
Department of Computer Science, Stanford University(计算机科学系,斯坦福大学)
;
The Bradley Department of Electrical and Computer Engineering, Virginia Tech(电气与计算机工程系,弗吉尼亚理工学院)