Reliability of Large Language Model Generated Clinical Reasoning in Assisted Reproductive Technology: Blinded Comparative Evaluation Study
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
机构 * Independent Researcher(独立研究者)
专题命中 安全评测 :RLHF(abstract);分类 cs.CL
Comments 38 pages. Manuscript submitted for review to the Journal of Computational Literary Studies (JCLS)
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
专题命中 安全评测 :safety(abstract);分类 cs.LG
Comments NeurIPS 2025 Track on Datasets and Benchmarks
机构 * Center for Data Science, AAIS, Peking University(数据科学中心,AAIS,北京大学) ; Wangxuan Institute of Computer Technology, Peking University(王轩计算机技术研究所,北京大学) ; Tsinghua University(清华大学) ; State Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
机构 * New York University(纽约大学)
专题命中 安全评测 :alignment(abstract)
专题命中 安全评测 :safety(abstract)
机构 * Imperial College London(伦敦帝国学院) ; University of Manchester(曼彻斯特大学) ; HiThink Research(HiThink研究院) ; Soochow University(苏州大学) ; Hong Kong Baptist University(香港 Baptist大学) ; Idiap Research Institute(Idiap研究 institute) ; Meta AI ; Fudan University(复旦大学)
专题命中 安全评测 :alignment(abstract)
Comments Project: https://github.com/HiThink-Research/NEXUS-O
专题命中 安全评测 :alignment(abstract)
专题命中 安全评测 :alignment(abstract)
专题命中 安全评测 :alignment(abstract)
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Carnegie Mellon University(卡内基梅隆大学) ; Squirrel Ai Learning
专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.AI、cs.LG
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI
专题命中 AI治理与伦理 :safety(abstract)
专题命中 AI治理与伦理 :alignment(abstract)
Comments Accepted by CIKM 2024
机构 * Seoul National University(首尔国立大学) ; Oracle(Oracle公司) ; Yonsei University(延世大学) ; Intel Labs(英特尔实验室)
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI
Comments Accepted at NeurIPS 2025 (poster). This is the camera-ready version
机构 * Universidade Estadual de Campinas(坎皮纳斯州立大学)
专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI、cs.LG
机构 * Lehigh University(莱维理工大学) ; University of Texas at Dallas(德克萨斯大学达拉斯分校) ; University of Pittsburgh(匹兹堡大学)
专题命中 其他安全 :alignment(title);safety(abstract)
机构 * Fudan University(复旦大学) ; National University of Singapore(新加坡国立大学) ; Singapore Management University(新加坡管理学院)
专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI
机构 * Rutgers University(罗格斯大学) ; Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学) ; University of Science and Technology of China(中国科学技术大学) ; Meta AI ; North Carolina State University(北卡罗来纳州立大学)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG
机构 * University of Sheffield(谢菲尔德大学) ; Hitachi, Ltd.(日立株式会社) ; University of Exeter(埃克塞特大学)
专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI
Comments Accepted to TMLR
机构 * Massachusetts Institute of Technology(麻省理工学院)
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.CY
Comments to be published in IEEE TALE 2025
机构 * Carnegie Mellon University, USA(卡内基梅隆大学,美国) ; KTH Royal Institute of Technology, Sweden(皇家理工学院)
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
机构 * University of Exeter Department of Computer Science(埃克塞特大学计算机科学系)
专题命中 其他安全 :alignment(abstract);分类 cs.CL
Comments 16 pages, 11 figures (including appendix). To be presented at the 5th Wordplay @ EMNLP workshop (2025)
机构 * MIIT Key Laboratory of Data Intelligence and Management, Beihang University(信息科技部数据智能与管理重点实验室,北京航空航天大学)
专题命中 其他安全 :alignment(abstract);分类 cs.LG
Journal ref NeurIPS 2025
机构 * Greenwich Vietnam FPT University(越南格林威治FPT大学)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
机构 * Mila – Québec AI Institute(魁北克人工智能研究所) ; École polytechnique Palaiseau, France(巴黎理工大学Palaiseau分校) ; Université du Québec à Rimouski Lévis, Québec, Canada(魁北克 Rimouski 大学 Lévis 分校) ; Institut Pierre-Simon Laplace, IRD Sorbonne Université Paris, France(皮埃尔-西蒙·拉普拉斯研究所,IRD 巴黎索邦大学) ; Université de Montréal Montréal, Québec, Canada(蒙特利尔大学)
专题命中 其他安全 :safety(abstract);分类 cs.LG
Comments Tackling Climate Change with Machine Learning: workshop at NeurIPS 2025
机构 * UC Berkeley(加州大学伯克利分校) ; NEC Labs America(NEC美国实验室) ; UC San Diego(加州大学圣地亚哥分校)
专题命中 其他安全 :safety(abstract);分类 cs.LG
Comments ICCV 2025
专题命中 其他安全 :alignment(abstract);分类 cs.AI