Bridging Privacy and Robustness for Trustworthy Machine Learning
机构 * Huazhong University of Science and Technology(华中科技大学)
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Huazhong University of Science and Technology(华中科技大学)
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG
机构 * Department of Electrical Engineering, University of Southern California(电气工程系,美国南加州大学) ; Department of Aeronautics and Astronautics, Stanford University(航空与宇航系,斯坦福大学) ; NVIDIA Research(NVIDIA研究)
专题命中 安全评测 :safety(title,abstract)
机构 * Oxford Internet Institute(牛津互联网研究所) ; University of Oxford(牛津大学)
专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.CY
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments accepted at ICCVW'25 - Systematic Trust in AI Models: Ensuring Fairness, Reliability, Explainability, and Accountability in Machine Learning Frameworks
机构 * Department of Computer Science & Engineering(计算机科学与工程系) ; The Pennsylvania State University(宾夕法尼亚州立大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL
Comments accepted in COLM2025, 9 pages
机构 * Verily Life Sciences(Verily 生物科技)
专题命中 安全评测 :trustworthy(abstract);分类 cs.LG
Comments 27 pages
机构 * Florida Institute of Technology(佛罗里达理工学院) ; Microsoft Research(微软研究院)
专题命中 安全评测 :alignment(abstract);分类 cs.LG
Comments Accepted to ICCV 2025, MARS2 Workshop. Total 14 pages, 12 figures and 3 tables
机构 * Khalifa University(卡利法大学) ; The Chinese University of Hong Kong(香港中文大学) ; Queen Mary University of London(伦敦大学玛丽女王学院) ; Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) ; City University of Hong Kong(香港城市大学)
专题命中 安全评测 :alignment(abstract)
Comments Accepted in ICCV 2025, Codebase: https://github.com/fesvhtr/TRIG