Cascade Reward Sampling for Efficient Decoding-Time Alignment
机构 * Department of Computer Science(计算机科学系)
专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Department of Computer Science(计算机科学系)
专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.LG
机构 * Florida State University(佛罗里达州立大学)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments Accepted for publication in the Proceedings of the 5th Workshop on Bias and Fairness in AI (BIAS 2025) at ECML PKDD
机构 * School of Computer Engineering, Jimei University, Xiamen, 361021, China(厦门大学计算机工程学院) ; College of Science, Mathematics and Technology, Wenzhou-Kean University, Wenzhou, 325060, China(温州-凯恩大学科学、数学与技术学院) ; The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, 511453, China(香港科学与技术大学(广州)) ; School of Professional Studies, New York University, New York, 10003, United States(纽约大学专业研究学院) ; School of Informatics, Xiamen University, Xiamen, 361102, China(厦门大学信息学院)
专题命中 偏好对齐 :RLHF(abstract);prompt injection(abstract);分类 cs.AI
机构 * Dalian University of Technology(大连理工大学) ; University of Surrey(Surrey大学) ; University of Oxford(牛津大学)
专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG
专题命中 安全训练 :alignment(title,abstract);DPO(abstract);safety(abstract);分类 cs.AI
专题命中 安全训练 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.LG
机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Department of Electrical and Computer Engineering(电气与计算机工程系) ; Department of Mechanical Science Engineering(机械科学与工程系) ; School of Software(软件学院)
专题命中 安全训练 :safety(title,abstract)
Comments Accepted to Humanoids 2025
专题命中 安全训练 :alignment(abstract);trustworthy(abstract);分类 cs.AI
机构 * Brown University(布朗大学) ; Columbia University(哥伦比亚大学) ; Massachusetts Institute of Technology(麻省理工学院)
专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.CY
Comments ICML 2025
专题命中 越狱攻击 :prompt injection(title,abstract)
Comments ACL 2025 Main
专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.AI
Comments 15 pages
机构 * University of Science and Technology of China(中国科学技术大学) ; The Hong Kong Polytechnic University(香港理工大学) ; University of Washington(华盛顿大学) ; Nanjing University(南京大学) ; Stanford University(斯坦福大学) ; University of Chinese Academy of Sciences(中国科学院大学)
专题命中 越狱攻击 :safety(abstract);jailbreak(abstract)
Comments ICCV 2025
专题命中 越狱攻击 :safety(abstract)
Comments This paper has been submitted to the Transportation Research Board (TRB) for consideration for presentation at the 2026 Annual Meeting
专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL
机构 * The University of Queensland(昆士兰大学) ; University of California, Merced(加州大学默塞德分校)
专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI
Comments Work in progress
机构 * Department of Electrical and Computer Engineering, Aristotle University of Thessaloniki(阿尔伯塔大学电气与计算机工程系) ; Information Technologies Institute, CERTH(信息科技研究所)
专题命中 幻觉与事实性 :alignment(abstract)
Comments ICCV2025
专题命中 幻觉与事实性 :trustworthy(abstract)
Journal ref IEEE Communications Magazine, July 2025
机构 * School of Cyber Engineering, Xidian University(西安电子科技大学电子工程学院) ; School of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院)
专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI
机构 * Singapore Management University(新加坡管理大学) ; The University of Georgia(佐治亚大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments 19 pages
机构 * King’s College London(伦敦国王学院) ; Imperial College London(帝国理工学院) ; Columbia University(哥伦比亚大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
Comments Submitted to the NeurIPS 2025 Workshop GenAI4Health. Conference website: https://aihealth.ischool.utexas.edu/GenAI4HealthNeurips2025/
机构 * Okinawa Institute of Science and Technology(冲绳科学技术研究所) ; Institute of Science Tokyo(东京科学研究所) ; Amazon(亚马逊)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
机构 * AI VIETNAM Lab(AI越南实验室) ; Carnegie Mellon University(卡内基梅隆大学) ; University of Wisconsin - Madison(威斯康星大学麦迪逊分校) ; University of Pittsburgh(匹兹堡大学) ; University of Alabama at Birmingham(阿拉巴马大学伯明翰分校) ; Northwestern University(西北大学)
专题命中 安全评测 :alignment(abstract);分类 cs.LG
Comments 11 pages, 5 figures. Accepted to VisionDocs @ ICCV 2025
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
Comments version 3
机构 * Urban Information Lab, The University of Texas at Austin(德克萨斯大学奥斯汀分校城市信息实验室)
专题命中 安全评测 :safety(abstract);分类 cs.LG
专题命中 安全评测 :alignment(abstract)
Comments Video for Figure 4: https://youtu.be/5nSml5F2oQk?si=X0q51SXcZiTRBdFS Video for Figure 6 (2D): https://youtu.be/ANkAZ9FMq0o?si=px8oSAeKkUwBR7uf Video for Figure 6 (3D): https://youtu.be/AUlcCH7P73U?si=QFYYeCGLIGdsJCad
Journal ref Journal of Computational Science, Volume 87, 2025, 102574
机构 * University of Sydney(悉尼大学) ; University of Wollongong(沃林戈大学) ; University of Adelaide(阿德莱德大学) ; First Clinical Medical College, Guangzhou University of Chinese Medicine(广州中医药大学第一临床学院)
专题命中 安全评测 :alignment(abstract)
机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学)
专题命中 安全评测 :alignment(abstract)
机构 * The Chinese University of Hong Kong, Department of Computer Science & Engineering(香港中文大学计算机科学与工程系)
专题命中 安全评测 :alignment(abstract)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY
Comments This work has been submitted to the IEEE for possible publication
机构 * Université de Montréal(蒙特利尔大学) ; Mila – Quebec AI Institute(魁北克人工智能研究院) ; Chandar Research Lab(Chandar研究实验室) ; IBM Research(IBM研究院) ; Fujitsu Research(富士通研究院) ; Polytechnique Montréal(蒙特利尔理工学院)
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG