Natural Spectral Fusion: p-Exponent Cyclic Scheduling and Early Decision-Boundary Alignment in First-Order Optimization
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG
机构 * Beihang University(北航大学) ; National University of Singapore(国立新加坡大学) ; King Abdullah University of Science and Technology(国王 Abdullah 科学与技术大学)
专题命中 偏好对齐 :alignment(title,abstract)
Comments Accepted by ICCV 2025
机构 * Caltech(加州理工学院) ; UIUC(伊利诺伊大学) ; University of Cambridge(剑桥大学)
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.CY
Comments We make public all code and source data at https://github.com/psychology-of-AI/Personality-Illusion for full reproducibility
专题命中 偏好对齐 :DPO(abstract)
机构 * Leibniz Institute for Resilience Research(莱比锡韧性研究所) ; University Medical Center Halle(哈雷医学院) ; German Center for Mental Health (DZPG)(德国心理健康中心(DZPG)) ; University Medical Center of the Johannes Gutenberg-University Mainz(美因茨约瑟夫·冯·拉贝大学医学院)
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.LG
机构 * University of California Berkeley(加州大学伯克利分校) ; Virginia Tech(弗吉尼亚理工大学)
专题命中 安全训练 :safety(title,abstract);分类 cs.LG
专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG
Comments 9 pages
Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 2025
机构 * Dept. of Computing Science University of Aberdeen(计算科学系阿伯丁大学)
专题命中 安全训练 :safety(abstract);分类 cs.CL
Comments The paper has 5 figures and 1 table
专题命中 安全训练 :trustworthy(abstract)
专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.LG
Comments Published at ICLR 2025
机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
专题命中 红队测试 :safety(abstract);分类 cs.CL、cs.LG
机构 * Harbin Institute of Technology, China(哈尔滨工业大学) ; Zhejiang University, China(浙江大学) ; Tsinghua University, Shenzhen, China(清华大学深圳研究院)
专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI
Comments Accepted to EMNLP 2025 Main Conference (Oral)
机构 * Department of Machine Learning Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) Abu Dhabi, UAE(机器学习系,Mohamed bin Zayed人工智能大学(MBZUAI),阿布扎比,阿联酋) ; Department of Computer Vision Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) Abu Dhabi, UAE(计算机视觉系,Mohamed bin Zayed人工智能大学(MBZUAI),阿布扎比,阿联酋)
专题命中 幻觉与事实性 :alignment(title,abstract)
机构 * Department of School of Engineering, Shenzhen MSU-BIT University, Shenzhen, China, 518000(深圳MSU-BIT大学工程学院部门) ; MSU-BIT-SMBU Joint Research Center of Applied Mathematics, Shenzhen MSU-BIT University, Shenzhen, China, 518000(应用数学联合研究中心)
专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI
机构 * Queen's University(女王大学)
专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI
机构 * OpenAI ; Georgia Tech(佐治亚理工学院)
专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL
机构 * Georgia Institute of Technology(佐治亚理工学院) ; Dolby Laboratories(杜比实验室)
专题命中 安全评测 :alignment(title,abstract);safety(title,abstract);分类 cs.AI
Comments Rejected by AAAI25-AIA. Accepted by ICML25. Authors are thankful to the anonymous reviewers from both AAAI25-AIA and ICML25
机构 * Department of Data Science & AI, Monash University(数据科学与人工智能系,莫纳什大学) ; College of Engineering and Computer Science, VinUniversity(工程与计算机科学学院,文大学)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL
Comments EMNLP 2025
机构 * University of Hull(赫尔大学)
专题命中 安全评测 :safety(title,abstract);分类 cs.LG
专题命中 安全评测 :alignment(title);分类 cs.CL、cs.LG
机构 * Dept. of Information Engineering University of Pisa, Italy Pisa, Italy(信息工程系 乌迪内大学 乌迪内,意大利)
专题命中 安全评测 :alignment(abstract);safety(abstract);trustworthy(abstract)
机构 * Lightcap ; Department of Future(未来系) ; Turkish Aeronautical Association(土耳其航空航天协会)
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI
Comments 13 pages
机构 * Department of Computer Science and Artificial Intelligence(计算机科学与人工智能系) ; Andalusian Institute of Data Science and Computational Intelligence(安达卢西亚数据科学与计算智能研究所) ; University of Granada(格拉纳达大学)
专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.CL
Comments 27 pages, 9 tables and 2 figures
机构 * NYU(纽约大学) ; Independent Researcher(独立研究者)
专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI
专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.CY
Comments 17 pages, 1 figure, reproduced with permission from Springer Nature from Coordination, Organizations, Institutions, Norms, and Ethics for Governance of Multi-Agent Systems XVIII (COINE 2025)
机构 * Oracle AI
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to EMNLP 2025
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments Accepted at IEEE-EMBS International Conference on Biomedical and Health Informatics (BHI'25). 7 pages, 5 figures, 3 tables
机构 * Vrije University of Amsterdam(阿姆斯特丹自由大学) ; University College London(伦敦大学学院) ; University of Amsterdam(阿姆斯特丹大学) ; National Institute of Informatics(日本信息处理学会) ; Shandong University(山东大学) ; Guangzhou University(广州大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments 14 pages, accepted by EMNLP 2025
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG
机构 * The Chinese University of Hong Kong(香港中文大学) ; Noah’s Ark Lab, Huawei(华为诺亚实验室)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments Submitted to 1st Open Conference on AI Agents for Science (agents4science 2025)