Reinforcement Learning Foundations for Deep Research Systems: A Survey
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI
Comments 39 pages, second version
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI
Comments 39 pages, second version
机构 * Harvard University(哈佛大学)
专题命中 偏好对齐 :alignment(abstract);safety(abstract);分类 cs.LG
Comments Accepted to NeurIPS 2025 (Spotlight Paper)
机构 * Cisco AI Threat Research & Security(思科AI威胁研究与安全)
专题命中 安全训练 :alignment(abstract);safety(abstract);jailbreak(abstract);prompt injection(abstract)
机构 * Tokyo Women’s Medical University(东京女子医科大学)
专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.CY
Comments 14 pages, 2 tables, 1 figure
专题命中 安全训练 :safety(abstract);分类 cs.AI
Comments Under review
专题命中 安全训练 :safety(abstract);分类 cs.AI
专题命中 越狱攻击 :jailbreak(title);red teaming(abstract);分类 cs.CL
专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI
Comments Submitted to AAAI 2026 Workshop on Trust and Control in Agentic AI (TrustAgent)
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; University of Oxford(牛津大学) ; Beijing Academy of Artificial Intelligence(北京人工智能研究院) ; The Hong Kong University of Science and Technology(香港科学与技术大学) ; Beijing University of Technology(北京工业大学) ; Tsinghua University(清华大学) ; City University of Hong Kong(香港城市大学)
专题命中 提示注入 :prompt injection(abstract);分类 cs.CL
Comments This paper is accepted by IJCAI2025 Workshop on Deepfake Detection, Localization, and Interpretability as Best Student Paper
机构 * UCLA(美国大学洛杉矶分校) ; CUHK(香港中文大学) ; HKUST (GZ)(香港科技大学(广州)) ; HKUST(香港科技大学) ; Tsinghua University(清华大学) ; ECNU(华东师范大学) ; NVIDIA(英伟达) ; UCL(伦敦大学学院) ; HKU(香港大学) ; University of Cambridge(剑桥大学)
专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY
Comments EMNLP 2025 (Findings)
机构 * University of Amsterdam(阿姆斯特丹大学)
专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL
Comments Presented at UncertaiNLP Workshop at EMNLP 2025 https://aclanthology.org/2025.uncertainlp-main.21.pdf
Journal ref UncertaiNLP Workshop at Empirical Methods in Natural Language Processing 2025 (EMNLP 2025)
机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校) ; University of Texas at Dallas(德克萨斯大学达拉斯分校)
专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL
Comments Accepted by UDM-KDD'24
机构 * Christian Doppler Laboratory for Dependable Intelligent Systems in Harsh Environments(可信智能系统在恶劣环境中的克里斯蒂安·多普勒实验室) ; Graz University of Technology(格拉茨技术大学) ; Know Center Research GmbH(Know Center研究有限责任公司)
专题命中 安全评测 :trustworthy(title);分类 cs.LG
Comments Published in Machine Learning (Springer), vol. 114, no. 12, Article 267, 2025
Journal ref Mach Learn 114, 267 (2025)
机构 * Department of AI Ethics, CEA Paris-Saclay(人工智能伦理系,CEA巴黎萨克雷)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY
Comments 17 pages, 2 Tables and 2 Pictures
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 18 pages, 2 figures. Category: cs.LG. Code and data: https://github.com/Course-Correct-Labs/mirror-loop
机构 * Walmart Global Tech(沃尔玛全球技术)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
Comments 4 page, NeurIPS 2025 Workshop: Evaluating the Evolving LLM Lifecycle
机构 * School of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院) ; Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
专题命中 安全评测 :trustworthy(abstract)
机构 * University of British Columbia(不列颠哥伦比亚大学) ; Vector Institute for AI(人工智能向量研究所)
专题命中 安全评测 :alignment(abstract)
机构 * Department of Computer Science University of Bath, United Kingdom(计算机科学系 英国巴斯大学)
专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.LG
Comments 51 pages, Masters thesis. 10 tables, 7 figures, project data & code here: https://github.com/robynwyrick/mirror-neuron-frog-and-toad
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY
Comments Updated to the peer-reviewed version accepted and published in Computational Intelligence, Volume 41, Issue 6 (Wiley, 2025)
Journal ref Computational Intelligence, Volume 41, Issue 6 (Wiley, 2025)
机构 * University of Amsterdam(阿姆斯特丹大学) ; QUvA-Lab(QUvA实验室) ; Qualcomm AI Research(高通人工智能研究)
专题命中 其他安全 :alignment(title,abstract);分类 cs.LG
Comments To appear in NeurIPS workshop: AI and ML for Next-Generation Wireless Communications (AI4NextG)
机构 * RIKEN AIP(日本理化学研究所Advanced Institute for Physical Research) ; The University of Tokyo(东京大学)
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL
Comments EMNLP 2025 Findings
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
机构 * Department of Computer Science and Engineering(计算机科学与工程系) ; AIGEN Sciences(AIGEN公司)
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
Comments Accepted to IEEE BIBM 2025
机构 * Apple(苹果公司) ; Johns Hopkins University(约翰霍普金斯大学) ; Stanford University(斯坦福大学)
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
Comments Accepted to 2025 International Conference on Space Robotics (iSpaRo). Presented at RSS 2025 Workshop on Space Robotics
机构 * University of Chinese Academy of Sciences(中国科学院大学) ; Huawei Technologies Co., Ltd.(华为技术有限公司)
专题命中 其他安全 :alignment(abstract);分类 cs.CL
Comments EMNLP2025 Industry Track
机构 * Institute of Computer Science, University of Tartu, Estonia(计算机科学研究所,塔尔图大学,爱沙尼亚)
专题命中 其他安全 :safety(abstract);分类 cs.LG
Comments ECAI 2025
Journal ref Frontiers in Artificial Intelligence and Applications 413 (ECAI 2025) 5027 - 5034