arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-27 至 2025-10-27 共收录 42 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 5 篇

2505.17859 2025-10-27 cs.LG cs.AI stat.ML 86%

Scalable Valuation of Human Feedback through Provably Robust Model Alignment

Masahiro Fujisawa, Masaki Adachi, Michael A. Osborne

机构 * The University of Osaka(大阪大学) Lattice Lab, Toyota Motor Corporation(丰田公司Lattice实验室) Machine Learning Research Group, University of Oxford(牛津大学机器学习研究组) RIKEN AIP(理化学研究所AIP)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI、cs.LG

Comments Accepted by the 39th Conference on Neural Information Processing Systems (NeurIPS2025), 49 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21798 2025-10-27 cs.CL cs.AI 81%

Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment

Hongbin Zhang, Kehai Chen, Xuefeng Bai, Yang Xiang, Min Zhang

机构 * Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen, China(计算与智能研究所,哈尔滨工业大学,深圳,中国) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Under review;Work in progress;

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12854 2025-10-27 cs.CL cs.AI 73%

TPO: Aligning Large Language Models with Multi-branch & Multi-step Preference Trees

Weibin Liao, Xu Chu, Yasha Wang

机构 * School of Computer Science, Peking University(北京大学计算机科学系) Center on Frontiers of Computing Studies, Peking University(北京大学前沿计算研究中心) National Research and Engineering Center of Software Engineering, Peking University(北京大学软件工程研究中心)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11475 2025-10-27 cs.CL cs.AI cs.LG 67%

HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages

Zhilin Wang, Jiaqi Zeng, Olivier Delalleau, Hoo-Chang Shin, Felipe Soares, Alexander Bukharin, Ellie Evans, Yi Dong, Oleksii Kuchaiev

机构 * NVIDIA

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

Comments NeurIPS 2025 Datasets and Benchmarks Track Camera Ready, 46 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11080 2025-10-27 cs.CL cs.AI cs.LG 67%

BLEUBERI: BLEU is a surprisingly effective reward for instruction following

Yapei Chang, Yekyung Kim, Michael Krumdick, Amir Zadeh, Chuan Li, Chris Tanner, Mohit Iyyer

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments neurips cam-ready

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 5 篇

2406.16258 2025-10-27 cs.RO cs.AI cs.LG 81%

MEReQ: Max-Ent Residual-Q Inverse RL for Sample-Efficient Alignment from Intervention

Yuxin Chen, Chen Tang, Jianglan Wei, Chenran Li, Ran Tian, Xiang Zhang, Wei Zhan, Peter Stone, Masayoshi Tomizuka

专题命中 安全训练 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21254 2025-10-27 cs.AI 79%

Out-of-Distribution Detection for Safety Assurance of AI and Autonomous Systems

Victoria J. Hodge, Colin Paterson, Ibrahim Habli

机构 * Department of Computer Science University of York(计算机科学系约克大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07736 2025-10-27 cs.AI 77%

RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards

Jingnan Zheng, Xiangtian Ji, Yijun Lu, Chenhang Cui, Weixiang Zhao, Gelei Deng, Zhenkai Liang, An Zhang, Tat-Seng Chua

专题命中 安全训练 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.AI

Journal ref 39th Conference on Neural Information Processing Systems (NeurIPS 2025). 39th Conference on Neural Information Processing Systems (NeurIPS 2025). 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10226 2025-10-27 cs.CV cs.AI cs.LG 62%

ScoreMix: Synthetic Data Generation by Score Composition in Diffusion Models Improves Recognition

Parsa Rahimi, Sebastien Marcel

机构 * EPFL(苏黎世联邦理工学院) Idiap(日内瓦智能感知研究院)

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

Comments Extended version of ICMLw25 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12733 2025-10-27 cs.RO cs.AI cs.LG 62%

HYPE: Hybrid Planning with Ego Proposal-Conditioned Predictions

Hang Yu, Julian Jordan, Julian Schmidt, Silvan Lindner, Alessandro Canevaro, Wilhelm Stork

机构 * Mercedes-Benz AG, Research & Development(梅赛德斯-奔驰集团,研发部) Karlsruhe Institute of Technology, ITIV(卡尔斯鲁厄理工学院,ITIV) University of Tübingen(图宾根大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to IEEE ITSC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 3 篇

2505.13763 2025-10-27 cs.AI cs.CL q-bio.NC 73%

Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal Activations

Li Ji-An, Hua-Dong Xiong, Robert C. Wilson, Marcelo G. Mattar, Marcus K. Benna

机构 * Neurosciences Graduate Program University of California San Diego(加州大学圣地亚哥分校神经科学研究生项目) School of Psychology Georgia Tech(佐治亚理工学院心理学系) Department of Psychology New York University(纽约大学心理学系) Department of Neurobiology University of California San Diego(加州大学圣地亚哥分校神经生物学系)

专题命中 越狱攻击 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21214 2025-10-27 cs.CR 50%

Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses

Xingwei Zhong, Kar Wai Fok, Vrizlynn L. L. Thing

专题命中 越狱攻击 :jailbreak(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21189 2025-10-27 cs.CR 50%

Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency

Yukun Jiang, Mingjie Li, Michael Backes, Yang Zhang

专题命中 越狱攻击 :jailbreak(abstract)

Comments Accepted in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 3 篇

2510.21049 2025-10-27 cs.CL cs.AI cs.LG 82%

Reasoning's Razor: Reasoning Improves Accuracy but Can Hurt Recall at Critical Operating Points in Safety and Hallucination Detection

Atoosa Chegini, Hamid Kazemi, Garrett Souza, Maria Safi, Yang Song, Samy Bengio, Sinead Williamson, Mehrdad Farajtabar

机构 * Apple(苹果公司)

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21323 2025-10-27 cs.CV cs.LG 79%

VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set

Shufan Shen, Junshu Sun, Qingming Huang, Shuhui Wang

机构 * Key Lab of Intell. Info. Process., Inst. of Comput. Tech., CAS(智能信息处理重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.LG

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06771 2025-10-27 cs.AI cs.CV cs.LG 62%

Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty

Meera Hahn, Wenjun Zeng, Nithish Kannen, Rich Galt, Kartikeya Badola, Been Kim, Zi Wang

机构 * Google DeepMind(谷歌DeepMind)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

Journal ref International Conference on Machine Learning, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 安全评测 12 篇

2510.21557 2025-10-27 cs.AI 79%

Co-Sight: Enhancing LLM-Based Agents via Conflict-Aware Meta-Verification and Trustworthy Reasoning with Structured Facts

Hongwei Zhang, Ji Lu, Shiqing Jiang, Chenxiang Zhu, Li Xie, Chen Zhong, Haoran Chen, Yurui Zhu, Yongsheng Du, Yanqin Gao, Lingjun Huang, Baoli Wang, Fang Tan, Peng Zou

机构 * Zhongxing Telecom Equipment (ZTE), China(中兴通讯设备(ZTE),中国)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21551 2025-10-27 cs.LG 79%

Interpretable Multimodal Zero-Shot ECG Diagnosis via Structured Clinical Knowledge Alignment

Jialu Tang, Hung Manh Pham, Ignace De Lathauwer, Henk S. Schipper, Yuan Lu, Dong Ma, Aaqib Saeed

机构 * Eindhoven University of Technology(埃因霍温理工大学) Singapore Management University(新加坡管理大学) Maxima Medical Center(马克斯医疗中心) Erasmus Medical Center(埃因霍温医学院)

专题命中 安全评测 :alignment(title);trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21606 2025-10-27 cs.CV 78%

Modest-Align: Data-Efficient Alignment for Vision-Language Models

Jiaxiang Liu, Yuan Wang, Jiawei Du, Joey Tianyi Zhou, Mingkun Xu, Zuozhu Liu

机构 * Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) ZJU-Angelalign R&D Center for Intelligence Healthcare(浙大天使align智能医疗研发中心) Centre for Frontier AI Research (CFAR)(前沿人工智能研究中心) Agency for Science, Technology and Research (A*STAR)(科技研究局) Institute of High Performance Computing (IHPC)(高性能计算研究所)

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21120 2025-10-27 cs.CV 78%

SafetyPairs: Isolating Safety Critical Image Features with Counterfactual Image Generation

Alec Helbling, Shruti Palaskar, Kundan Krishna, Polo Chau, Leon Gatys, Joseph Yitan Cheng

机构 * Georgia Tech(佐治亚理工学院) Apple(苹果公司)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21524 2025-10-27 cs.AI 70%

EU-Agent-Bench: Measuring Illegal Behavior of LLM Agents Under EU Law

Ilija Lichkovski, Alexander Müller, Mariam Ibrahim, Tiwai Mhundwa

机构 * AI Safety Initiative Groningen(格罗宁根人工智能安全倡议)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI

Comments Accepted at the Workshop on Regulatable ML at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21133 2025-10-27 cs.CR cs.AI 70%

Quantifying CBRN Risk in Frontier Models

Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal, Prashanth Harshangi

机构 * Enkrypt AI

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05201 2025-10-27 cs.LG cs.AI cs.CL 67%

FAITH: A Framework for Assessing Intrinsic Tabular Hallucinations in Finance

Mengao Zhang, Jiayu Fu, Tanya Warrier, Yuwen Wang, Tianhui Tan, Ke-wei Huang

机构 * Asian Institute of Digital Finance, National University of Singapore(亚洲数字金融研究所,新加坡国立大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 9 pages, AMC ICAIF'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21090 2025-10-27 cs.CL cs.AI cs.LG 67%

Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only

Qingru Zhang, Liang Qiu, Ilgee Hong, Zhenghao Xu, Tianyi Liu, Shiyang Li, Rongzhi Zhang, Zheng Li, Lihong Li, Bing Yin, Chao Zhang, Jianshu Chen, Haoming Jiang, Tuo Zhao

机构 * Georgia Institute of Technology(佐治亚理工学院) Amazon(亚马逊)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21389 2025-10-27 cs.LG cs.AI cs.HC 62%

Assessing the Real-World Utility of Explainable AI for Arousal Diagnostics: An Application-Grounded User Study

Stefan Kraft, Andreas Theissler, Vera Wienhausen-Wilke, Gjergji Kasneci, Hendrik Lensch

机构 * IT-Designers Gruppe Esslingen am Neckar(埃斯林根内卡尔IT设计组) University of Tübingen(图宾根大学) University of Giessen(吉森大学) Klinikum Esslingen, Klinik für Kardiologie, Pneumologie und Angilologie(埃斯林根克林克医院,心内科、呼吸科和内科) Technical University of Munich(慕尼黑技术大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03317 2025-10-27 cs.CV cs.AI 57%

Photorealistic Inpainting for Perturbation-based Explanations in Ecological Monitoring

Günel Aghakishiyeva, Jiayi Zhou, Saagar Arya, Julian Dale, James David Poling, Holly R. Houliston, Jamie N. Womble, Gregory D. Larsen, David W. Johnston, Brinnae Bent

机构 * Duke University(杜克大学) University of Agder(阿格德大学) University of Cambridge(剑桥大学) U.S. National Park Service Department of Interior(美国国家公园管理局) Alaska Spatial Science(阿拉斯加空间科学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments NeurIPS 2025 Imageomics Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04675 2025-10-27 cs.AI cs.LO 57%

HypRL: Reinforcement Learning of Control Policies for Hyperproperties

Tzu-Han Hsu, Arshia Rafieioskouei, Borzoo Bonakdarpour

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20976 2025-10-27 cs.LG 57%

L^2M^3OF: A Large Language Multimodal Model for Metal-Organic Frameworks

Jiyu Cui, Fang Wu, Haokai Zhao, Minggao Feng, Xenophon Evangelopoulos, Andrew I. Cooper, Yejin Choi

机构 * Department of Chemistry, University of Liverpool(利兹大学化学系) Leverhulme Research Centre for Functional Materials Design, University of Liverpool(利兹大学功能性材料设计研究所以) Department of Computer Science, University of Stanford(斯坦福大学计算机科学系) School of Computer Science and Engineering, University of New South Wales(新南威尔士大学计算机科学与工程学院)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

6. AI治理与伦理 4 篇

2506.21584 2025-10-27 cs.CL cs.AI cs.CY 82%

Empirical Evidence for Alignment Faking in a Small LLM and Prompt-Based Mitigation Techniques

Jeanice Koorndijk

机构 * Seraphion Technology(塞拉菲昂技术)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments NeurIPS RegML Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16836 2025-10-27 cs.CL cs.AI 62%

Misspellings in Natural Language Processing: A survey

Gianluca Sperduti, Alejandro Moreo

机构 * Istituto di Scienza e Tecnologie dell’Informazione, Consiglio Nazionale delle Ricerche(信息科学与技术研究所,国家研究理事会)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏