arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-16 至 2025-10-16 共收录 51 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 7 篇

2510.13512 2025-10-16 cs.LG cs.AI 84%

Offline and Online KL-Regularized RLHF under Differential Privacy

Yulian Wu, Rushil Thareja, Praneeth Vepakomma, Francesco Orabona

机构 * King Abdullah University of Science and Technology (KAUST)(卡布尔大学科学与技术学院) Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学) Massachusetts Institute of Technology (MIT)(麻省理工学院)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13694 2025-10-16 cs.LG 83%

Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking

Yuchun Miao, Liang Ding, Sen Zhang, Rong Bao, Lefei Zhang, Dacheng Tao

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments 46 pages, 36 figures, submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03459 2025-10-16 cs.LG 74%

Can DPO Learn Diverse Human Values? A Theoretical Scaling Law

Shawn Im, Sharon Li

机构 * Department of Computer Sciences University of Wisconsin-Madison(计算机科学系威斯康星大学麦迪逊分校)

专题命中 偏好对齐 :DPO(title);分类 cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12041 2025-10-16 cs.CL 70%

Improving Text-to-Image Generation with Input-Side Inference-Time Scaling

Ruibo Chen, Jiacheng Pan, Heng Huang, Zhenheng Yang

机构 * TikTok University of Maryland, College Park(马里兰大学)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13501 2025-10-16 cs.AI 57%

Confidence as a Reward: Transforming LLMs into Reward Models

He Du, Bowen Li, Chengxing Xie, Chang Gao, Kai Chen, Dacheng Tao

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) Xidian University(西安电子科技大学) The Chinese University of Hong Kong(香港中文大学) Nanyang Technological University(南洋理工大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23144 2025-10-16 cs.AI cond-mat.stat-mech cs.MA nlin.AO physics.soc-ph 57%

Coordination Requires Simplification: Thermodynamic Bounds on Multi-Objective Compromise in Natural and Artificial Intelligence

Atma Anand

机构 * Department of Physics and Astronomy, University of Rochester(物理与天文学系,罗切斯特大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

Comments 15 pages, 1 figure, 9 pages supplementary material, submitted to Journal of Physics: Complexity

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10013 2025-10-16 cs.CV cs.CL 57%

Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki Effect

Tom Kouwenhoven, Kiana Shahrasbi, Tessa Verhoef

机构 * Leiden Institute of Advanced Computer Science(莱顿先进计算机科学研究所) Leiden University(莱顿大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments Presented at the Thirty-Ninth Annual Conference on Neural Information Processing Systems (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 1 篇

2510.13093 2025-10-16 stat.ML cs.AI cs.LG 62%

A Multi-dimensional Semantic Surprise Framework Based on Low-Entropy Semantic Manifolds for Fine-Grained Out-of-Distribution Detection

Ningkang Peng, Yuzhe Mao, Yuhao Zhang, Linjin Qian, Qianfeng Yu, Yanhui Gu, Yi Chen, Li Kong

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 2 篇

2510.13190 2025-10-16 cs.CL 70%

SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs

Juan Ren, Mark Dras, Usman Naseem

机构 * School of Computing, Macquarie University(计算机学院,麦考瑞大学)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13451 2025-10-16 cs.CR 50%

Toward Efficient Inference Attacks: Shadow Model Sharing via Mixture-of-Experts

Li Bai, Qingqing Ye, Xinwei Zhang, Sen Zhang, Zi Liang, Jianliang Xu, Haibo Hu

专题命中 越狱攻击 :alignment(abstract)

Comments To appear in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 提示注入 1 篇

2510.13543 2025-10-16 cs.CR cs.AI 79%

In-Browser LLM-Guided Fuzzing for Real-Time Prompt Injection Testing in Agentic AI Browsers

Avihay Cohen

机构 * Avihay Cohen

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments 37 pages , 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 3 篇

2505.19234 2025-10-16 cs.AI cs.CL cs.MA 62%

GUARDIAN: Safeguarding LLM Multi-Agent Collaborations with Temporal Graph Modeling

Jialong Zhou, Lichao Wang, Xiao Yang

机构 * King’s College London(伦敦国王学院) Beijing Institute of Technology(北京理工大学) Tsinghua University(清华大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12470 2025-10-16 cs.CL 57%

Reasoning on a Spectrum: Aligning LLMs to System 1 and System 2 Thinking

Alireza S. Ziabari, Nona Ghazizadeh, Zhivar Sourati, Farzan Karimi-Malekabadi, Payam Piray, Morteza Dehghani

机构 * University of Southern California(南加州大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13183 2025-10-16 cs.CL 57%

DSCD: Large Language Model Detoxification with Self-Constrained Decoding

Ming Dong, Jinkui Zhang, Bolong Zheng, Xinhui Tu, Po Hu, Tingting He

机构 * Hubei Provincial Key Laboratory of Artificial Intelligence and Smart Learning(湖北人工智能与智能学习省级重点实验室) National Language Resources Monitoring and Research Center for Network Media(网络媒体语言资源监测与研究国家中心) Central China Normal University(中央财经大学) Wuhan University of Technology(武汉理工大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

Comments Accepted at EMNLP 2025 MainConference

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 隐私与版权 2 篇

2510.13606 2025-10-16 cs.LG 57%

Towards Robust Knowledge Removal in Federated Learning with High Data Heterogeneity

Riccardo Santi, Riccardo Salami, Simone Calderara

机构 * AImageLab, University of Modena and Reggio Emilia(AImageLab,摩德纳和雷吉奥艾米利亚大学)

专题命中 隐私与版权 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13136 2025-10-16 cs.CR 50%

Privacy-Aware Framework of Robust Malware Detection in Indoor Robots: Hybrid Quantum Computing and Deep Neural Networks

Tan Le, Van Le, Sachin Shetty

专题命中 隐私与版权 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 安全评测 14 篇

2510.13653 2025-10-16 cs.CY 88%

International AI Safety Report 2025: First Key Update: Capabilities and Risk Implications

Yoshua Bengio, Stephen Clare, Carina Prunkl, Shalaleh Rismani, Maksym Andriushchenko, Ben Bucknall, Philip Fox, Tiancheng Hu, Cameron Jones, Sam Manning, Nestor Maslej, Vasilios Mavroudis, Conor McGlynn, Malcolm Murray, Charlotte Stix, Lucia Velasco, Nicole Wheeler, Daniel Privitera, Sören Mindermann, Daron Acemoglu, Thomas G. Dietterich, Fredrik Heintz, Geoffrey Hinton, Nick Jennings, Susan Leavy, Teresa Ludermir, Vidushi Marda, Helen Margetts, John McDermid, Jane Munga, Arvind Narayanan, Alondra Nelson, Clara Neppel, Gopal Ramchurn, Stuart Russell, Marietje Schaake, Bernhard Schölkopf, Alavaro Soto, Lee Tiedrich, Gaël Varoquaux, Andrew Yao, Ya-Qin Zhang, Leandro Aguirre, Olubunmi Ajala, Fahad Albalawi Noora AlMalek, Christian Busch, André Carvalho, Jonathan Collas, Amandeep Gill, Ahmet Hatip, Juha Heikkilä, Chris Johnson, Gill Jolly, Ziv Katzir, Mary Kerema, Hiroaki Kitano, Antonio Krüger, Aoife McLysaght, Oleksii Molchanovskyi, Andrea Monti, Kyoung Mu Lee, Mona Nemer, Nuria Oliver, Raquel Pezoa, Audrey Plonk, José Portillo, Balaraman Ravindran, Hammam Riza, Crystal Rugege, Haroon Sheikh, Denise Wong, Yi Zeng, Liming Zhu

专题命中 安全评测 :safety(title,abstract);AI safety(title,abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13351 2025-10-16 cs.CL cs.AI 86%

Protect: Towards Robust Guardrailing Stack for Trustworthy Enterprise LLM Systems

Karthik Avinash, Nikhil Pareek, Rishav Hada

机构 * FutureAGI Inc.(未来人工智能公司)

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);prompt injection(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09710 2025-10-16 cs.CL cs.AI 81%

SeCon-RAG: A Two-Stage Semantic Filtering and Conflict-Free Framework for Trustworthy RAG

Xiaonan Si, Meilin Zhu, Simeng Qin, Lijia Yu, Lijun Zhang, Shuaitong Liu, Xinfeng Li, Ranjie Duan, Yang Liu, Xiaojun Jia

机构 * Institute of Software Chinese Academy of Sciences Beijing China(中国科学院软件研究所) Key Laboratory of System Software (Chinese Academy of Sciences) and State Key Laboratory of Computer Science, Institute of Software, Chinese Academy of Sciences, Beijing, China(中国科学院系统软件重点实验室和计算机科学国家重点实验室) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学) Northeast University China(东北大学) Institute of Ai For industries Nanjing China(人工智能产业研究院) Southwest University China(西南大学) Nanyang Technological University Singapore(新加坡南洋理工大学) Alibaba China(阿里巴巴(中国))

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17244 2025-10-16 cs.CL cs.AI 81%

ReasoningShield: Safety Detection over Reasoning Traces of Large Reasoning Models

Changyi Li, Jiayi Wang, Xudong Pan, Geng Hong, Min Yang

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12864 2025-10-16 cs.AI cs.CL cs.LG 75%

From Literal to Liberal: A Meta-Prompting Framework for Eliciting Human-Aligned Exception Handling in Large Language Models

Imran Khan

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 13 pages. Code and data are available at https://github.com/strongSoda/LITERAL-TO-LIBERAL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13804 2025-10-16 cs.CV cs.AI cs.CL 62%

Generative Universal Verifier as Multimodal Meta-Reasoner

Xinchen Zhang, Xiaoying Zhang, Youbin Wu, Yanbin Cao, Renrui Zhang, Ruihang Chu, Ling Yang, Yujiu Yang

机构 * Tsinghua University(清华大学) ByteDance Seed(字节跳动种子) Princeton University(普林斯顿大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13106 2025-10-16 cs.SE cs.AI cs.CL 62%

TRUSTVIS: A Multi-Dimensional Trustworthiness Evaluation Framework for Large Language Models

Ruoyu Sun, Da Song, Jiayang Song, Yuheng Huang, Lei Ma

机构 * University of Alberta(阿尔伯塔大学) Mila - Quebec Artificial Intelligence Institute(魁北克人工智能研究所) The University of Tokyo(东京大学) Macau University of Science and Technology(澳门科学技术大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments 4 pages, 2 figures, To appear in ASE 2025 Demo Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07575 2025-10-16 cs.AI cs.LG 62%

Benchmarking is Broken -- Don't Let AI be its Own Judge

Zerui Cheng, Stella Wohnig, Ruchika Gupta, Samiul Alam, Tassallah Abdullahi, João Alves Ribeiro, Christian Nielsen-Garcia, Saif Mir, Siran Li, Jason Orender, Seyed Ali Bahrainian, Daniel Kirste, Aaron Gokaslan, Mikołaj Glinka, Carsten Eickhoff, Ruben Wolff

机构 * Princeton University(普林斯顿大学) CISPA Helmholtz Center for Information Security(CISPA海德堡信息安全中心) Michigan State University(密歇根州立大学) Ohio State University(俄亥俄州立大学) Brown University(布朗大学) Massachusetts Institute of Technology(麻省理工学院) University of California, Los Angeles(加州大学洛杉矶分校) University of Tübingen(图宾根大学) Old Dominion University(旧 Dominion 大学) Technical University of Munich(慕尼黑技术大学) Cornell University(康奈尔大学) Forest AI(森林AI)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 14 pages; Accepted to NeurIPS 2025. Link to poster: https://neurips.cc/virtual/2025/poster/121919; Link to project website: https://www.peerbench.ai/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02594 2025-10-16 cs.HC cs.AI cs.CL 62%

A Risk Taxonomy and Reflection Tool for Large Language Model Adoption in Public Health

Jiawei Zhou, Amy Z. Chen, Darshi Shah, Laura M. Schwab Reese, Munmun De Choudhury

机构 * Georgia Institute of Technology(佐治亚理工学院) Purdue University(普渡大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12981 2025-10-16 cs.LG 57%

Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check

Sungjun Cho, Dasol Hwang, Frederic Sala, Sangheum Hwang, Kyunghyun Cho, Sungmin Cha

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) LG AI Research(LG人工智能研究) Seoul National University of Science and Technology(首尔科学技术大学) New York University(纽约大学) Genentech(基因泰克)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 20 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12931 2025-10-16 cs.CV cs.CL 57%

Unifying Vision-Language Latents for Zero-label Image Caption Enhancement

Sanghyun Byun, Jung Ick Guack, Mohanad Odema, Baisub Lee, Jacob Song, Woo Seong Chung

机构 * LG Electronics USA(LG电子美国公司)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted to PMLR and NeurIPS 2025 UniReps

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09558 2025-10-16 cs.CL 57%

AutoPR: Let's Automate Your Academic Promotion!

Qiguang Chen, Zheng Yan, Mingda Yang, Libo Qin, Yixin Yuan, Hanjing Li, Jinhao Liu, Yiyan Ji, Dengyun Peng, Jiannan Guan, Mengkang Hu, Yantao Du, Wanxiang Che

机构 * LARG Research Center for Social Computing and Interactive Robotics(社会计算与交互机器人研究室) Harbin Institute of Technology(哈尔滨工业大学) School of Computer Science and Engineering(计算机科学与工程学院) Central South University(中南大学) The University of Hong Kong(香港大学) ByteDance China (Seed)(字节跳动中国(种子))

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Preprint. Code: https://github.com/LightChen233/AutoPR . Benchmark: https://huggingface.co/datasets/yzweak/PRBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12995 2025-10-16 cs.CV 50%

Brought a Gun to a Knife Fight: Modern VFM Baselines Outgun Specialized Detectors on In-the-Wild AI Image Detection

Yue Zhou, Xinan He, Kaiqing Lin, Bing Fan, Feng Ding, Jinhua Zeng, Bin Li

机构 * Guangdong Provincial Key Laboratory of Intelligent Information Processing(广东省智能信息处理重点实验室) Shenzhen Key Laboratory of Media Security(深圳媒体安全重点实验室) SZU AFS Joint Innovation Center for AI Technology, Shenzhen University(深圳大学AFS人工智能技术联合创新中心) University of North Texas(北卡罗来纳州立大学) Academy of Forensic Science(法医科学研究院) Nanchang University(南昌大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16902 2025-10-16 cs.CV cs.RO 50%

RealEngine: Simulating Autonomous Driving in Realistic Context

Junzhe Jiang, Nan Song, Jingyu Li, Xiatian Zhu, Li Zhang

机构 * School of Data Science, Fudan University(复旦大学数据科学学院) University of Surrey(萨里大学)

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏