arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-14 至 2025-11-14 共收录 45 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2507.20067 2025-11-14 cs.AI cs.CL cs.LG 82%

PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training

Sarat Chandra Bobbili, Ujwal Dinesha, Dheeraj Narasimha, Srinivas Shakkottai

机构 * Texas A&M University(德克萨斯A&M大学) Inria(法国国家信息与自动化研究所)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21597 2025-11-14 cs.CL cs.AI cs.LG 67%

Reducing the Scope of Language Models

David Yunis, Siyu Huo, Chulaka Gunasekara, Danish Contractor

机构 * Toyota Technological Institute at Chicago(丰田技术研究所(芝加哥)) IBM(IBM公司)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Appears in AAAI 2026 in the Main Technical Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10032 2025-11-14 cs.HC cs.AI cs.CY 66%

Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback

Vijay Keswani, Cyrus Cousins, Breanna Nguyen, Vincent Conitzer, Hoda Heidari, Jana Schaich Borg, Walter Sinnott-Armstrong

机构 * Duke University(杜克大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 偏好对齐 :alignment(abstract,comments);分类 cs.AI、cs.CY

Comments To appear in the AAAI 2026 Alignment Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10279 2025-11-14 cs.CV 50%

PROPA: Toward Process-level Optimization in Visual Reasoning via Reinforcement Learning

Yanbei Jiang, Chao Lei, Yihao Ding, Krista Ehinger, Jey Han Lau

机构 * University of Melbourne(墨尔本大学)

专题命中 偏好对齐 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 3 篇

2509.11816 2025-11-14 cs.LG cs.AI cs.CL 67%

Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning

Filip Sondej, Yushi Yang

机构 * Jagiellonian University(杰尔乔夫大学) University of Oxford(牛津大学)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10586 2025-11-14 eess.SY cs.RO cs.SY 50%

Safe Planning in Interactive Environments via Iterative Policy Updates and Adversarially Robust Conformal Prediction

Omid Mirzaeedodangeh, Eliot Shekhtman, Nikolai Matni, Lars Lindemann

机构 * Automatic Control Laboratory (IfA)(自动控制实验室)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09813 2025-11-14 cs.HC 50%

I've Seen Enough: Measuring the Toll of Content Moderation on Mental Health

Gabrielle M Gauthier, Eesha Ali, Amna Asim, Sarah Cornell-Maier, Lori A. Zoellner

专题命中 安全训练 :red teaming(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 5 篇

2511.09880 2025-11-14 cs.CL cs.CR 88%

EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models

Jialin Wu, Kecen Li, Zhicong Huang, Xinfeng Li, Xiaofeng Wang, Cheng Hong

机构 * Ant Group(蚂蚁集团) Nanyang Technological University(南洋理工大学)

专题命中 越狱攻击 :alignment(title,abstract);safety(title,abstract);分类 cs.CL

Comments Accepted by IEEE Symposium on Security and Privacy (S&P) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14250 2025-11-14 cs.CL cs.AI cs.CR 86%

Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors

Yi Zhao, Youzhi Zhang

机构 * Department of Computing The Hong Kong Polytechnic University Hong Kong SAR, China Centre for Artificial Intelligence Robotics Hong Kong Institute of Science \& Innovation Chinese Academy of Sciences Hong Kong SAR, China

专题命中 越狱攻击 :jailbreak(title,abstract);DPO(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments Accepted at ACSAC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18638 2025-11-14 cs.CR cs.AI cs.CL 84%

Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation

Daniel Schwartz, Dmitriy Bespalov, Zhe Wang, Ninad Kulkarni, Yanjun Qi

机构 * Amazon Bedrock Science(亚马逊Bedrock科学) Drexel University(德雷塞尔大学) University of Virginia(弗吉尼亚大学)

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract,comments);分类 cs.CL、cs.AI

Comments 14 pages, 5 figures; published in EMNLP 2025 ; Code at: https://github.com/dsbuddy/GAP-LLM-Safety

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10519 2025-11-14 cs.CL cs.AI 84%

Say It Differently: Linguistic Styles as Jailbreak Vectors

Srikant Panda, Avinash Rai

机构 * Independent Researcher(独立研究者)

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10088 2025-11-14 cs.LG cs.AI cs.CV 62%

eXIAA: eXplainable Injections for Adversarial Attack

Leonardo Pesce, Jiawen Wei, Gianmarco Mengaldo

机构 * Department of Mechanical Engineering National University of Singapore(机械工程系国立新加坡大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 1 篇

2507.19110 2025-11-14 cs.CV 50%

LISA: A Layer-wise Integration and Suppression Approach for Hallucination Mitigation in Multimodal Large Language Models

Zhihui Guo, Xin Man, Hui Xu, Jie Shao, Zhiguo Jiang, Xianchao Zhang, Heng Tao Shen

专题命中 幻觉与事实性 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 2 篇

2511.09895 2025-11-14 cs.LG cs.AI 62%

Simulator and Experience Enhanced Diffusion Model for Comprehensive ECG Generation

Xiaoda Wang, Kaiqiao Han, Yuhao Xu, Xiao Luo, Yizhou Sun, Wei Wang, Carl Yang

机构 * Emory University(埃默里大学) University of California, Los Angeles(加州大学洛杉矶分校) University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

专题命中 隐私与版权 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02573 2025-11-14 cs.CL 57%

Guess or Recall? Training CNNs to Classify and Localize Memorization in LLMs

Jérémie Dentan, Davide Buscaldi, Sonia Vanier

专题命中 隐私与版权 :alignment(abstract);分类 cs.CL

Comments This paper has been accepted for publication at AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 11 篇

2502.08045 2025-11-14 cs.CL cs.AI cs.CY 82%

Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs

Mohsinul Kabir, Ajwad Abrar, Sophia Ananiadou

机构 * Department of Computer Science, National Center for Text Mining, The University of Manchester(计算机科学系,文本挖掘国家中心,曼彻斯特大学) Department of Computer Science and Engineering, Islamic University of Technology(计算机科学与工程系,伊斯兰技术大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted at EMNLP 2025 (Main)

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09964 2025-11-14 cs.SE cs.AI cs.PL 79%

EnvTrace: Simulation-Based Semantic Evaluation of LLM Code via Execution Trace Alignment -- Demonstrated at Synchrotron Beamlines

Noah van der Vleuten, Anthony Flores, Shray Mathur, Max Rakitin, Thomas Hopkins, Kevin G. Yager, Esther H. R. Tsai

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07871 2025-11-14 cs.CL 79%

AlignSurvey: A Comprehensive Benchmark for Human Preferences Alignment in Social Surveys

Chenxi Lin, Weikang Yuan, Zhuoren Jiang, Biao Huang, Ruitao Zhang, Jianan Ge, Yueqian Xu, Jianxing Yu

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15434 2025-11-14 cs.CV cs.LG 79%

Semantic4Safety: Causal Insights from Zero-shot Street View Imagery Segmentation for Urban Road Safety

Huan Chen, Ting Han, Siyu Chen, Zhihao Guo, Yiping Chen, Meiliu Wu

机构 * School of Geospatial Engineering and Science, Sun Yat-sen University(地理空间工程与科学学院,中山大学) School of Geographical and Earth Sciences, University of Glasgow(地理与地球科学学院,格拉斯哥大学) School of Economics and Management, Shanxi University(经济学与管理学院,山西大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

Comments 11 pages, 10 figures, The 8th ACM SIGSPATIAL International Workshop on AI for Geographic Knowledge Discovery (GeoAI '25), November 3--6, 2025, Minneapolis, MN, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09855 2025-11-14 cs.LG 74%

Unlearning Imperative: Securing Trustworthy and Responsible LLMs through Engineered Forgetting

James Jin Kang, Dang Bui, Thanh Pham, Huo-Chong Ling

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments 14 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09748 2025-11-14 cs.CL cs.AI 62%

How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation

Muskaan Chopra, Lorenz Sparrenberg, Sarthak Khanna, Rafet Sifa

机构 * Fraunhofer IAIS - Department of Media Engineering(弗劳恩霍夫人工智能研究所-媒体工程部门) University of Bonn - Department of Computer Science(波恩大学-计算机科学系) Lamarr Institute for Machine Learning(拉马尔人工智能与机器学习研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted in IEEE BigData 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10203 2025-11-14 cs.CV cs.AI cs.RO 57%

VISTA: A Vision and Intent-Aware Social Attention Framework for Multi-Agent Trajectory Prediction

Stephane Da Silva Martins, Emanuel Aldea, Sylvie Le Hégarat-Mascle

机构 * SATIE - CNRS UMR 8029 Paris-Saclay University, France(巴黎-萨克雷大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Paper accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07127 2025-11-14 cs.LG 57%

REACT-LLM: A Benchmark for Evaluating LLM Integration with Causal Features in Clinical Prognostic Tasks

Linna Wang, Zhixuan You, Qihui Zhang, Jiunan Wen, Ji Shi, Yimin Chen, Yusen Wang, Fanqi Ding, Ziliang Feng, Li Lu

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03553 2025-11-14 cs.CL 57%

CCD-Bench: Probing Cultural Conflict in Large Language Model Decision-Making

Hasibur Rahman, Hanan Salam

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05796 2025-11-14 cs.CV cs.AI 57%

Dual-Mode Deep Anomaly Detection for Medical Manufacturing: Structural Similarity and Feature Distance

Julio Zanon Diaz, Georgios Siogkas, Peter Corcoran

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 12 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09742 2025-11-14 cs.CV cs.AI 57%

Feature Quality and Adaptability of Medical Foundation Models: A Comparative Evaluation for Radiographic Classification and Segmentation

Frank Li, Theo Dapamede, Mohammadreza Chavoshi, Young Seok Jeon, Bardia Khosravi, Abdulhameed Dere, Beatrice Brown-Mulry, Rohan Satya Isaac, Aawez Mansuri, Chiratidzo Sanyika, Janice Newsome, Saptarshi Purkayastha, Imon Banerjee, Hari Trivedi, Judy Gichoya

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 7 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 3 篇

2511.09663 2025-11-14 cs.CY cs.AI cs.HC 81%

Alignment Debt: The Hidden Work of Making AI Usable

Cumi Oyemike, Elizabeth Akpan, Pierre Hervé-Berdys

机构 * YUX Design(YUX设计)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments 19 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10573 2025-11-14 cs.LG cs.AI cs.CL cs.HC cs.MA 80%

Towards Emotionally Intelligent and Responsible Reinforcement Learning

Garapati Keerthana, Manik Gupta

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05901 2025-11-14 cs.CL cs.AI 73%

Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations

Rui Yang, Matthew Yu Heng Wong, Huitao Li, Xin Li, Wentao Zhu, Jingchi Liao, Kunyu Yu, Jonathan Chong Kai Liew, Weihao Xuan, Yingjian Chen, Yuhe Ke, Jasmine Chiat Ling Ong, Douglas Teodoro, Chuan Hong, Daniel Shi Wei Ting, Nan Liu

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 16 篇

2511.10215 2025-11-14 cs.CL cs.AI 81%

Persona-Aware Alignment Framework for Personalized Dialogue Generation

Guanrong Li, Xinyu Liu, Zhen Wu, Xinyu Dai

机构 * National Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室) Nanjing University(南京大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Pre-MIT Press publication version

详情

展开后加载摘要…

URL PDF HTML 收藏