arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-21 至 2025-10-21 共收录 98 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 11 篇

2510.16167 2025-10-21 cs.LG cs.CL 86%

Alignment is Localized: A Causal Probe into Preference Layers

Archie Chaudhury

机构 * Independent(独立研究者)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01203 2025-10-21 cs.LG stat.ML 83%

KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity

Gholamali Aminian, Amir R. Asadi, Idan Shenfeld, Youssef Mroueh

机构 * The Alan Turing Institute(艾伦·图灵研究所) Statistical Laboratory(统计实验室) University of Cambridge(剑桥大学) Massachusetts Institute of Technology(麻省理工学院) IBM Research USA(IBM美国研究)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments Extra experiments are added in new version

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17041 2025-10-21 cs.CV cs.AI cs.LG 81%

Free$^2$Guide: Training-Free Text-to-Video Alignment using Image LVLM

Jaemin Kim, Bryan Sangwoo Kim, Jong Chul Ye

机构 * Graduate School of AI, KAIST(人工智能研究生院,韩国科学技术院)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments ICCV 2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05831 2025-10-21 cs.CL 79%

Leveraging Robust Optimization for LLM Alignment under Distribution Shifts

Mingye Zhu, Yi Liu, Zheren Fu, Yongdong Zhang, Zhendong Mao

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Communication Content Cognition(通信内容认知国家重点实验室)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15065 2025-10-21 cs.LG 77%

Direct Preference Optimization With Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences

Keertana Chidambaram, Karthik Vinay Seetharaman, Vasilis Syrgkanis

机构 * Stanford University(斯坦福大学)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14400 2025-10-21 cs.CL cs.AI cs.IR 76%

MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question Answering

Yingpeng Ning, Yuanyuan Sun, Ling Luo, Yanhua Wang, Yuchen Pan, Hongfei Lin

机构 * College of Computer Science and Technology, Dalian University of Technology(大连理工大学计算机科学与技术学院) Air Force Communications NCO Academy(空军通信NCO学院)

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI

Comments Accepted as a short paper at BlBM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20548 2025-10-21 cs.LG cs.AI cs.CL 75%

$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training

Jin Peng Zhou, Kaiwen Wang, Jonathan Chang, Zhaolin Gao, Nathan Kallus, Kilian Q. Weinberger, Kianté Brantley, Wen Sun

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12081 2025-10-21 cs.CV cs.AI cs.CL 73%

VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models

Haidong Xu, Guangwei Xu, Zhedong Zheng, Xiatian Zhu, Wei Ji, Xiangtai Li, Ruijie Guo, Meishan Zhang, Min zhang, Hao Fei

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) University of Macau(澳门大学) University of Surrey(Surrey大学) Nanjing University(南京大学) Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI

Comments Accepted by NeurIPS 2025; Project Page: https://walkermitty.github.io/VimoRAG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12943 2025-10-21 cs.CL 57%

The Curious Case of Curiosity across Human Cultures and LLMs

Angana Borah, Zhijing Jin, Rada Mihalcea

机构 * University of Michigan - Ann Arbor(密歇根大学安阿伯分校) University of Toronto(多伦多大学) Vector Institute(向量研究所) MPI for Intelligent Systems, Tubingen, Germany(图宾根德国智能系统研究所)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments Preprint (Paper under review)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15993 2025-10-21 q-fin.PM cs.LG q-fin.ST 57%

Aligning Language Models with Investor and Market Behavior for Financial Recommendations

Fernando Spadea, Oshani Seneviratne

机构 * Rensselaer Polytechnic Institute(拉特兰理工学院)

专题命中 偏好对齐 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06023 2025-10-21 cs.CV 50%

Dual Caption Preference Optimization for Diffusion Models

Amir Saeidi, Yiran Luo, Agneet Chatterjee, Shamanthak Hegde, Bimsara Pathiraja, Yezhou Yang, Chitta Baral

机构 * School of Computing and Augmented Intelligence(计算与增强智能学院) Arizona State University(亚利桑那州立大学)

专题命中 偏好对齐 :DPO(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 8 篇

2510.17402 2025-10-21 cs.CL cs.AI cs.LG 75%

Leveraging Group Relative Policy Optimization to Advance Large Language Models in Traditional Chinese Medicine

Jiacheng Xie, Shuai Zeng, Yang Yu, Xiaoting Tang, Guanghui An, Dong Xu

专题命中 安全训练 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16144 2025-10-21 cs.NI cs.AI cs.MA 70%

Agentic AI for Ultra-Modern Networks: Multi-Agent Framework for RAN Autonomy and Assurance

Sukhdeep Singh, Avinash Bhat, Shweta M, Subhash K Singh, Moonki Hong, Madhan Raj K, Kandeepan Sithamparanathan, Sunder A. Khowaja, Kapal Dev

专题命中 安全训练 :safety(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17369 2025-10-21 cs.RO cs.AI cs.LG 62%

Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots

Haochen Su, Cristian Meo, Francesco Stella, Andrea Peirone, Kai Junge, Josie Hughes

机构 * EPFL(苏黎世联邦理工学院) LatentWorlds AI TUDelft(代尔夫特理工大学) Embodied AI SA(具身人工智能股份有限公司)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by NeurIPS 2025 SpaVLE workshop. 4 pages, 2 figures(in main paper, excluding references and supplements)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15896 2025-10-21 cs.HC cs.AI cs.CY 62%

From Coordination to Personalization: A Trust-Aware Simulation Framework for Emergency Department Decision Support

Zoi Lygizou, Dimitris Kalles

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16255 2025-10-21 cs.CR cs.AI 57%

Detecting Adversarial Fine-tuning with Auditing Agents

Sarah Egler, John Schulman, Nicholas Carlini

机构 * MATS & Anthropic Fellows Program(MATS与Anthropic Fellow项目) Thinking Machines Lab(Thinking Machines实验室) Anthropic

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15937 2025-10-21 q-fin.RM q-fin.TR 50%

Tail-Safe Stochastic-Control SPX-VIX Hedging: A White-Box Bridge Between AI Sensitivities and Arbitrage-Free Market Dynamics

Jian'an Zhang

专题命中 安全训练 :safety(abstract)

Comments 52 pages; 3 figures; PRIMEarxiv template; fully reproducible artifact (code, configs, plots)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00682 2025-10-21 cs.RO 50%

Immersive Explainability: Visualizing Robot Navigation Decisions through XAI Semantic Scene Projections in Virtual Reality

Jorge de Heuvel, Sebastian Müller, Marlene Wessels, Aftab Akhtar, Christian Bauckhage, Maren Bennewitz

机构 * University of Bonn(波恩大学) Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔人工智能与机器学习研究所) Center for Robotics(机器人中心) University of Mainz(美因茨大学) Fraunhofer Institute for Intelligent Analysis and Information Systems IAIS(弗劳恩霍夫智能分析与信息系统研究所)

专题命中 安全训练 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12861 2025-10-21 cs.RO 50%

Safe Multi-Agent Reinforcement Learning for Behavior-Based Cooperative Navigation

Murad Dawood, Sicong Pan, Nils Dengler, Siqi Zhou, Angela P. Schoellig, Maren Bennewitz

机构 * Humanoid Robots Lab, University of Bonn(波恩大学人形机器人实验室)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 6 篇

2510.15430 2025-10-21 cs.CV cs.AI 85%

Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models

Shuang Liang, Zhihao Xu, Jialing Tao, Hui Xue, Xiting Wang

专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);safety(abstract);分类 cs.AI

Comments Withdrawn due to an accidental duplicate submission. This paper (arXiv:2510.15430) was unintentionally submitted as a new entry instead of a new version of our previous work (arXiv:2508.09201)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17006 2025-10-21 cs.CL 83%

Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization

Masahiro Kaneko, Zeerak Talat, Timothy Baldwin

机构 * MBZUAI(马克斯·普朗克人工智能研究所) University of Edinburgh(爱丁堡大学)

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17000 2025-10-21 cs.CR cs.CL cs.LG 73%

Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs

Masahiro Kaneko, Timothy Baldwin

机构 * MBZUAI Abu Dhabi, UAE(阿布扎赫德MBZUAI)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.LG

Comments NeurIPS 2025 (spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19360 2025-10-21 cs.CL cs.AI 62%

Semantic Representation Attack against Aligned Large Language Models

Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang, Shaohui Mei, Lap-Pui Chau

机构 * Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(香港理工大学电子与电气工程系) School of Electronics and Information, Northwestern Polytechnical University(西北工业大学电子与信息学院)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19304 2025-10-21 cs.CY cs.AI cs.CE 62%

Epistemic Trade-Off: An Analysis of the Operational Breakdown and Ontological Limits of "Certainty-Scope" in AI

Generoso Immediato

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.CY

Comments Preprint V3 (October 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16794 2025-10-21 cs.CR cs.LG 57%

Black-box Optimization of LLM Outputs by Asking for Directions

Jie Zhang, Meng Ding, Yang Liu, Jue Hong, Florian Tramèr

机构 * ETH Zurich(苏黎世联邦理工学院) University at Buffalo(布法罗大学) Bytedance, Security Research(字节跳动安全研究)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 提示注入 3 篇

2510.16128 2025-10-21 cs.CR cs.CY 79%

Prompt injections as a tool for preserving identity in GAI image descriptions

Kate Glazko, Jennifer Mankoff

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CY

Comments Accepted as a poster to Soups 2025

Journal ref The Twenty-First Symposium on Usable Privacy and Security (SOUPS 2025) Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06493 2025-10-21 cs.CR cs.AI 70%

System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection

Zongze Li, Jiawei Guo, Haipeng Cai

机构 * University at Buffalo(布法罗大学)

专题命中 提示注入 :jailbreak(abstract);prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02736 2025-10-21 cs.OS cs.SE 50%

AgentSight: System-Level Observability for AI Agents Using eBPF

Yusheng Zheng, Yanpeng Hu, Tong Yu, Andi Quinn

专题命中 提示注入 :prompt injection(abstract)

Journal ref PACMI'2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 6 篇

2510.16257 2025-10-21 cs.CL 79%

Towards Low-Resource Alignment to Diverse Perspectives with Sparse Feedback

Chu Fei Luo, Samuel Dahan, Xiaodan Zhu

机构 * Department of Electrical and Computer Engineering & Ingenuity Labs Research Institute(电气与计算机工程系及创新实验室研究机构) Conflict Analytics Lab, Queen’s University(冲突分析实验室,女王大学) Cornell Law School(康奈尔法学院)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL

Comments Findings of EMNLP 2025, 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17598 2025-10-21 cs.LG cs.AI cs.CL 67%

Hallucination Detection in LLMs Using Spectral Features of Attention Maps

Jakub Binkowski, Denis Janiak, Albert Sawczyn, Bogdan Gabrys, Tomasz Kajdanowicz

机构 * Wroclaw University of Science and Technology(沃拉日-克拉夫大学科学与技术学院) University of Technology Sydney(悉尼技术大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to EMNLP 2025. Code available at https://github.com/graphml-lab-pwr/lapeigvals

详情

展开后加载摘要…

URL PDF HTML 收藏