arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-29 至 2025-08-29 共收录 31 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3 篇

2508.21016 2025-08-29 cs.LG cs.AI 84%

Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance

Luozhijie Jin, Zijie Qiu, Jie Liu, Zijie Diao, Lifeng Qiao, Ning Ding, Alex Lamb, Xipeng Qiu

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20181 2025-08-29 cs.CV cs.AI cs.CL cs.MM 73%

Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization

Alberto Compagnoni, Davide Caffagni, Nicholas Moratelli, Lorenzo Baraldi, Marcella Cornia, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI

Comments BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07818 2025-08-29 cs.CV 67%

DanceGRPO: Unleashing GRPO on Visual Generation

Zeyue Xue, Jie Wu, Yu Gao, Fangyuan Kong, Lingting Zhu, Mengzhao Chen, Zhiheng Liu, Wei Liu, Qiushan Guo, Weilin Huang, Ping Luo

机构 * The University of Hong Kong(香港大学)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract)

Comments Project Page: https://dancegrpo.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 4 篇

2508.20766 2025-08-29 cs.CL cs.AI cs.LG 89%

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection

Harethah Abu Shairah, Hasan Abed Al Kader Hammoud, George Turkiyyah, Bernard Ghanem

机构 * King Abdullah University of Science and Technology (KAUST)(卡布尔大学科学与技术学院)

专题命中 安全训练 :alignment(title,abstract);safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16441 2025-08-29 cs.RO cs.AI math.GN 79%

Safe and Efficient Social Navigation through Explainable Safety Regions Based on Topological Features

Victor Toscano-Duran, Sara Narteni, Alberto Carlevaro, Jérôme Guzzi Rocio Gonzalez-Diaz, Maurizio Mongelli

机构 * Department of Applied Mathematics I, University of Seville(应用数学系,塞维利亚大学) CNR-IEIIT Genoa, Italy(意大利热那亚CNR-IEIIT) SUPSI, IDSIA Lugano, Switzerland(瑞士卢加诺SUPSI,IDSIA)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20151 2025-08-29 cs.AI 70%

IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement

Yuanzhe Shen, Zisu Huang, Zhengkang Guo, Yide Liu, Guanxu Chen, Ruicheng Yin, Xiaoqing Zheng, Xuanjing Huang

专题命中 安全训练 :safety(abstract);jailbreak(abstract);分类 cs.AI

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02997 2025-08-29 cs.CL 57%

CoCoTen: Detecting Adversarial Inputs to Large Language Models through Latent Space Features of Contextual Co-occurrence Tensors

Sri Durga Sai Sowmya Kadali, Evangelos E. Papalexakis

机构 * University of California, Riverside(加州大学河滨分校)

专题命中 安全训练 :jailbreak(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 3 篇

2503.06989 2025-08-29 cs.CR cs.CV 85%

Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application

Wenzhuo Xu, Zhipeng Wei, Xiongtao Sun, Zonghao Ying, Deyue Zhang, Dongdong Yang, Xiangzheng Zhang, Quanchen Zou

专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22957 2025-08-29 cs.CL cs.AI cs.CY cs.MA 80%

Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models

Younwoo Choi, Changling Li, Yongjin Yang, Zhijing Jin

专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20848 2025-08-29 cs.CR cs.AI 79%

JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring

Junjie Chu, Mingjie Li, Ziqing Yang, Ye Leng, Chenhao Lin, Chao Shen, Michael Backes, Yun Shen, Yang Zhang

机构 * CISPA Helmholtz Center for Information Security(CISPA海德堡信息安全中心) Xi’an Jiaotong University(西安交通大学)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI

Comments 17 pages, 5 figures. For the code and data supporting this work, see https://trustairlab.github.io/jades.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 红队测试 1 篇

2508.20411 2025-08-29 cs.AI cs.CR cs.CY 86%

Governable AI: Provable Safety Under Extreme Threat Models

Donglin Wang, Weiyun Liang, Chunyuan Chen, Jing Xu, Yulong Fu

机构 * College of Artificial Intelligence, Nankai University(人工智能学院,南开大学) Sursen Corp.(苏尔森公司) China Academy of Electronics and Information Technology(中国电子信息技术研究院)

专题命中 红队测试 :safety(title,abstract);alignment(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 2 篇

2508.11030 2025-08-29 cs.HC 78%

Families' Vision of Generative AI Agents for Household Safety Against Digital and Physical Threats

Zikai Wen, Lanjing Liu, Yaxing Yao

专题命中 隐私与版权 :safety(title,abstract)

Comments Accepted in Proc. ACM Hum.-Comput. Interact. 9, 7, Article CSCW

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18674 2025-08-29 cs.CV 50%

Image-guided topic modeling for interpretable privacy classification

Alina Elena Baia, Andrea Cavallaro

专题命中 隐私与版权 :alignment(abstract)

Comments Paper accepted at the eXCV Workshop at ECCV 2024. Supplementary material included. Code available at https://github.com/idiap/itm

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 7 篇

2508.20776 2025-08-29 cs.CV cs.AI 70%

Safer Skin Lesion Classification with Global Class Activation Probability Map Evaluation and SafeML

Kuniko Paxton, Koorosh Aslansefat, Amila Akagić, Dhavalkumar Thakker, Yiannis Papadopoulos

机构 * School of Computer Science, University of Hull(赫尔大学计算机科学学院) Faculty of Electrical Engineering, University of Sarajevo(萨拉热窝大学电气工程学院)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20737 2025-08-29 cs.SE cs.AI 70%

Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol

Wei Ma, Yixiao Yang, Qiang Hu, Shi Ying, Zhi Jin, Bo Du, Zhenchang Xing, Tianlin Li, Junjie Shi, Yang Liu, Linxiao Jiang

机构 * Singapore Management University Singapore Capital Normal University Beijing China Tianjin University Tianjin China Wuhan University China CSIRO's Data61 \& Australian National University Australia Nanyang Technological University Singapore Singapore Management University Capital Normal University Tianjin University Wuhan University CSIRO's Data61 \& Australian National University Nanyang Technological University

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18076 2025-08-29 cs.CL 70%

Neither Valid nor Reliable? Investigating the Use of LLMs as Judges

Khaoula Chehbouni, Mohammed Haddou, Jackie Chi Kit Cheung, Golnoosh Farnadi

机构 * McGill University(麦吉尔大学) Mila - Quebec AI Institute(魁北克AI研究所) Statistics Canada(加拿大统计局)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL

Comments Prepared for conference submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21061 2025-08-29 cs.HC cs.AI cs.LG 62%

OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language Models

Adam Coscia, Shunan Guo, Eunyee Koh, Alex Endert

机构 * Georgia Institute of Technology(佐治亚理工学院) Adobe Research(Adobe研究)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to UIST 2025. 18 pages, 9 figures, 2 tables. For a demo video, see https://youtu.be/uobhmxo6EIE

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20416 2025-08-29 cs.CL cs.AI 62%

DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding

Hengchuan Zhu, Yihuan Xu, Yichen Li, Zijie Meng, Zuozhu Liu

机构 * Zhejiang University(浙江大学) ZJU-Angelalign R&D Center for Intelligence Healthcare(浙江大学智能医疗研发中心)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20288 2025-08-29 eess.SY cs.LG cs.SY 57%

Neural Spline Operators for Risk Quantification in Stochastic Systems

Zhuoyuan Wang, Raffaele Romagnoli, Kamyar Azizzadenesheli, Yorie Nakahira

机构 * Department of Electrical and Computering Engineering, Carnegie Mellon University(电气与计算机工程系,卡内基梅隆大学) School of Science and Engineering, Department of Mathematics and Computer Science, Duquesne University(科学与工程学院,数学与计算机科学系,杜克森大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20851 2025-08-29 cs.CV 50%

PathMR: Multimodal Visual Reasoning for Interpretable Pathology Diagnosis

Ye Zhang, Yu Zhou, Jingwen Qi, Yongbing Zhang, Simon Puettmann, Finn Wichmann, Larissa Pereira Ferreira, Lara Sichward, Julius Keyl, Sylvia Hartmann, Shuo Zhao, Hongxiao Wang, Xiaowei Xu, Jianxu Chen

机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) Leibniz-Institut für Analytische Wissenschaften – ISAS – e.V.(莱比锡分析科学研究所(ISAS)) Department of Pathology, The Sixth Affiliated Hospital, Sun Yat-sen University(中山大学第六附属医院病理科部) Institute of Pathology, University Hospital Essen(埃森大学医院病理科研究所) Academy for Multidisciplinary Studies, Capital Normal University(首都师范大学多学科研究学院)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 1 篇

2508.20333 2025-08-29 cs.LG cs.AI cs.CL cs.DC 85%

Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs

Md Abdullah Al Mamun, Ihsen Alouani, Nael Abu-Ghazaleh

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 10 篇

2508.20130 2025-08-29 q-bio.QM cs.AI cs.LG 81%

Artificial Intelligence for CRISPR Guide RNA Design: Explainable Models and Off-Target Safety

Alireza Abbaszadeh, Armita Shahlai

机构 * Department of Computer Engineering, Ma.C., Islamic Azad University(计算机工程系,伊斯兰阿兹德大学) Department of Biological Sciences and Technologies, Faculty of Basic Sciences, Islamic Azad University(基础科学学院生物科学与技术系,伊斯兰阿兹德大学)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 29 pages, 5 figures, 2 tables, 42 cited references

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20465 2025-08-29 q-bio.NC 78%

On the possibility of deep alignment

Alex B. Kiefer

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19512 2025-08-29 cs.CL 70%

Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging

Hua Farn, Hsuan Su, Shachi H Kumar, Saurav Sahay, Shang-Tse Chen, Hung-yi Lee

机构 * National Taiwan University(国立台湾大学) Intel Lab(英特尔实验室)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20217 2025-08-29 cs.CL cs.AI 62%

Prompting Strategies for Language Model-Based Item Generation in K-12 Education: Bridging the Gap Between Small and Large Language Models

Mohammad Amini, Babak Ahmadi, Xiaomeng Xiong, Yilin Zhang, Christopher Qiao

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20722 2025-08-29 cs.CL 57%

rStar2-Agent: Agentic Reasoning Technical Report

Ning Shang, Yifei Liu, Yi Zhu, Li Lyna Zhang, Weijiang Xu, Xinyu Guan, Buze Zhang, Bingcheng Dong, Xudong Zhou, Bowen Zhang, Ying Xin, Ziming Miao, Scarlett Li, Fan Yang, Mao Yang

机构 * Microsoft Research(微软研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20374 2025-08-29 cs.AI 57%

TCIA: A Task-Centric Instruction Augmentation Method for Instruction Finetuning

Simin Ma, Shujian Liu, Jun Tan, Yebowen Hu, Song Wang, Sathish Reddy Indurthi, Sanqiang Zhao, Liwei Wu, Jianbing Han, Kaiqiang Song

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15576 2025-08-29 cs.CV cs.LG 57%

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models

Xin Huang, Ruibin Li, Tong Jia, Wei Zheng, Ya Wang

机构 * School of Artificial Intelligence and Software Engineering, Nanyang Normal University, Henan, China(人工智能与软件工程学院,南阳师范学院,河南) Institute for Artificial Intelligence, Peking University, Beijing, China(人工智能研究院,北京大学,北京) Collaborative Innovation Center of Intelligent Explosion-proof Equipment, Henan, China(智能防爆设备协同创新中心,河南)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Accepted at the International Joint Conference on Artificial Intelligence (IJCAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00258 2025-08-29 cs.AI q-bio.NC 57%

Possible Principles for Aligned Structure Learning Agents

Lancelot Da Costa, Tomáš Gavenčiak, David Hyland, Mandana Samiei, Cristian Dragos-Manta, Candice Pattisapu, Adeel Razi, Karl Friston

机构 * VERSES AI Research Lab(VERSES AI研究实验室) Charles University(查尔斯大学) University of Oxford(牛津大学) Mila, Quebec AI Institute(魁北克人工智能研究院) McGill University(麦吉尔大学) University of Montreal(蒙特利尔大学) University College London(伦敦大学学院) Monash University(莫纳什大学) CIFAR Azrieli Global Scholars Program(CIFAR阿兹里埃利全球学者计划)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 24 pages of content, 33 with references; accepted version

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20623 2025-08-29 cs.CV 50%

AvatarBack: Back-Head Generation for Complete 3D Avatars from Front-View Images

Shiqi Xin, Xiaolin Zhang, Yanbin Liu, Peng Zhang, Caifeng Shan

机构 * College of Electrical Engineering and Automation, Shandong University of Science and Technology(山东科技大学电气工程与自动化学院) Department of Data Science and Artificial Intelligence, Auckland University of Technology(奥克兰大学数据科学与人工智能系) College of Computer Science and Engineering, Shandong University of Science and Technology(山东科技大学计算机科学与工程学院)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏