arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-08 至 2025-09-08 共收录 45 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2509.04713 2025-09-08 cs.LG 79%

Natural Spectral Fusion: p-Exponent Cyclic Scheduling and Early Decision-Boundary Alignment in First-Order Optimization

Gongyue Zhang, Honghai Liu

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20309 2025-09-08 cs.CV 78%

Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs

Zitian Wang, Yue Liao, Kang Rong, Fengyun Rao, Yibo Yang, Si Liu

机构 * Beihang University(北航大学) National University of Singapore(国立新加坡大学) King Abdullah University of Science and Technology(国王 Abdullah 科学与技术大学)

专题命中 偏好对齐 :alignment(title,abstract)

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03730 2025-09-08 cs.AI cs.CL cs.CY cs.LG stat.ML 77%

The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs

Pengrui Han, Rafal Kocielnik, Peiyang Song, Ramit Debnath, Dean Mobbs, Anima Anandkumar, R. Michael Alvarez

机构 * Caltech(加州理工学院) UIUC(伊利诺伊大学) University of Cambridge(剑桥大学)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.CY

Comments We make public all code and source data at https://github.com/psychology-of-AI/Personality-Illusion for full reproducibility

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05946 2025-09-08 eess.SY cs.SY 50%

InstructMPC: A Human-LLM-in-the-Loop Framework for Context-Aware Control

Ruixiang Wu, Jiahao Ai, Tongxin Li

专题命中 偏好对齐 :DPO(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 5 篇

2509.04512 2025-09-08 cs.CL cs.LG 84%

Scaling behavior of large language models in emotional safety classification across sizes and tasks

Edoardo Pinzuti, Oliver Tüscher, André Ferreira Castro

机构 * Leibniz Institute for Resilience Research(莱比锡韧性研究所) University Medical Center Halle(哈雷医学院) German Center for Mental Health (DZPG)(德国心理健康中心(DZPG)) University Medical Center of the Johannes Gutenberg-University Mainz(美因茨约瑟夫·冯·拉贝大学医学院)

专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12391 2025-09-08 cs.LG 79%

Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL

Junyu Guo, Zhi Zheng, Donghao Ying, Ming Jin, Shangding Gu, Costas Spanos, Javad Lavaei

机构 * University of California Berkeley(加州大学伯克利分校) Virginia Tech(弗吉尼亚理工大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04535 2025-09-08 cs.RO cs.AI cs.LG 62%

In-Context Policy Adaptation via Cross-Domain Skill Diffusion

Minjong Yoo, Woo Kyung Kim, Honguk Woo

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

Comments 9 pages

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04478 2025-09-08 cs.CL 57%

An End-to-End System for Culturally-Attuned Driving Feedback using a Dual-Component NLG Engine

Iniakpokeikiye Peter Thompson, Yi Dewei, Reiter Ehud

机构 * Dept. of Computing Science University of Aberdeen(计算科学系阿伯丁大学)

专题命中 安全训练 :safety(abstract);分类 cs.CL

Comments The paper has 5 figures and 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04714 2025-09-08 cs.SI 50%

ThumbnailTruth: A Multi-Modal LLM Approach for Detecting Misleading YouTube Thumbnails Across Diverse Cultural Settings

Wajiha Naveed, Zartash Afzal Uzmi, Zafar Ayyub Qazi

专题命中 安全训练 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 1 篇

2412.13341 2025-09-08 cs.LG cs.CR 70%

Concept-ROT: Poisoning Concepts in Large Language Models with Model Editing

Keltin Grimes, Marco Christiani, David Shriver, Marissa Connor

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.LG

Comments Published at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 红队测试 1 篇

2509.04615 2025-09-08 cs.CL cs.CR cs.LG 62%

Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs

Brennen Hill, Surendra Parla, Venkata Abhijeeth Balabhadruni, Atharv Prajod Padmalayam, Sujay Chandra Shekara Sharma

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 红队测试 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 5 篇

2508.15442 2025-09-08 eess.AS cs.AI cs.SD 79%

Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets

Chenlin Liu, Minghui Fang, Patrick Zhang, Wei Zhou, Jie Gao, Jiqing Han

机构 * Harbin Institute of Technology, China(哈尔滨工业大学) Zhejiang University, China(浙江大学) Tsinghua University, Shenzhen, China(清华大学深圳研究院)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

Comments Accepted to EMNLP 2025 Main Conference (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15298 2025-09-08 cs.CV 78%

TPA: Temporal Prompt Alignment for Fetal Congenital Heart Defect Classification

Darya Taratynova, Alya Almsouti, Beknur Kalmakhanbet, Numan Saeed, Mohammad Yaqub

机构 * Department of Machine Learning Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) Abu Dhabi, UAE(机器学习系,Mohamed bin Zayed人工智能大学(MBZUAI),阿布扎比,阿联酋) Department of Computer Vision Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) Abu Dhabi, UAE(计算机视觉系,Mohamed bin Zayed人工智能大学(MBZUAI),阿布扎比,阿联酋)

专题命中 幻觉与事实性 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00461 2025-09-08 cs.CL cs.AI 62%

TECP: Token-Entropy Conformal Prediction for LLMs

Beining Xu, Yongming Lu

机构 * Department of School of Engineering, Shenzhen MSU-BIT University, Shenzhen, China, 518000(深圳MSU-BIT大学工程学院部门) MSU-BIT-SMBU Joint Research Center of Applied Mathematics, Shenzhen MSU-BIT University, Shenzhen, China, 518000(应用数学联合研究中心)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04735 2025-09-08 cs.CV cs.AI 57%

Enhancing Self-Driving Segmentation in Adverse Weather Conditions: A Dual Uncertainty-Aware Training Approach to SAM Optimization

Dharsan Ravindran, Kevin Wang, Zhuoyuan Cao, Saleh Abdelrahman, Jeffery Wu

机构 * Queen's University(女王大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04664 2025-09-08 cs.CL 57%

Why Language Models Hallucinate

Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, Edwin Zhang

机构 * OpenAI Georgia Tech(佐治亚理工学院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 17 篇

2408.09600 2025-09-08 cs.AI cs.CR 88%

Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning

Tiansheng Huang, Gautam Bhattacharya, Pratik Joshi, Josh Kimball, Ling Liu

机构 * Georgia Institute of Technology(佐治亚理工学院) Dolby Laboratories(杜比实验室)

专题命中 安全评测 :alignment(title,abstract);safety(title,abstract);分类 cs.AI

Comments Rejected by AAAI25-AIA. Accepted by ICML25. Authors are thankful to the anonymous reviewers from both AAAI25-AIA and ICML25

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08613 2025-09-08 cs.CL 79%

Assessing the Sensitivity and Alignment of FOL Closeness Metrics

Ramya Keerthy Thatikonda, Wray Buntine, Ehsan Shareghi

机构 * Department of Data Science & AI, Monash University(数据科学与人工智能系,莫纳什大学) College of Engineering and Computer Science, VinUniversity(工程与计算机科学学院,文大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04536 2025-09-08 cs.LG math.QA math.ST stat.TH 79%

Q-SafeML: Safety Assessment of Quantum Machine Learning via Quantum Distance Metrics

Oliver Dunn, Koorosh Aslansefat, Yiannis Papadopoulos

机构 * University of Hull(赫尔大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17114 2025-09-08 cs.CL cs.CV cs.LG cs.MM 76%

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language

Subrata Biswas, Mohammad Nur Hossain Khan, Bashima Islam

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05042 2025-09-08 cs.RO 75%

Shared Autonomy through LLMs and Reinforcement Learning for Applications to Ship Hull Inspections

Cristiano Caissutti, Estelle Gerbier, Ehsan Khorrambakht, Paolo Marinelli, Andrea Munafo', Andrea Caiti

机构 * Dept. of Information Engineering University of Pisa, Italy Pisa, Italy(信息工程系 乌迪内大学 乌迪内,意大利)

专题命中 安全评测 :alignment(abstract);safety(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04549 2025-09-08 cs.CL cs.AI 73%

Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions

Faruk Alpay, Taylan Alpay

机构 * Lightcap Department of Future(未来系) Turkish Aeronautical Association(土耳其航空航天协会)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05199 2025-09-08 cs.CL 70%

Triadic Fusion of Cognitive, Functional, and Causal Dimensions for Explainable LLMs: The TAXAL Framework

David Herrera-Poyatos, Carlos Peláez-González, Cristina Zuheros, Virilo Tejedor, Rosana Montes, Francisco Herrera

机构 * Department of Computer Science and Artificial Intelligence(计算机科学与人工智能系) Andalusian Institute of Data Science and Computational Intelligence(安达卢西亚数据科学与计算智能研究所) University of Granada(格拉纳达大学)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.CL

Comments 27 pages, 9 tables and 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04979 2025-09-08 cs.AI 70%

Internet 3.0: Architecture for a Web-of-Agents with it's Algorithm for Ranking Agents

Rajesh Tembarai Krishnamachari, Srividya Rajesh

机构 * NYU(纽约大学) Independent Researcher(独立研究者)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20986 2025-09-08 cs.CY econ.TH 70%

MAD Chairs: A new tool to evaluate AI

Chris Santos-Lang

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.CY

Comments 17 pages, 1 figure, reproduced with permission from Springer Nature from Coordination, Organizations, Institutions, Norms, and Ethics for Governance of Multi-Agent Systems XVIII (COINE 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04657 2025-09-08 cs.CL cs.AI cs.DB cs.LG 67%

Evaluating NL2SQL via SQL2NL

Mohammadtaher Safarzadeh, Afshin Oroojlooyjadid, Dan Roth

机构 * Oracle AI

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04752 2025-09-08 cs.HC cs.AI cs.LG 62%

SePA: A Search-enhanced Predictive Agent for Personalized Health Coaching

Melik Ozolcer, Sang Won Bae

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted at IEEE-EMBS International Conference on Biomedical and Health Informatics (BHI'25). 7 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05528 2025-09-08 cs.AI cs.CL 62%

Conversational Education at Scale: A Multi-LLM Agent Workflow for Procedural Learning and Pedagogic Quality Assessment

Jiahuan Pei, Fanghua Ye, Xin Sun, Wentao Deng, Koen Hindriks, Junxiao Wang

机构 * Vrije University of Amsterdam(阿姆斯特丹自由大学) University College London(伦敦大学学院) University of Amsterdam(阿姆斯特丹大学) National Institute of Informatics(日本信息处理学会) Shandong University(山东大学) Guangzhou University(广州大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 14 pages, accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23903 2025-09-08 cs.AI cs.LG 62%

Neural Network Verification with PyRAT

Augustin Lemesle, Julien Lehmann, Tristan Le Gall

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04504 2025-09-08 cs.CL cs.AI 62%

Behavioral Fingerprinting of Large Language Models

Zehua Pei, Hui-Ling Zhen, Ying Zhang, Zhiyuan Yang, Xing Li, Xianzhi Yu, Mingxuan Yuan, Bei Yu

机构 * The Chinese University of Hong Kong(香港中文大学) Noah’s Ark Lab, Huawei(华为诺亚实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Submitted to 1st Open Conference on AI Agents for Science (agents4science 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏