arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-19 至 2025-08-19 共收录 58 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 6 篇

2505.19743 2025-08-19 cs.CL cs.LG 86%

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models

Yang Zhang, Yu Yu, Bo Tang, Yu Zhu, Chuxiong Sun, Wenqiang Wei, Jie Hu, Zipeng Xie, Zhiyu Li, Feiyu Xiong, Edward Chung

机构 * Hong Kong Polytechnic University(香港理工大学) MemTensor (Shanghai) Technology Co., Ltd(MemTensor(上海)科技有限公司) University of Science and Technology of China(中国科学技术大学) China Telecom Corporation Limited Beijing Research Institute(中国电信北京研究院) Nanjing University of Information Science and Technology(南京信息工程大学)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.LG

Comments Accepted to 34th International Joint Conference on Artificial Intelligence (IJCAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00845 2025-08-19 cs.LG cs.CL 84%

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment

Yizhuo Zhang, Heng Wang, Shangbin Feng, Zhaoxuan Tan, Xinyun Liu, Yulia Tsvetkov

机构 * University of Washington(华盛顿大学) Xi’an Jiaotong University(西安交通大学) University of Notre Dame(圣母大学) Google(谷歌)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.LG

Comments 8 pages, 1 figures, 2 tables. Experimental code and results are publicly available at https://anonymous.4open.science/r/Graph_RL-BF08/readme.md

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12458 2025-08-19 cs.CL 77%

M3PO: Multimodal-Model-Guided Preference Optimization for Visual Instruction Following

Ruirui Gao, Emily Johnson, Bowen Tan, Yanfei Qian

机构 * University of Massachusetts, Amherst(马萨诸塞大学阿姆赫斯特分校)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07970 2025-08-19 cs.LG cs.AI 62%

WeChat-YATT: A Scalable, Simple, Efficient, and Production Ready Training Library

Junyu Wu, Weiming Chang, Xiaotao Liu, Guanyou He, Tingfeng Xian, Haoqiang Hong, Boqi Chen, Hongtao Tian, Tao Yang, Yunsheng Shi, Feng Lin, Ting Yao, Jiatao Xu

机构 * Tencent(腾讯)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG

Comments arXiv admin note: substantial text overlap with arXiv:2507.22789

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12165 2025-08-19 cs.AI 57%

RLNVR: Reinforcement Learning from Non-Verified Real-World Rewards

Rohit Krishnan, Jon Evans

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15714 2025-08-19 cs.CL 57%

Chinchunmei at SemEval-2025 Task 11: Boosting the Large Language Model's Capability of Emotion Perception using Contrastive Learning

Tian Li, Yujian Sun, Huizhi Liang

机构 * School of Computing, Newcastle University(计算学院,新卡斯尔大学) Shumei AI Research Institute(舒美人工智能研究院)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL

Journal ref Association for Computational Linguistics, Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), 319-330

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 5 篇

2508.12531 2025-08-19 cs.LG cs.AI 81%

Rethinking Safety in LLM Fine-tuning: An Optimization Perspective

Minseon Kim, Jin Myung Kwak, Lama Alssum, Bernard Ghanem, Philip Torr, David Krueger, Fazl Barez, Adel Bibi

机构 * Microsoft Research(微软研究院) KAIST(韩国科学技术院) KAUST(韩国科学技术大学) Université de Montréal(蒙特利尔大学) Mila(蒙特利尔人工智能研究院) WhiteBox(WhiteBox公司) University of Oxford(牛津大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14850 2025-08-19 cs.LG cs.AI cs.RO 81%

Hierarchical Multi-Agent Reinforcement Learning with Control Barrier Functions for Safety-Critical Autonomous Systems

H. M. Sabbir Ahmad, Ehsan Sabouni, Alexander Wasilkoff, Param Budhraja, Zijian Guo, Songyuan Zhang, Chuchu Fan, Christos Cassandras, Wenchao Li

机构 * Boston University(波士顿大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12920 2025-08-19 cs.AI cs.MA 70%

Do Large Language Model Agents Exhibit a Survival Instinct? An Empirical Study in a Sugarscape-Style Simulation

Atsushi Masumori, Takashi Ikegami

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02903 2025-08-19 cs.CE cs.LG cs.NA math.NA 57%

Predicting Open-Hole Laminates Failure Using Support Vector Machines With Classical and Quantum Kernels

Giorgio Tosti Balducci, Boyang Chen, Matthias Möller, Marc Gerritsma, Roeland De Breuker

机构 * Delft University of Technology(代尔夫特理工大学)

专题命中 安全训练 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11799 2025-08-19 math.OC cs.RO 50%

Scaling Robust Optimization for Swarms: A Distributed Perspective

Arshiya Taj Abdul, Augustinos D. Saravanos, Evangelos A. Theodorou

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 4 篇

2411.01077 2025-08-19 cs.CL cs.LG 81%

Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection

Zhipeng Wei, Yuqi Liu, N. Benjamin Erichson

机构 * International Computer Science Institute, CA, USA(国际计算机科学研究所) Lawrence Berkeley National Laboratory, CA, USA(伯克利国家实验室)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05934 2025-08-19 cs.CR cs.AI 79%

Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models

Ma Teng, Jia Xiaojun, Duan Ranjie, Li Xinfeng, Huang Yihao, Jia Xiaoshuang, Chu Zhixuan, Ren Wenqi

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20841 2025-08-19 cs.CL 57%

Concealment of Intent: A Game-Theoretic Analysis

Xinbo Wu, Abhishek Umrawal, Lav R. Varshney

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13048 2025-08-19 cs.CR 50%

MAJIC: Markovian Adaptive Jailbreaking via Iterative Composition of Diverse Innovative Strategies

Weiwei Qi, Shuo Shao, Wei Gu, Tianhang Zheng, Puning Zhao, Zhan Qin, Kui Ren

专题命中 越狱攻击 :jailbreak(abstract)

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 提示注入 1 篇

2508.12175 2025-08-19 cs.CR 50%

Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous

Ben Nassi, Stav Cohen, Or Yair

专题命中 提示注入 :prompt injection(abstract)

Comments https://sites.google.com/view/invitation-is-all-you-need/home

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 1 篇

2506.12609 2025-08-19 cs.CV 50%

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation

Lexiang Tang, Xianwei Zhuang, Bang Yang, Zhiyuan Hu, Hongxiang Li, Lu Ma, Jinghan Ru, Yuexian Zou

专题命中 幻觉与事实性 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 隐私与版权 3 篇

2508.12727 2025-08-19 cs.LG 79%

FedSODA: Federated Fine-tuning of LLMs via Similarity Group Pruning and Orchestrated Distillation Alignment

Manning Zhu, Songtao Guo, Pengzhan Zhou, Yansong Ning, Chang Han, Dewen Qiao

机构 * Chongqing University(重庆大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Third Military Medical University(第三军医大学)

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12158 2025-08-19 cs.CL 74%

LLM-as-a-Judge for Privacy Evaluation? Exploring the Alignment of Human and LLM Perceptions of Privacy in Textual Data

Stephen Meisenbacher, Alexandra Klymenko, Florian Matthes

机构 * Technical University of Munich School of Computation, Information(慕尼黑技术大学计算、信息与技术学院)

专题命中 隐私与版权 :alignment(title);分类 cs.CL

Comments 13 pages, 3 figures, 4 tables. Accepted to HAIPS @ CCS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11690 2025-08-19 cs.CY cs.AI 62%

Real Time Child Abduction And Detection System

Tadisetty Sai Yashwanth, Yangalasetty Sruthi Royal, Vankayala Rajeshwari Shreya, Mayank Kashyap, Divyaprabha K N

机构 * Dept. of CSE PES University Bangalore, India(计算机科学与工程系,PES大学,印度班加罗尔)

专题命中 隐私与版权 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 安全评测 20 篇

2505.11247 2025-08-19 cs.AI cs.LG cs.RO 81%

LD-Scene: LLM-Guided Diffusion for Controllable Generation of Adversarial Safety-Critical Driving Scenarios

Mingxing Peng, Yuting Xie, Xusen Guo, Ruoyu Yao, Hai Yang, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12624 2025-08-19 cs.CL cs.AI 81%

Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, Dieuwke Hupkes

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Meta

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments https://aclanthology.org/2025.gem-1.33/

Journal ref Proceedings of the Fourth Workshop on Generation Evaluation and Metrics GEM2 2025 pages 404 to 430; July 31 August 1 2025; 2025 Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10872 2025-08-19 cs.CV cs.AI cs.ET 79%

V-RoAst: Visual Road Assessment. Can VLM be a Road Safety Assessor Using the iRAP Standard?

Natchapon Jongwiriyanurak, Zichao Zeng, June Moh Goo, Xinglei Wang, Ilya Ilyankou, Kerkritt Sriroongvikrai, Nicola Christie, Meihui Wang, Huanfa Chen, James Haworth

机构 * University College London(伦敦大学学院) Chulalongkorn University(朱拉隆梭大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01408 2025-08-19 cs.RO cs.CV 78%

From Shadows to Safety: Occlusion Tracking and Risk Mitigation for Urban Autonomous Driving

Korbinian Moller, Luis Schwarzmeier, Johannes Betz

专题命中 安全评测 :safety(title,abstract)

Comments 8 Pages. Submitted to the IEEE Intelligent Vehicles Symposium (IV 2025), Romania

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16867 2025-08-19 cs.CV 78%

ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering

Kaisi Guan, Zhengfeng Lai, Yuchong Sun, Peng Zhang, Wei Liu, Kieran Liu, Meng Cao, Ruihua Song

机构 * Renmin University of China(中国人民大学) Apple(苹果公司)

专题命中 安全评测 :alignment(title,abstract)

Comments International Conference on Computer Vision 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11957 2025-08-19 cs.MA cs.AI cs.LG 73%

A Comprehensive Review of AI Agents: Transforming Possibilities in Technology and Beyond

Xiaodong Qu, Andrews Damoah, Joshua Sherwood, Peiyan Liu, Christian Shun Jin, Lulu Chen, Minjie Shen, Nawwaf Aleisa, Zeyuan Hou, Chenyu Zhang, Lifu Gao, Yanshu Li, Qikai Yang, Qun Wang, Cristabelle De Souza

机构 * University of Maryland, College Park(马里兰大学学院公园分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Brown University(布朗大学) San Francisco State University(旧金山州立大学) Stanford University(斯坦福大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06225 2025-08-19 cs.AI 70%

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution

Zailong Tian, Zhuoheng Han, Yanzhe Chen, Haozhe Xu, Xi Yang, Richeng Xuan, Houfeng Wang, Lizi Liao

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11647 2025-08-19 cs.LO cs.AI 70%

Categorical Construction of Logically Verifiable Neural Architectures

Logan Nye

机构 * Carnegie Mellon University School of Computer Science(卡内基梅隆大学计算机科学学院)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13152 2025-08-19 cs.CL cs.AI 62%

RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns

Xin Chen, Junchao Wu, Shu Yang, Runzhe Zhan, Zeyu Wu, Ziyang Luo, Di Wang, Min Yang, Lidia S. Chao, Derek F. Wong

机构 * NLP(自然语言处理) CT Lab, Department of Computer and Information Science, University of Macau(计算机与信息科学系,澳门大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Provable Responsible AI and Data Analytics Lab, KAUST(可证明责任AI与数据分析实验室,卡尔斯兰大学) Hong Kong Baptist University(香港 Baptist 大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted to TACL 2025. This version is a pre-MIT Press publication version

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12212 2025-08-19 cs.LG cs.AI q-bio.QM 62%

ProtTeX-CC: Activating In-Context Learning in Protein LLM via Two-Stage Instruction Compression

Chuanliu Fan, Zicheng Ma, Jun Gao, Nan Yu, Jun Zhang, Ziqiang Cao, Yi Qin Gao, Guohong Fu

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏