arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-19 至 2025-11-19 共收录 54 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 7 篇

2509.12179 2025-11-19 cs.AI cs.MA 85%

Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation

Yubo Li, Weiyi Song

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02709 2025-11-19 cs.LG 83%

Preference Robustness for DPO with Applications to Public Health

Cheol Woo Kim, Shresth Verma, Mauricio Tec, Milind Tambe

专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14221 2025-11-19 cs.IR cs.AI 70%

LLM-Aligned Geographic Item Tokenization for Local-Life Recommendation

Hao Jiang, Guoquan Wang, Donglin Zhou, Sheng Yu, Yang Zeng, Wencong Zeng, Kun Gai, Guorui Zhou

机构 * Kuaishou Technology(快手科技)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13765 2025-11-19 cs.LG cs.AI 62%

PROF: An LLM-based Reward Code Preference Optimization Framework for Offline Imitation Learning

Shengjie Sun, Jiafei Lyu, Runze Liu, Mengbei Yan, Bo Liu, Deheng Ye, Xiu Li

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13787 2025-11-19 cs.LG cs.AI 62%

Preference Learning with Lie Detectors can Induce Honesty or Evasion

Chris Cundy, Adam Gleave

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13689 2025-11-19 cs.CL cs.CV 57%

Crossing Borders: A Multimodal Challenge for Indian Poetry Translation and Image Generation

Sofia Jamil, Kotla Sai Charan, Sriparna Saha, Koustava Goswami, Joseph K J

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14760 2025-11-19 cs.CV 50%

UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning

Rui Tian, Mingfei Gao, Haiming Gang, Jiasen Lu, Zhe Gan, Yinfei Yang, Zuxuan Wu, Afshin Dehghan

机构 * Institute of Trustworthy Embodied AI, Fudan University(可信具身人工智能研究院,复旦大学) Apple(苹果公司)

专题命中 偏好对齐 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 7 篇

2511.14433 2025-11-19 cs.LO cs.RO cs.SE 78%

Safe-ROS: An Architecture for Autonomous Robots in Safety-Critical Domains

Diana C. Benjumea, Marie Farrell, Louise A. Dennis

机构 * Department of Computer Science The University of Manchester Manchester, UK(计算机科学系曼彻斯特大学曼彻斯特英国) University of Manchester Manchester, UK(曼彻斯特大学曼彻斯特英国)

专题命中 安全训练 :safety(title,abstract)

Comments In Proceedings FMAS 2025, arXiv:2511.13245

Journal ref EPTCS 436, 2025, pp. 48-68

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14428 2025-11-19 cs.LO cs.AI cs.CY 73%

Context-aware, Ante-hoc Explanations of Driving Behaviour

Dominik Grundt, Ishan Saxena, Malte Petersen, Bernd Westphal, Eike Möhlmann

专题命中 安全训练 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

Comments In Proceedings FMAS 2025, arXiv:2511.13245

Journal ref EPTCS 436, 2025, pp. 114-135

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14435 2025-11-19 cs.SE cs.AI cs.LO 57%

Watchdogs and Oracles: Runtime Verification Meets Large Language Models for Autonomous Systems

Angelo Ferrando

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments In Proceedings FMAS 2025, arXiv:2511.13245

Journal ref EPTCS 436, 2025, pp. 80-87

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13288 2025-11-19 cs.AI 57%

Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO

Haoyang Hong, Jiajun Yin, Yuan Wang, Jingnan Liu, Zhe Chen, Ailing Yu, Ji Li, Zhiling Ye, Hansong Xiao, Yefei Chen, Hualei Zhou, Yun Yue, Minghui Yang, Chunxiao Guo, Junwei Liu, Peng Wei, Jinjie Gu

机构 * Ant Group(蚂蚁集团) Imperial College London(伦敦帝国理工学院)

专题命中 安全训练 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14378 2025-11-19 physics.soc-ph 50%

Emergent Cooperative Driving Strategies for Stop-and-Go Wave Mitigation via Multi-Agent Reinforcement Learning

Raphael Korbmacher, Daniel Straub, Antoine Tordeux, Claudia Totzeck

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11567 2025-11-19 eess.SY cs.RO cs.SY 50%

Who Moved My Distribution? Conformal Prediction for Interactive Multi-Agent Systems

Allen Emmanuel Binny, Anushri Dixit

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12274 2025-11-19 cs.RO 50%

Hierarchical LLMs In-the-Loop Optimization for Real-Time Multi-Robot Target Tracking under Unknown Hazards

Yuwei Wu, Yuezhan Tao, Peihan Li, Guangyao Shi, Gaurav S. Sukhatme, Vijay Kumar, Lifeng Zhou

机构 * GRASP Lab, University of Pennsylvania(宾夕法尼亚大学GRASP实验室) Department of Electrical and Computer Engineering, Drexel University(德雷塞尔大学电气与计算机工程系) Department of Computer Science, University of Southern California(南加州大学计算机科学系)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 3 篇

2511.14423 2025-11-19 cs.CL 85%

Unified Defense for Large Language Models against Jailbreak and Fine-Tuning Attacks in Education

Xin Yi, Yue Li, Dongsheng Shi, Linlin Wang, Xiaoling Wang, Liang He

机构 * Shanghai Institute of Artificial Intelligence for Education, East China Normal University, Shanghai 200062, China(教育人工智能研究所,华东师范大学,上海200062,中国) School of Computer Science and Technology, East China Normal University, Shanghai 200062, China(计算机科学与技术学院,华东师范大学,上海200062,中国) School of Computer Science(计算机科学学院)

专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14140 2025-11-19 cs.CR 82%

Beyond Fixed and Dynamic Prompts: Embedded Jailbreak Templates for Advancing LLM Security

Hajun Kim, Hyunsik Na, Daeseon Choi

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13535 2025-11-19 cs.CV 50%

Accuracy is Not Enough: Poisoning Interpretability in Federated Learning via Color Skew

Farhin Farhad Riya, Shahinul Hoque, Jinyuan Stella Sun, Olivera Kotevska

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 提示注入 2 篇

2508.01249 2025-11-19 cs.CR cs.AI cs.CL cs.LG cs.SE 82%

AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection

Peiran Wang, Yang Liu, Yunfei Lu, Yifeng Cai, Hongbo Chen, Qingyou Yang, Jie Zhang, Jue Hong, Ye Wu

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00447 2025-11-19 cs.CR cs.AI 79%

DRIP: Defending Prompt Injection via Token-wise Representation Editing and Residual Instruction Fusion

Ruofan Liu, Yun Lin, Zhiyong Huang, Jin Song Dong

机构 * National University of Singapore(新加坡国立大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 1 篇

2510.10205 2025-11-19 cs.AI 70%

PIXEL: Adaptive Steering Via Position-wise Injection with eXact Estimated Levels under Subspace Calibration

Manjiang Yu, Hongji Li, Priyanka Singh, Xue Li, Di Wang, Lijie Hu

机构 * University of Queensland(昆士兰大学) Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学) Provable Responsible AI and Data Analytics (PRADA) Lab(可证责任AI与数据分析师实验室) King Abdullah University of Science and Technology (KAUST)(国王 Abdullah 科学与技术大学)

专题命中 幻觉与事实性 :alignment(abstract);trustworthy(abstract);分类 cs.AI

Comments 20 pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 17 篇

2511.11590 2025-11-19 cs.CY cs.AI cs.HC 84%

Embedding Explainable AI in NHS Clinical Safety: The Explainability-Enabled Clinical Safety Framework (ECSF)

Robert Gigiu

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.AI、cs.CY

Comments 33 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02175 2025-11-19 cs.SD cs.CL eess.AS 79%

Hidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment through Latent Acoustic Pattern Triggers

Liang Lin, Miao Yu, Kaiwen Luo, Yibo Zhang, Lilan Peng, Dexian Wang, Xuehai Tang, Yuanhe Zhang, Xikang Yang, Zhenhong Zhou, Kun Wang, Yang Liu

专题命中 安全评测 :alignment(title);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09904 2025-11-19 cs.AI 70%

CTRL-ALT-DECEIT: Sabotage Evaluations for Automated AI R&D

Francis Rhys Ward, Teun van der Weij, Hanna Gábor, Sam Martin, Raja Mehta Moreno, Harel Lidar, Louis Makower, Thomas Jodrell, Lauren Robson

机构 * LawZero Apollo Research Imperial College London(帝国理工学院伦敦校区)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

Comments 53 pages, 21 figures, 8 tables. Accepted as a spotlight at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21359 2025-11-19 cs.CL cs.AI cs.CY cs.LG econ.GN q-fin.EC 70%

Can Machines Think Like Humans? A Behavioral Evaluation of LLM Agents in Dictator Games

Ji Ma

机构 * Gradel Institute of Charity, New College, University of Oxford(格拉德学院,牛津大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23982 2025-11-19 cs.CV cs.RO 67%

StyleDrive: Towards Driving-Style Aware Benchmarking of End-To-End Autonomous Driving

Ruiyang Hao, Bowen Jing, Haibao Yu, Zaiqing Nie

机构 * AIR, Tsinghua University(空气动力学研究所,清华大学) King’s College London(伦敦国王学院) The University of Manchester(曼彻斯特大学) The University of Hong Kong(香港大学)

专题命中 安全评测 :alignment(abstract);safety(abstract)

Comments 25 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23229 2025-11-19 cs.CL cs.AI cs.CY 67%

MCTSr-Zero: Self-Reflective Psychological Counseling Dialogues Generation via Principles and Adaptive Exploration

Hao Lu, Yanchi Gu, Haoyuan Huang, Yulin Zhou, Ningxin Zhu, Chen Li

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 48 pages, 3 figures. Accepted in AAAI-2026 (Main Technical Track). For code and model, see this https://github.com/JianChengXingYun/Mctsr-Zero

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14166 2025-11-19 cs.CL cs.AI 66%

Selective Weak-to-Strong Generalization

Hao Lang, Fei Huang, Yongbin Li

专题命中 安全评测 :alignment(abstract,comments);分类 cs.CL、cs.AI

Comments AAAI2025 Special Track on AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13948 2025-11-19 cs.CV cs.CL cs.LG 62%

EchoAgent: Guideline-Centric Reasoning Agent for Echocardiography Measurement and Interpretation

Matin Daghyani, Lyuyang Wang, Nima Hashemi, Bassant Medhat, Baraa Abdelsamad, Eros Rojas Velez, XiaoXiao Li, Michael Y. C. Tsang, Christina Luong, Teresa S. M. Tsang, Purang Abolmaesumi

机构 * University of British Columbia(不列颠哥伦比亚大学) Vancouver General Hospital(温哥华总医院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

Comments 12 pages, Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13809 2025-11-19 cs.LG cs.AI 62%

ScoresActivation: A New Activation Function for Model Agnostic Global Explainability by Design

Emanuel Covaci, Fabian Galis, Radu Balan, Daniela Zaharie, Darian Onchis

机构 * West University of Timișoara, Romania(蒂米șоara西大学) University of Maryland, United States(马里兰大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Paper submitted to ECAI 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13753 2025-11-19 cs.LG cs.AI cs.CR 62%

Robustness of LLM-enabled vehicle trajectory prediction under data security threats

Feilong Wang, Fuqiang Liu

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 20 pages, 2 figures, 11 tables, working paper

详情

展开后加载摘要…

URL PDF HTML 收藏