arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-19 至 2025-11-19 共收录 7 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 7 篇

2511.14433 2025-11-19 cs.LO cs.RO cs.SE 78%

Safe-ROS: An Architecture for Autonomous Robots in Safety-Critical Domains

Diana C. Benjumea, Marie Farrell, Louise A. Dennis

机构 * Department of Computer Science The University of Manchester Manchester, UK(计算机科学系曼彻斯特大学曼彻斯特英国) University of Manchester Manchester, UK(曼彻斯特大学曼彻斯特英国)

专题命中 安全训练 :safety(title,abstract)

Comments In Proceedings FMAS 2025, arXiv:2511.13245

Journal ref EPTCS 436, 2025, pp. 48-68

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14428 2025-11-19 cs.LO cs.AI cs.CY 73%

Context-aware, Ante-hoc Explanations of Driving Behaviour

Dominik Grundt, Ishan Saxena, Malte Petersen, Bernd Westphal, Eike Möhlmann

专题命中 安全训练 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

Comments In Proceedings FMAS 2025, arXiv:2511.13245

Journal ref EPTCS 436, 2025, pp. 114-135

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14435 2025-11-19 cs.SE cs.AI cs.LO 57%

Watchdogs and Oracles: Runtime Verification Meets Large Language Models for Autonomous Systems

Angelo Ferrando

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments In Proceedings FMAS 2025, arXiv:2511.13245

Journal ref EPTCS 436, 2025, pp. 80-87

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13288 2025-11-19 cs.AI 57%

Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO

Haoyang Hong, Jiajun Yin, Yuan Wang, Jingnan Liu, Zhe Chen, Ailing Yu, Ji Li, Zhiling Ye, Hansong Xiao, Yefei Chen, Hualei Zhou, Yun Yue, Minghui Yang, Chunxiao Guo, Junwei Liu, Peng Wei, Jinjie Gu

机构 * Ant Group(蚂蚁集团) Imperial College London(伦敦帝国理工学院)

专题命中 安全训练 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14378 2025-11-19 physics.soc-ph 50%

Emergent Cooperative Driving Strategies for Stop-and-Go Wave Mitigation via Multi-Agent Reinforcement Learning

Raphael Korbmacher, Daniel Straub, Antoine Tordeux, Claudia Totzeck

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11567 2025-11-19 eess.SY cs.RO cs.SY 50%

Who Moved My Distribution? Conformal Prediction for Interactive Multi-Agent Systems

Allen Emmanuel Binny, Anushri Dixit

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12274 2025-11-19 cs.RO 50%

Hierarchical LLMs In-the-Loop Optimization for Real-Time Multi-Robot Target Tracking under Unknown Hazards

Yuwei Wu, Yuezhan Tao, Peihan Li, Guangyao Shi, Gaurav S. Sukhatme, Vijay Kumar, Lifeng Zhou

机构 * GRASP Lab, University of Pennsylvania(宾夕法尼亚大学GRASP实验室) Department of Electrical and Computer Engineering, Drexel University(德雷塞尔大学电气与计算机工程系) Department of Computer Science, University of Southern California(南加州大学计算机科学系)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏