arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-21 至 2025-10-21 共收录 8 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 8 篇

2510.17402 2025-10-21 cs.CL cs.AI cs.LG 75%

Leveraging Group Relative Policy Optimization to Advance Large Language Models in Traditional Chinese Medicine

Jiacheng Xie, Shuai Zeng, Yang Yu, Xiaoting Tang, Guanghui An, Dong Xu

专题命中 安全训练 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16144 2025-10-21 cs.NI cs.AI cs.MA 70%

Agentic AI for Ultra-Modern Networks: Multi-Agent Framework for RAN Autonomy and Assurance

Sukhdeep Singh, Avinash Bhat, Shweta M, Subhash K Singh, Moonki Hong, Madhan Raj K, Kandeepan Sithamparanathan, Sunder A. Khowaja, Kapal Dev

专题命中 安全训练 :safety(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17369 2025-10-21 cs.RO cs.AI cs.LG 62%

Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots

Haochen Su, Cristian Meo, Francesco Stella, Andrea Peirone, Kai Junge, Josie Hughes

机构 * EPFL(苏黎世联邦理工学院) LatentWorlds AI TUDelft(代尔夫特理工大学) Embodied AI SA(具身人工智能股份有限公司)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by NeurIPS 2025 SpaVLE workshop. 4 pages, 2 figures(in main paper, excluding references and supplements)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15896 2025-10-21 cs.HC cs.AI cs.CY 62%

From Coordination to Personalization: A Trust-Aware Simulation Framework for Emergency Department Decision Support

Zoi Lygizou, Dimitris Kalles

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16255 2025-10-21 cs.CR cs.AI 57%

Detecting Adversarial Fine-tuning with Auditing Agents

Sarah Egler, John Schulman, Nicholas Carlini

机构 * MATS & Anthropic Fellows Program(MATS与Anthropic Fellow项目) Thinking Machines Lab(Thinking Machines实验室) Anthropic

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15937 2025-10-21 q-fin.RM q-fin.TR 50%

Tail-Safe Stochastic-Control SPX-VIX Hedging: A White-Box Bridge Between AI Sensitivities and Arbitrage-Free Market Dynamics

Jian'an Zhang

专题命中 安全训练 :safety(abstract)

Comments 52 pages; 3 figures; PRIMEarxiv template; fully reproducible artifact (code, configs, plots)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00682 2025-10-21 cs.RO 50%

Immersive Explainability: Visualizing Robot Navigation Decisions through XAI Semantic Scene Projections in Virtual Reality

Jorge de Heuvel, Sebastian Müller, Marlene Wessels, Aftab Akhtar, Christian Bauckhage, Maren Bennewitz

机构 * University of Bonn(波恩大学) Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔人工智能与机器学习研究所) Center for Robotics(机器人中心) University of Mainz(美因茨大学) Fraunhofer Institute for Intelligent Analysis and Information Systems IAIS(弗劳恩霍夫智能分析与信息系统研究所)

专题命中 安全训练 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12861 2025-10-21 cs.RO 50%

Safe Multi-Agent Reinforcement Learning for Behavior-Based Cooperative Navigation

Murad Dawood, Sicong Pan, Nils Dengler, Siqi Zhou, Angela P. Schoellig, Maren Bennewitz

机构 * Humanoid Robots Lab, University of Bonn(波恩大学人形机器人实验室)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏