arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-06 至 2025-11-06 共收录 2 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 2 篇

2511.03271 2025-11-06 cs.CR cs.CL 79%

Let the Bees Find the Weak Spots: A Path Planning Perspective on Multi-Turn Jailbreak Attacks against LLMs

Yize Liu, Yunyun Hou, Aina Sui

专题命中 越狱攻击 :jailbreak(title);red teaming(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03434 2025-11-06 cs.HC cs.AI cs.MA cs.NI cs.SI 57%

Inter-Agent Trust Models: A Comparative Study of Brief, Claim, Proof, Stake, Reputation and Constraint in Agentic Web Protocol Design-A2A, AP2, ERC-8004, and Beyond

Botao 'Amber' Hu, Helena Rong

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI

Comments Submitted to AAAI 2026 Workshop on Trust and Control in Agentic AI (TrustAgent)

详情

展开后加载摘要…

URL PDF HTML 收藏