arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-13 至 2025-11-13 共收录 42 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 3 篇

2511.05927 2025-11-13 cs.CY cs.AI econ.GN q-fin.EC 62%

Artificial intelligence and the Gulf Cooperation Council workforce adapting to the future of work

Mohammad Rashed Albous, Melodena Stephens, Odeh Rashed Al-Jayyousi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Journal ref Humanit Soc Sci Commun 12, 1649 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08641 2025-11-13 cs.CR cs.CY cs.MA 57%

QOC DAO -- Stepwise Development Towards an AI Driven Decentralized Autonomous Organization

Marc Jansen, Christophe Verdot

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他安全 10 篇

2511.08842 2025-11-13 cs.AR cs.AI cs.CR 88%

3D Guard-Layer: An Integrated Agentic AI Safety System for Edge Artificial Intelligence

Eren Kurshan, Yuan Xie, Paul Franzon

专题命中 其他安全 :safety(title,abstract);AI safety(title,abstract);分类 cs.AI

Comments Resubmitting Re: Arxiv Committee Approval

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07419 2025-11-13 cs.LG 79%

Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs

Zhongyang Li, Ziyue Li, Tianyi Zhou

机构 * Johns Hopkins University(约翰霍普金斯大学) University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04834 2025-11-13 cs.LG cs.AI cs.CV 76%

Prompt-Based Safety Guidance Is Ineffective for Unlearned Text-to-Image Diffusion Models

Jiwoo Shin, Byeonghu Na, Mina Kang, Wonhyeok Choi, Il-Chul Moon

机构 * KAIST(韩国科学技术院)

专题命中 其他安全 :safety(title);分类 cs.AI、cs.LG

Comments Accepted at NeurIPS 2025 Workshop on Generative and Protective AI for Content Creation

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23749 2025-11-13 astro-ph.IM cs.LG 57%

Re-envisioning Euclid Galaxy Morphology: Identifying and Interpreting Features with Sparse Autoencoders

John F. Wu, Michael Walmsley

机构 * Space Telescope Science Institute(太空望远镜科学研究所) Johns Hopkins University(约翰霍普金斯大学) University of Toronto(多伦多大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Authors contributed equally to this work. Accepted to NeurIPS Machine Learning and the Physical Sciences Workshop. See trained model at https://huggingface.co/mwalmsley/euclid-rr2-mae, HuggingFace demo at https://huggingface.co/spaces/mwalmsley/euclid_masked_autoencoder, and code at https://github.com/jwuphysics/euclid-galaxy-morphology-saes

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08802 2025-11-13 cs.LG 57%

The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?

Denis Sutter, Julian Minder, Thomas Hofmann, Tiago Pimentel

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments NeurIPS 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08593 2025-11-13 cs.CL 57%

Knowledge Graph Analysis of Legal Understanding and Violations in LLMs

Abha Jha, Abel Salinas, Fred Morstatter

专题命中 其他安全 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09410 2025-11-13 cs.DC cs.DS cs.PF 50%

No Cords Attached: Coordination-Free Concurrent Lock-Free Queues

Yusuf Motiwala

专题命中 其他安全 :safety(abstract)

Comments 10 pages, 2 figures, 3 tables. Lock-free concurrent queue with coordination-free memory reclamation

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25054 2025-11-13 econ.GN q-fin.EC 50%

Signaling in the Age of AI: Evidence from Cover Letters

Jingyi Cui, Gabriel Dias, Justin Ye

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15503 2025-11-13 cs.CV 50%

Domain Adaptation from Generated Multi-Weather Images for Unsupervised Maritime Object Classification

Dan Song, Shumeng Huo, Wenhui Li, Lanjun Wang, Chao Xue, An-An Liu

机构 * The School of Electrical and Information Engineering, Tianjin University, China(天津大学电气与信息工程学院) Tiandy Technologies Co., Ltd, Tianjin, China(天津天盾科技有限公司)

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16602 2025-11-13 cs.CV cs.GR 50%

Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models

Ronghuan Wu, Wanchao Su, Jing Liao

机构 * City University of Hong Kong(香港城市大学) Monash University(墨尔本大学)

专题命中 其他安全 :alignment(abstract)

Comments Project Page: https://chat2svg.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏