arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-22 至 2025-10-22 共收录 11 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 11 篇

2510.18154 2025-10-22 cs.AI cs.CY 88%

Annotating the Chain-of-Thought: A Behavior-Labeled Dataset for AI Safety

Antonio-Gabriel Chacón Menke, Phan Xuan Tan, Eiji Kamioka

专题命中 其他安全 :safety(title,abstract);AI safety(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15886 2025-10-22 cs.CY cs.AI 87%

Combining Cost-Constrained Runtime Monitors for AI Safety

Tim Tian Hua, James Baskerville, Henri Lemoine, Mia Hopman, Aryan Bhatt, Tyler Tracy

机构 * MARS Redwood Research

专题命中 其他安全 :safety(title,abstract);AI safety(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17833 2025-10-22 q-bio.NC cs.AI 79%

Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage

Ángela López-Cardona, Sebastián Idesis, Mireia Masias-Bruns, Sergi Abadal, Ioannis Arapakis

机构 * Universitat Politècnica de Catalunya(加泰罗尼亚理工大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14359 2025-10-22 cs.CV 78%

Dual Data Alignment Makes AI-Generated Image Detector Easier Generalizable

Ruoxin Chen, Junwei Xi, Zhiyuan Yan, Ke-Yue Zhang, Shuang Wu, Jingyi Xie, Xu Chen, Lei Xu, Isabel Guan, Taiping Yao, Shouhong Ding

机构 * Tencent YouTu Lab(腾讯优图实验室) East China University of Science and Technology(东华大学) Peking University(北京大学) Renmin University of China(中国人民大学) Shenzhen University(深圳大学) Hong Kong University of Science and Technology(香港科技大学)

专题命中 其他安全 :alignment(title,abstract)

Comments NeurIPS 2025 Spotlight. 13 Pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18502 2025-10-22 cs.CV cs.AI cs.CL cs.LG 67%

Zero-Shot Vehicle Model Recognition via Text-Based Retrieval-Augmented Generation

Wei-Chia Chang, Yan-Ann Chen

机构 * Yuan Ze University(元智大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by The 38th Conference of Open Innovations Association FRUCT, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01472 2025-10-22 cs.CL cs.AI 62%

FALCON: Fine-grained Activation Manipulation by Contrastive Orthogonal Unalignment for Large Language Model

Jinwei Hu, Zhenglin Huang, Xiangyu Yin, Wenjie Ruan, Guangliang Cheng, Yi Dong, Xiaowei Huang

机构 * School of Computer Science and Informatics, University of Liverpool, UK(计算机科学与信息学学院,利物浦大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted at NeurIPS 2025 with minor revisions

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18745 2025-10-22 cs.CL 61%

Topoformer: brain-like topographic organization in Transformer language models through spatial querying and reweighting

Taha Binhuraib, Greta Tuckute, Nicholas Blauch

机构 * Novus Technologies MIT(麻省理工学院) Harvard University(哈佛大学)

专题命中 其他安全 :alignment(abstract,comments);分类 cs.CL

Comments ICLR 2024 Workshop on Representational Alignment (Re-Align) Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18583 2025-10-22 cs.CV cs.LG 57%

CovMatch: Cross-Covariance Guided Multimodal Dataset Distillation with Trainable Text Encoder

Yongmin Lee, Hye Won Chung

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00363 2025-10-22 cs.IR cs.CL 57%

Adapting General-Purpose Embedding Models to Private Datasets Using Keyword-based Retrieval

Yubai Wei, Jiale Han, Yi Yang

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Link: https://github.com/BaileyWei/BMEmbed

Journal ref Findings of the Association for Computational Linguistics ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17909 2025-10-22 cs.CL 57%

Atomic Literary Styling: Mechanistic Manipulation of Prose Generation in Neural Language Models

Tsogt-Ochir Enkhbayar

机构 * Mongol AI(蒙古AI)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments 12 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18113 2025-10-22 cs.CR 50%

Investigating the Impact of Dark Patterns on LLM-Based Web Agents

Devin Ersoy, Brandon Lee, Ananth Shreekumar, Arjun Arunasalam, Muhammad Ibrahim, Antonio Bianchi, Z. Berkay Celik

专题命中 其他安全 :safety(abstract)

Comments At IEEE S&P 2026

详情

展开后加载摘要…

URL PDF HTML 收藏