arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-14 至 2025-11-14 共收录 45 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 16 篇

2510.03186 2025-11-14 cs.LG 79%

Superposition disentanglement of neural representations reveals hidden alignment

André Longon, David Klindt, Meenakshi Khosla

机构 * UC San Diego(圣迭戈大学) Cold Spring Harbor Laboratory(冷泉港实验室)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10334 2025-11-14 cs.CV 78%

Learning to Tell Apart: Weakly Supervised Video Anomaly Detection via Disentangled Semantic Alignment

Wenti Yin, Huaxin Zhang, Xiang Wang, Yuqing Lu, Yicheng Zhang, Bingquan Gong, Jialong Zuo, Li Yu, Changxin Gao, Nong Sang

专题命中 其他安全 :alignment(title,abstract)

Comments Accepted to AAAI 2026. Code is available at https://github.com/lessiYin/DSANet

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10211 2025-11-14 cs.CV 78%

HeatV2X: Scalable Heterogeneous Collaborative Perception via Efficient Alignment and Interaction

Yueran Zhao, Zhang Zhang, Chao Sun, Tianze Wang, Chao Yue, Nuoran Li

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 其他安全 :alignment(title,abstract)

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10627 2025-11-14 cs.AI cs.CV cs.FL cs.LG 62%

Querying Labeled Time Series Data with Scenario Programs

Edward Kim, Devan Shanker, Varun Bharadwaj, Hongbeen Park, Jinkyu Kim, Hazem Torfah, Daniel J Fremont, Sanjit A Seshia

机构 * University of California, Berkeley(加州大学伯克利分校) Korea University(韩国大学) Chalmers University of Technology(查尔姆斯理工大学) University of Gothenburg(哥德堡大学) University of California, Santa Cruz(加州大学圣克ruz分校)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Journal ref NASA Formal Methods Conference 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10583 2025-11-14 cs.CL cs.AI 62%

Evaluating Prompting Strategies with MedGemma for Medical Order Extraction

Abhinand Balachandran, Bavana Durgapraveen, Gowsikkan Sikkan Sudhagar, Vidhya Varshany J S, Sriram Rajkumar

机构 * EXL Health AI Lab at MEDIQA-OE 2025(EXL健康AI实验室)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments 2 figures 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09796 2025-11-14 cs.CL cs.AI 62%

Predicate-Argument Structure Divergences in Chinese and English Parallel Sentences and their Impact on Language Transfer

Rocco Tripodi, Xiaoyu Liu

机构 * Department of Environmental Sciences, Informatics and Statistics(环境科学、信息学与统计学系) Department of Linguistic Sciences And Foreign Literatures(语言科学与外国文学系)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10229 2025-11-14 cs.CL 57%

LangGPS: Language Separability Guided Data Pre-Selection for Joint Multilingual Instruction Tuning

Yangfan Ye, Xiaocheng Feng, Xiachong Feng, Lei Huang, Weitao Ma, Qichen Hong, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments AAAI2026 Main Track Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10020 2025-11-14 cs.CV cs.AI 57%

Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly Generation

Yuxin Jiang, Wei Luo, Hui Zhang, Qiyu Chen, Haiming Yao, Weiming Shen, Yunkang Cao

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09984 2025-11-14 cs.CL 57%

Language Drift in Multilingual Retrieval-Augmented Generation: Characterization and Decoding-Time Mitigation

Bo Li, Zhenghua Xu, Rui Xie

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments AAAI'26, Oral Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09973 2025-11-14 cs.CV cs.AI 57%

Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models

Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Taiga Yamane, Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08238 2025-11-14 cs.CV cs.AI 57%

Remodeling Semantic Relationships in Vision-Language Fine-Tuning

Xiangyang Wu, Liu Liu, Baosheng Yu, Jiayan Qiu, Zhenwei Shi

机构 * Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院) School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) Nanyang Technological University(南洋理工大学) University of Leicester(莱斯特大学) School of Astronautics, Beihang University(北京航空航天大学航天学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05909 2025-11-14 cs.CL 57%

Beyond Perplexity: Let the Reader Select Retrieval Summaries via Spectrum Projection Score

Zhanghao Hu, Qinglin Zhu, Siya Qi, Yulan He, Hanqi Yan, Lin Gui

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted by AAAI 2026 Oral. Project link: https://zhanghao-aaai2026-sps.github.io/AAAI2026-SPS/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22712 2025-11-14 cs.LG stat.ML 57%

Generalized Linear Mode Connectivity for Transformers

Alexander Theus, Alessandro Cabodi, Sotiris Anagnostidis, Antonio Orvieto, Sidak Pal Singh, Valentina Boeva

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Accepted to NeurIPS 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09844 2025-11-14 cs.LG cs.PF 57%

Steering Pretrained Drafters during Speculative Decoding

Frédéric Berdoz, Peer Rheinboldt, Roger Wattenhofer

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Accepted at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09709 2025-11-14 cs.CL 57%

Contextual morphologically-guided tokenization for Latin encoder models

Marisa Hudspeth, Patrick J. Burns, Brendan O'Connor

机构 * Manning College of Information & Computer Sciences, University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校信息与计算机科学学院) Institute for the Study of the Ancient World, New York University(纽约大学古代世界研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏