arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-17 至 2025-11-17 共收录 44 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 2 篇

2503.16851 2025-11-17 cs.CR cs.CL 70%

Interpretable LLM Guardrails via Sparse Representation Steering

Zeqing He, Zhibo Wang, Huiyu Xu, Hejun Lin, Wenhui Zhang, Zhixuan Chu

机构 * The State Key Laboratory of Blockchain and Data Security, Zhejiang University, China(区块链与数据安全国家重点实验室,浙江大学,中国) School of Cyber Science and Technology, Zhejiang University, China(网络安全与技术学院,浙江大学,中国) College of Computer and Information Sciences, Fujian Agriculture and Forestry University, China(计算机与信息科学学院,福建农林大学,中国)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22896 2025-11-17 physics.soc-ph 50%

Modelling vehicle and pedestrian collective dynamics: Challenges and advances

Antoine Tordeux, Cécile Appert-Rolland, Alexandre Nicolas, Armin Seyfried, Denis Ullmo

专题命中 AI治理与伦理 :safety(abstract)

Comments 20 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他安全 12 篇

2406.03442 2025-11-17 cs.CL cs.AI 79%

Are language models rational? The case of coherence norms and belief revision

Thomas Hofweber, Peter Hase, Elias Stengel-Eskin, Mohit Bansal

机构 * Department of Philosophy University of North Carolina at Chapel Hill(哲学系北卡罗来纳大学教堂山分校) Department of Computer Science University of North Carolina at Chapel Hill(计算机科学系北卡罗来纳大学教堂山分校) Department of Computer Science University of Texas at Austin(计算机科学系德克萨斯大学奥斯汀分校)

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments substantial expansions of sections 4 and 5, updated references, numerous smaller additions and clarifications

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07334 2025-11-17 cs.CV 78%

Unleashing the Potential of Large Language Models for Text-to-Image Generation through Autoregressive Representation Alignment

Xing Xie, Jiawei Liu, Ziyue Lin, Huijie Fan, Zhi Han, Yandong Tang, Liangqiong Qu

专题命中 其他安全 :alignment(title,abstract)

Comments Accepted by AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10850 2025-11-17 cs.CL cs.AI cs.LG 67%

Leveraging Parameter Space Symmetries for Reasoning Skill Transfer in LLMs

Stefan Horoi, Sangwoo Cho, Supriyo Chakraborty, Shi-Xiong Zhang, Sambit Sahu, Guy Wolf, Genta Indra Winata

机构 * Université de Montréal(蒙特利尔大学) Mila – Quebec AI Institute(魁北克人工智能研究所) Capital One(Capital One公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10846 2025-11-17 cs.CL cs.AI cs.CY 67%

Reinforcing Stereotypes of Anger: Emotion AI on African American Vernacular English

Rebecca Dorn, Christina Chance, Casandra Rusti, Charles Bickham, Kai-Wei Chang, Fred Morstatter, Kristina Lerman

机构 * University of Southern California, Information Science Institute(南加州大学信息科学研究所) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10628 2025-11-17 cs.CL cs.AI cs.LG 67%

Instella: Fully Open Language Models with Stellar Performance

Jiang Liu, Jialian Wu, Xiaodong Yu, Yusheng Su, Prakamya Mishra, Gowtham Ramesh, Sudhanshu Ranjan, Chaitanya Manem, Ximeng Sun, Ze Wang, Pratik Prabhanjan Brahma, Zicheng Liu, Emad Barsoum

机构 * AMD

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20856 2025-11-17 cs.LG cs.AI 62%

Strada-LLM: Graph LLM for traffic prediction

Seyed Mohamad Moghadas, Bruno Cornelis, Alexandre Alahi, Adrian Munteanu

机构 * Department of Electronics and Informatics, Vrije Universiteit Brussel(布鲁塞尔自由大学电子与信息系) VITA Lab, EPFL(EPFL VITA实验室)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10675 2025-11-17 cs.CL cs.AI cs.IR 62%

Learn to Select: Exploring Label Distribution Divergence for In-Context Demonstration Selection in Text Classification

Ye Jiang, Taihang Wang, Youzheng Liu, Yimin Wang, Yuhan Xia, Yunfei Long

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16657 2025-11-17 cs.CV cs.AI cs.CL 62%

DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation

Zun Wang, Jialu Li, Han Lin, Jaehong Yoon, Mohit Bansal

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments AAAI 2026, Project website: https://zunwang1.github.io/DreamRunner

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11526 2025-11-17 cs.CV 50%

Bridging Hidden States in Vision-Language Models

Benjamin Fein-Ashley, Jacob Fein-Ashley

机构 * University of Southern California(南加州大学)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11168 2025-11-17 cs.CV 50%

CATS-V2V: A Real-World Vehicle-to-Vehicle Cooperative Perception Dataset with Complex Adverse Traffic Scenarios

Hangyu Li, Bofeng Cao, Zhaohui Liang, Wuzhen Li, Juyoung Oh, Yuxuan Chen, Shixiao Liang, Hang Zhou, Chengyuan Ma, Jiaxi Liu, Zheng Li, Peng Zhang, KeKe Long, Maolin Liu, Jackson Jiang, Chunlei Yu, Shengxiang Liu, Hongkai Yu, Xiaopeng Li

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) wuwen-ai Cleveland State University(克利夫兰州立大学)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10948 2025-11-17 cs.CV cs.HC 50%

DEFT-LLM: Disentangled Expert Feature Tuning for Micro-Expression Recognition

Ren Zhang, Huilai Li, Chao qi, Guoliang Xu, Tianyu Zhou, Wei wei, Jianqin Yin

机构 * College of Intelligent Engineering and Automation, Beijing University of Posts and Telecommunications(智能工程与自动化学院,北京邮电大学)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10774 2025-11-17 cs.CV 50%

Frequency-Aware Vision-Language Multimodality Generalization Network for Remote Sensing Image Classification

Junjie Zhang, Feng Zhao, Hanqiang Liu, Jun Yu

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏