arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-17 至 2025-11-17 共收录 12 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 12 篇

2406.03442 2025-11-17 cs.CL cs.AI 79%

Are language models rational? The case of coherence norms and belief revision

Thomas Hofweber, Peter Hase, Elias Stengel-Eskin, Mohit Bansal

机构 * Department of Philosophy University of North Carolina at Chapel Hill(哲学系北卡罗来纳大学教堂山分校) Department of Computer Science University of North Carolina at Chapel Hill(计算机科学系北卡罗来纳大学教堂山分校) Department of Computer Science University of Texas at Austin(计算机科学系德克萨斯大学奥斯汀分校)

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments substantial expansions of sections 4 and 5, updated references, numerous smaller additions and clarifications

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07334 2025-11-17 cs.CV 78%

Unleashing the Potential of Large Language Models for Text-to-Image Generation through Autoregressive Representation Alignment

Xing Xie, Jiawei Liu, Ziyue Lin, Huijie Fan, Zhi Han, Yandong Tang, Liangqiong Qu

专题命中 其他安全 :alignment(title,abstract)

Comments Accepted by AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10850 2025-11-17 cs.CL cs.AI cs.LG 67%

Leveraging Parameter Space Symmetries for Reasoning Skill Transfer in LLMs

Stefan Horoi, Sangwoo Cho, Supriyo Chakraborty, Shi-Xiong Zhang, Sambit Sahu, Guy Wolf, Genta Indra Winata

机构 * Université de Montréal(蒙特利尔大学) Mila – Quebec AI Institute(魁北克人工智能研究所) Capital One(Capital One公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10846 2025-11-17 cs.CL cs.AI cs.CY 67%

Reinforcing Stereotypes of Anger: Emotion AI on African American Vernacular English

Rebecca Dorn, Christina Chance, Casandra Rusti, Charles Bickham, Kai-Wei Chang, Fred Morstatter, Kristina Lerman

机构 * University of Southern California, Information Science Institute(南加州大学信息科学研究所) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10628 2025-11-17 cs.CL cs.AI cs.LG 67%

Instella: Fully Open Language Models with Stellar Performance

Jiang Liu, Jialian Wu, Xiaodong Yu, Yusheng Su, Prakamya Mishra, Gowtham Ramesh, Sudhanshu Ranjan, Chaitanya Manem, Ximeng Sun, Ze Wang, Pratik Prabhanjan Brahma, Zicheng Liu, Emad Barsoum

机构 * AMD

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20856 2025-11-17 cs.LG cs.AI 62%

Strada-LLM: Graph LLM for traffic prediction

Seyed Mohamad Moghadas, Bruno Cornelis, Alexandre Alahi, Adrian Munteanu

机构 * Department of Electronics and Informatics, Vrije Universiteit Brussel(布鲁塞尔自由大学电子与信息系) VITA Lab, EPFL(EPFL VITA实验室)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10675 2025-11-17 cs.CL cs.AI cs.IR 62%

Learn to Select: Exploring Label Distribution Divergence for In-Context Demonstration Selection in Text Classification

Ye Jiang, Taihang Wang, Youzheng Liu, Yimin Wang, Yuhan Xia, Yunfei Long

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16657 2025-11-17 cs.CV cs.AI cs.CL 62%

DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation

Zun Wang, Jialu Li, Han Lin, Jaehong Yoon, Mohit Bansal

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments AAAI 2026, Project website: https://zunwang1.github.io/DreamRunner

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11526 2025-11-17 cs.CV 50%

Bridging Hidden States in Vision-Language Models

Benjamin Fein-Ashley, Jacob Fein-Ashley

机构 * University of Southern California(南加州大学)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11168 2025-11-17 cs.CV 50%

CATS-V2V: A Real-World Vehicle-to-Vehicle Cooperative Perception Dataset with Complex Adverse Traffic Scenarios

Hangyu Li, Bofeng Cao, Zhaohui Liang, Wuzhen Li, Juyoung Oh, Yuxuan Chen, Shixiao Liang, Hang Zhou, Chengyuan Ma, Jiaxi Liu, Zheng Li, Peng Zhang, KeKe Long, Maolin Liu, Jackson Jiang, Chunlei Yu, Shengxiang Liu, Hongkai Yu, Xiaopeng Li

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) wuwen-ai Cleveland State University(克利夫兰州立大学)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10948 2025-11-17 cs.CV cs.HC 50%

DEFT-LLM: Disentangled Expert Feature Tuning for Micro-Expression Recognition

Ren Zhang, Huilai Li, Chao qi, Guoliang Xu, Tianyu Zhou, Wei wei, Jianqin Yin

机构 * College of Intelligent Engineering and Automation, Beijing University of Posts and Telecommunications(智能工程与自动化学院,北京邮电大学)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10774 2025-11-17 cs.CV 50%

Frequency-Aware Vision-Language Multimodality Generalization Network for Remote Sensing Image Classification

Junjie Zhang, Feng Zhao, Hanqiang Liu, Jun Yu

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏