arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-04 至 2025-11-04 共收录 66 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 16 篇

2506.10707 2025-11-04 cs.LG cs.AI 62%

ConTextTab: A Semantics-Aware Tabular In-Context Learner

Marco Spinaci, Marek Polewczyk, Maximilian Schambach, Sam Thelin

机构 * SAP France(SAP法国分公司) SAP SE(SAP德国分公司)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted as spotlight at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14622 2025-11-04 cs.CR cs.AI cs.LG 62%

Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection

Yihao Guo, Haocheng Bian, Liutong Zhou, Ze Wang, Zhaoyi Zhang, Francois Kawala, Milan Dean, Ian Fischer, Yuantao Peng, Noyan Tokgozoglu, Ivan Barrientos, Riyaaz Shaik, Rachel Li, Chandru Venkataraman, Reza Shifteh Far, Moses Pawar, Venkat Sundaranatha, Michael Xu, Frank Chu

机构 * Apple(苹果公司) Cohere(Cohere公司) DeepMind(深度思维公司) Meta MongoDB(MongoDB公司)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01226 2025-11-04 cs.LG 57%

WindMiL: Equivariant Graph Learning for Wind Loading Prediction

Themistoklis Vargiemezis, Charilaos Kanatsoulis, Catherine Gorlé

机构 * Department of Civil & Environmental Engineering, Stanford, CA, USA(土木与环境工程系,斯坦福大学) Department of Computer Science, Stanford, CA, USA(计算机科学系,斯坦福大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00122 2025-11-04 cs.AI 57%

Engineering.ai: A Platform for Teams of AI Engineers in Computational Design

Ran Xu, Yupeng Qi, Jingsen Feng, Xu Chu

机构 * Faculty for Aerospace Engineering and Geodesy, University of Stuttgart, Stuttgart, Germany(航空航天工程与大地测量学系,斯图加特大学) Cluster of Excellence SimTech, University of Stuttgart, Stuttgart, Germany(卓越中心SimTech,斯图加特大学) Faculty of Environment, Science and Economy, University of Exeter, Exeter EX4 4QF, United Kingdom(环境、科学与经济学院,埃克塞特大学) University of Stuttgart, Stuttgart, Germany(斯图加特大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17267 2025-11-04 cs.CL 57%

GreekBarBench: A Challenging Benchmark for Free-Text Legal Reasoning and Citations

Odysseas S. Chlapanis, Dimitrios Galanis, Nikolaos Aletras, Ion Androutsopoulos

机构 * Department of Informatics, Athens University of Economics and Business(信息学院,雅典经济与商业大学) Archimedes, Athena Research Center(阿提卡研究中心-阿基米德) Athena Research Center(阿提卡研究中心) University of Sheffield(谢菲尔德大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 19 pages, 17 figures, accepted in EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01341 2025-11-04 cs.CL 57%

AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding

Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang, Suyuchen Wang, Chao Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu, Nicolas Chapados, Yoshua Bengio, Enamul Hoque, Christopher Pal, Issam H. Laradji, David Vazquez, Perouz Taslakian, Spandana Gella, Sai Rajeswar

机构 * ServiceNow York University(约克大学) Mila – Quebec AI Institute(魁北克人工智能研究院) École de Technologie Supérieure(魁北克高等技术学院) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) University of Waterloo(滑铁卢大学) CIFAR AI Chair(CIFAR人工智能 chair) Polytechnique Montréal(蒙特利尔理工学院) University of British Columbia(不列颠哥伦比亚大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01757 2025-11-04 cs.SE 50%

Towards LLM-Powered Task-Aware Retrieval of Scientific Workflows for Galaxy

Shamse Tasnim Cynthia, Banani Roy

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05887 2025-11-04 eess.AS 50%

Aligning Speech to Languages to Enhance Code-switching Speech Recognition

Hexin Liu, Xiangyu Zhang, Haoyang Zhang, Leibny Paola Garcia, Andy W. H. Khong, Eng Siong Chng, Shinji Watanabe

专题命中 安全评测 :alignment(abstract)

Comments Accepted to IEEE Trans. Audio Speech Lang. Process., copyright has been transferred to IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00659 2025-11-04 eess.SY cs.SY 50%

Unveiling Uniform Shifted Power Law in Stochastic Human and Autonomous Driving Behavior

Wang Chen, Heye Huang, Ke Ma, Hangyu Li, Shixiao Liang, Hang Zhou, Xiaopeng Li

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00613 2025-11-04 cs.CV 50%

CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World

Yating Yu, Congqi Cao, Zhaoying Wang, Weihua Meng, Jie Li, Yuxin Li, Zihao Wei, Zhongpei Shen, Jiajun Zhang

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14918 2025-11-04 cs.CV 50%

Semantic-Aware Representation Learning via Conditional Transport for Multi-Label Image Classification

Ren-Dong Xie, Zhi-Fen He, Bo Li, Bin Liu, Jin-Yan Hu

机构 * School of Mathematics and Information Science(数学与信息科学学院) Key Laboratory of Jiangxi Province for Image Processing and Pattern Recognition(江西省图像处理与模式识别重点实验室)

专题命中 安全评测 :alignment(abstract)

Comments The paper is under consideration at Pattern Recognition Letters

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 7 篇

2511.00379 2025-11-04 cs.AI cs.CL 81%

Diverse Human Value Alignment for Large Language Models via Ethical Reasoning

Jiahao Wang, Songkai Xue, Jinghui Li, Xiaozhen Wang

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by AIES 2025, camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09606 2025-11-04 cs.CY 70%

Local US officials' views on the impacts and governance of AI: Evidence from 2022 and 2023 survey waves

Sophia Hatz, Noemi Dreksler, Kevin Wei, Baobao Zhang

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

Journal ref PLoS One 20(10): e0332919, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00024 2025-11-04 cs.CY cs.AI cs.CL cs.LG stat.AP 70%

Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model

Haotian Hang, Yueyang Shen, Vicky Zhu, Jose Cruz, Michelle Li

机构 * University of Southern California(南加州大学) University of Michigan(密歇根大学) Babson College(巴布森学院) University of Connecticut(康涅狄格大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20432 2025-11-04 cs.AI cs.CY cs.GT cs.LG 67%

LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory

Jingru Jia, Zehua Yuan, Junhao Pan, Paul E. McNamara, Deming Chen

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01550 2025-11-04 cs.AI 57%

Analyzing Sustainability Messaging in Large-Scale Corporate Social Media

Ujjwal Sharma, Stevan Rudinac, Ana Mićković, Willemijn van Dolen, Marcel Worring

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14886 2025-11-04 cs.CV 50%

Surgical Scene Understanding in the Era of Foundation AI Models: A Comprehensive Review

Ufaq Khan, Umair Nawaz, Adnan Qayyum, Shazad Ashraf, Yutong Xie, Muhammad Haris Khan, Muhammad Bilal, Junaid Qadir

机构 * MBZ University of AI(MBZ人工智能大学) Hamad Bin Khalifa University(哈马德·本·卡西姆大学) Birmingham City University(伯明翰城市大学) Qatar University(卡塔尔大学) University Hospitals Birmingham and University of Birmingham(伯明翰大学医院和伯明翰大学)

专题命中 AI治理与伦理 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00026 2025-11-04 cs.RO 50%

Gen AI in Automotive: Applications, Challenges, and Opportunities with a Case study on In-Vehicle Experience

Chaitanya Shinde, Divya Garikapati

专题命中 AI治理与伦理 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 18 篇

2510.08872 2025-11-04 cs.AI cs.GT cs.HC cs.LG cs.MA 81%

GTAlign: Game-Theoretic Alignment of LLM Assistants for Social Welfare

Siqi Zhu, David Zhang, Pedro Cisneros-Velarde, Jiaxuan You

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) VMware Research(VMware研究)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 31 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19739 2025-11-04 cs.CV 78%

FUSE: Label-Free Image-Event Joint Monocular Depth Estimation via Frequency-Decoupled Alignment and Degradation-Robust Fusion

Pihai Sun, Junjun Jiang, Yuanqi Yao, Youyu Chen, Wenbo Zhao, Kui Jiang, Xianming Liu

机构 * Faculty of Computing, Harbin Institute of Technology(计算机学院,哈尔滨工业大学) Zhengzhou Research Institute, Harbin Institute of Technology(郑州研究院,哈尔滨工业大学)

专题命中 其他安全 :alignment(title,abstract)

Comments [IROS 2025, camera ready version]: 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06164 2025-11-04 cs.LG cs.AI 76%

Model Alignment Search

Satchel Grant

机构 * Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14500 2025-11-04 cs.AI cs.MA cs.NE 70%

The Digital Ecosystem of Beliefs: does evolution favour AI over humans?

David M. Bossens, Shanshan Feng, Yew-Soon Ong

机构 * Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR) Centre for Frontier AI Research (CFAR), Agency for Science, Technology and Research (A*STAR)(高性能计算研究所(IHPC)、科技研究局(A*STAR)前沿人工智能研究中心(CFAR)、科技研究局(A*STAR)) School of Computer Science Wuhan University(计算机科学学院 武汉大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14846 2025-11-04 cs.AI cs.CL cs.LO 62%

Where to Search: Measure the Prior-Structured Search Space of LLM Agents

Zhuo-Yang Song

机构 * School of Physics, Peking University(物理学院,北京大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments 11 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02184 2025-11-04 stat.ML cs.AI cs.CV cs.LG math.ST stat.TH 62%

Double Descent Meets Out-of-Distribution Detection: Theoretical Insights and Empirical Analysis on the role of model complexity

Mouïn Ben Ammar, David Brellmann, Arturo Mendoza, Antoine Manzanera, Gianni Franchi

机构 * U2IS Lab ENSTA Paris(ENSTA巴黎大学U2IS实验室) Safran Tech(萨弗兰技术)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted at NeurIPS 2025 (Conference on Neural Information Processing Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01258 2025-11-04 cs.CL cs.AI 62%

Measuring Algorithmic Partisanship via Zero-Shot Classification and Its Implications on Political Discourse

Nathan Junzi Chen

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 19 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00002 2025-11-04 cs.LG cs.AI cs.CV 62%

VRScout: Towards Real-Time, Autonomous Testing of Virtual Reality Games

Yurun Wu, Yousong Sun, Burkhard Wunsche, Jia Wang, Elliott Wen

机构 * School of Computer Science University of Auckland(计算机科学学院 奥克兰大学) School of Advanced Technology Xi'an Jiaotong-Liverpool University(先进科技学院 西交利物浦大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01380 2025-11-04 cs.CL 57%

Confounding Factors in Relating Model Performance to Morphology

Wessel Poelman, Thomas Bauwens, Miryam de Lhoneux

机构 * NLP, Department of Computer Science, KU Leuven(自然语言处理,计算机科学系,鲁文大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments EMNLP 2025: Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01289 2025-11-04 cs.CL 57%

FirstAidQA: A Synthetic Dataset for First Aid and Emergency Response in Low-Connectivity Settings

Saiyma Sittul Muna, Rezwan Islam Salvi, Mushfiqur Rahman Mushfique, Ajwad Abrar

机构 * Islamic University of Technology(伊斯兰技术大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL

Comments Accepted at the 5th Muslims in Machine Learning (MusIML) Workshop, co-located with NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15049 2025-11-04 cs.RO cs.AI cs.MA 57%

HAD-Gen: Human-like and Diverse Driving Behavior Modeling for Controllable Scenario Generation

Cheng Wang, Lingxin Kong, Massimiliano Tamborski, Stefano V. Albrecht

机构 * School of Engineering and Physical Sciences, Heriot-Watt University(赫瑞斯泰德大学工程与物理科学学院) School of Automation and Software Engineering, Shanxi University(山西大学自动化与软件工程学院) School of Informatics, University of Edinburgh(爱丁堡大学信息学院)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07537 2025-11-04 cs.LG 57%

Accident Impact Prediction based on a deep convolutional and recurrent neural network model

Pouyan Sajadi, Mahya Qorbani, Sobhan Moosavi, Erfan Hassannayebi

机构 * Department of Industrial Engineering, Sharif University of Technology(谢里夫理工大学工业工程系) School of Industrial and System Engineering, Georgia Institute of Technology(佐治亚理工学院工业与系统工程学院) Department of Computer Science and Engineering, Ohio State University(俄亥俄州立大学计算机科学与工程系)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments 28 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏