arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2511.05000 2025-11-10 cs.IR cs.AI 57%

Query Generation Pipeline with Enhanced Answerability Assessment for Financial Information Retrieval

Hyunkyu Kim, Yeeun Yoo, Youngjun Kwak

机构 * Kakaobank Seongnam-si Republic of Korea(韩国首尔市韩国 Kakao银行)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted(Oral) by ICAIF 2025. Hyunkyu Kim and Yeeun Yoo contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04690 2025-11-10 cs.MM cs.CL 57%

Automatización de Informes Geotécnicos para Macizos Rocosos con IA

Christofer Valencia, Alexis Llumigusín, Silvia Alvarez, Abrahan Arias, Christian Mejia-Escobar

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments 17 pages, in Spanish language

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21069 2025-11-10 cs.CV cs.AI cs.HC 57%

GAITEX: Human motion dataset of impaired gait and rehabilitation exercises using inertial and optical sensors

Andreas Spilz, Heiko Oppel, Jochen Werner, Kathrin Stucke-Straub, Felix Capanni, Michael Munz

机构 * AI for Sensor Data Analytics Research Group(人工智能传感器数据解析研究组) Ulm University of Applied Sciences(乌尔姆应用科学大学) Biomechatronic Research Group(生物机械研究组) Institute of Computer Science(计算机科学研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11137 2025-11-10 cs.CL 57%

Scalable Medication Extraction and Discontinuation Identification from Electronic Health Records Using Large Language Models

Chong Shao, Douglas Snyder, Chiran Li, Bowen Gu, Kerry Ngan, Chun-Ting Yang, Jiageng Wu, Richard Wyss, Kueiyu Joshua Lin, Jie Yang

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17086 2025-11-10 cs.CL 57%

Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM Reviews

Hyungyu Shin, Jingyu Tang, Yoonjoo Lee, Nayoung Kim, Hyunseung Lim, Ji Yong Cho, Hwajung Hong, Moontae Lee, Juho Kim

机构 * KAIST(韩国科学技术院) Huazhong University of Science and Technology(华中科技大学) LG AI Research(LG人工智能研究) University of Illinois Chicago(伊利诺伊大学香槟分校)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments EMNLP 2025 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04538 2025-11-07 cs.CL 57%

From Model to Breach: Towards Actionable LLM-Generated Vulnerabilities Reporting

Cyril Vallez, Alexander Sternfeld, Andrei Kucharavy, Ljiljana Dolamic

机构 * IEM, HES-SO Valais-Wallis(IEM,HES-SO瓦莱-达沃斯) II, HES-SO Valais-Wallis(II,HES-SO瓦莱-达沃斯) Cyber-Defence Campus, armasuisse(网络安全防御校区,armasuisse)

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06991 2025-11-07 cs.AI cs.GT cs.HC 57%

Evaluating LLM-Contaminated Crowdsourcing Data Without Ground Truth

Yichi Zhang, Jinlong Pang, Zhaowei Zhu, Yang Liu

机构 * DIMACS, Rutgers University(Rutgers大学DIMACS研究中心)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 32 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17793 2025-11-07 cs.CL 57%

Compression Hacking: A Supplementary Perspective on Informatics Properties of Language Models from Geometric Distortion

Jianxiang Zang, Meiling Ning, Yongda Wei, Shihan Dou, Jiazheng Zhang, Nijia Mo, Binhong Li, Tao Gui, Qi Zhang, Xuanjing Huang

机构 * Computation and Artificial Intelligence Innovative College, Fudan University(复旦大学计算与人工智能创新学院) Beijing University of Posts and Telecommunications(北京邮电大学) Shanghai University of International Business and Economics(上海国际商务经济学院) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05712 2025-11-07 cs.LG cs.CV q-bio.NC 57%

Scaling Laws for Task-Optimized Models of the Primate Visual Ventral Stream

Abdulkadir Gokce, Martin Schrimpf

机构 * EPFL(苏黎世联邦理工学院)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Published at ICML25 as a spotlight paper - 9 pages for the main paper, 22 pages in total. 7 main figures and 7 supplementary figures. Code, model weights, and benchmark results can be accessed at https://github.com/epflneuroailab/scaling-primate-vvs

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13406 2025-11-07 cs.AI cs.CE cs.MA 57%

Collaboration Dynamics and Reliability Challenges of Multi-Agent LLM Systems in Finite Element Analysis

Chuan Tian, Yilei Zhang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03051 2025-11-06 cs.AI cs.IR 57%

No-Human in the Loop: Agentic Evaluation at Scale for Recommendation

Tao Zhang, Kehui Yao, Luyi Ma, Jiao Chen, Reza Yousefi Maragheh, Kai Zhao, Jianpeng Xu, Evren Korpeoglu, Sushant Kumar, Kannan Achan

机构 * Walmart Global Tech(沃尔玛全球技术)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 4 page, NeurIPS 2025 Workshop: Evaluating the Evolving LLM Lifecycle

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20462 2025-11-06 cs.AI 57%

TAMO: Fine-Grained Root Cause Analysis via Tool-Assisted LLM Agent with Multi-Modality Observation Data in Cloud-Native Systems

Xiao Zhang, Qi Wang, Mingyi Li, Yuan Yuan, Mengbai Xiao, Fuzhen Zhuang, Dongxiao Yu

机构 * School of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院) Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02495 2025-11-05 cs.CV cs.CL 57%

DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding

Zixuan Liu, Siavash H. Khajavi, Guangkai Jiang

机构 * Department of Computer Science(计算机科学系) Tulane University(Tulane 大学) Department of Industrial Engineering and Management(工业工程与管理系) Aalto University(Aalto 大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments Advances in Neural Information Processing Systems 2025 (NeurIPS 2025), Poster, https://neurips.cc/virtual/2025/loc/san-diego/poster/121400

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01226 2025-11-04 cs.LG 57%

WindMiL: Equivariant Graph Learning for Wind Loading Prediction

Themistoklis Vargiemezis, Charilaos Kanatsoulis, Catherine Gorlé

机构 * Department of Civil & Environmental Engineering, Stanford, CA, USA(土木与环境工程系,斯坦福大学) Department of Computer Science, Stanford, CA, USA(计算机科学系,斯坦福大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00122 2025-11-04 cs.AI 57%

Engineering.ai: A Platform for Teams of AI Engineers in Computational Design

Ran Xu, Yupeng Qi, Jingsen Feng, Xu Chu

机构 * Faculty for Aerospace Engineering and Geodesy, University of Stuttgart, Stuttgart, Germany(航空航天工程与大地测量学系,斯图加特大学) Cluster of Excellence SimTech, University of Stuttgart, Stuttgart, Germany(卓越中心SimTech,斯图加特大学) Faculty of Environment, Science and Economy, University of Exeter, Exeter EX4 4QF, United Kingdom(环境、科学与经济学院,埃克塞特大学) University of Stuttgart, Stuttgart, Germany(斯图加特大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17267 2025-11-04 cs.CL 57%

GreekBarBench: A Challenging Benchmark for Free-Text Legal Reasoning and Citations

Odysseas S. Chlapanis, Dimitrios Galanis, Nikolaos Aletras, Ion Androutsopoulos

机构 * Department of Informatics, Athens University of Economics and Business(信息学院,雅典经济与商业大学) Archimedes, Athena Research Center(阿提卡研究中心-阿基米德) Athena Research Center(阿提卡研究中心) University of Sheffield(谢菲尔德大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 19 pages, 17 figures, accepted in EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01341 2025-11-04 cs.CL 57%

AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding

Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang, Suyuchen Wang, Chao Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu, Nicolas Chapados, Yoshua Bengio, Enamul Hoque, Christopher Pal, Issam H. Laradji, David Vazquez, Perouz Taslakian, Spandana Gella, Sai Rajeswar

机构 * ServiceNow York University(约克大学) Mila – Quebec AI Institute(魁北克人工智能研究院) École de Technologie Supérieure(魁北克高等技术学院) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) University of Waterloo(滑铁卢大学) CIFAR AI Chair(CIFAR人工智能 chair) Polytechnique Montréal(蒙特利尔理工学院) University of British Columbia(不列颠哥伦比亚大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27521 2025-11-03 cs.HC cs.CY 57%

Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk

Nick Judd, Alexandre Vaz, Kevin Paeth, Layla Inés Davis, Milena Esherick, Jason Brand, Inês Amaro, Tony Rousmaniere

专题命中 安全评测 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27244 2025-11-03 cs.SE cs.AI 57%

Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes

Ora Nova Fandina, Gal Amram, Eitan Farchi, Shmulik Froimovich, Raviv Gal, Wesam Ibraheem, Rami Katan, Alice Podolsky, Orna Raz

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27065 2025-11-03 cs.LG cs.PF 57%

MLPerf Automotive

Radoyeh Shojaei, Predrag Djurdjevic, Mostafa El-Khamy, James Goel, Kasper Mecklenburg, John Owens, Pınar Muyan-Özçelik, Tom St. John, Jinho Suh, Arjun Suresh

机构 * University of California, Davis(加州大学戴维斯分校) Arm(ARM公司) Samsung(三星) Qualcomm(高通) California State University, Sacramento(加州州立大学萨克拉门托分校) Gilmet Labs(Gilmet实验室) NVIDIA(英伟达) AMD(超微半导体)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 16 pages, 5 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26830 2025-11-03 cs.LG cs.CR 57%

SmoothGuard: Defending Multimodal Large Language Models with Noise Perturbation and Clustering Aggregation

Guangzhi Su, Shuchang Huang, Yutong Ke, Zhuohang Liu, Long Qian, Kaizhu Huang

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25761 2025-11-03 cs.CL 57%

DiagramEval: Evaluating LLM-Generated Diagrams via Graphs

Chumeng Liang, Jiaxuan You

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26052 2025-10-31 cs.CV cs.AI 57%

Dynamic VLM-Guided Negative Prompting for Diffusion Models

Hoyeon Chang, Seungjin Kim, Yoonseok Choi

机构 * KAIST(韩国科学技术院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: The First Workshop on Generative and Protective AI for Content Creation

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05747 2025-10-31 cs.CL 57%

SEA-LION: Southeast Asian Languages in One Network

Raymond Ng, Thanh Ngan Nguyen, Yuli Huang, Ngee Chia Tai, Wai Yi Leong, Wei Qi Leong, Xianbin Yong, Jian Gang Ngui, Yosephine Susanto, Nicholas Cheng, Hamsawardhini Rengarajan, Peerat Limkonchotiwat, Adithya Venkatadri Hulagadri, Kok Wai Teng, Yeo Yeow Tong, Bryan Siow, Wei Yi Teo, Wayne Lau, Choon Meng Tan, Brandon Ong, Zhi Hao Ong, Jann Railey Montalan, Adwin Chan, Sajeban Antonyrex, Ren Lee, Esther Choa, David Ong Tat-Wee, Bing Jie Darius Liu, William Chandra Tjhi, Erik Cambria, Leslie Teo

机构 * AI Singapore National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted at IJCNLP-AACL 2025 (Main Track). We released our model at https://huggingface.co/collections/aisingapore/sea-lionv3-672589a39cdadd6a5b199581

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04852 2025-10-31 cs.CV cs.LG 57%

CAUSAL3D: A Comprehensive Benchmark for Causal Learning from Visual Data

Disheng Liu, Yiran Qiao, Wuche Liu, Yiren Lu, Yunlai Zhou, Tuo Liang, Yu Yin, Jing Ma

机构 * Case Western Reserve University(凯斯西储大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Datasets link: https://huggingface.co/datasets/LLDDSS/Causal3D_Dataset

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25571 2025-10-30 cs.LG cs.DS cs.NA math.NA math.SP math.ST stat.TH 57%

Perturbation Bounds for Low-Rank Inverse Approximations under Noise

Phuc Tran, Nisheeth K. Vishnoi

机构 * Yale University(耶鲁大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25413 2025-10-30 cs.CL 57%

Seeing, Signing, and Saying: A Vision-Language Model-Assisted Pipeline for Sign Language Data Acquisition and Curation from Social Media

Shakib Yazdani, Yasser Hamidullah, Cristina España-Bonet, Josef van Genabith

机构 * German Research Center for Artificial Intelligence (DFKI GmbH)(德国人工智能研究中心(DFKI GmbH)) Saarland Informatics Campus(萨尔兰州信息技术校区) Barcelona Supercomputing Center (BSC-CNS)(巴塞罗那超级计算中心(BSC-CNS))

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted by RANLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25091 2025-10-30 cs.AI 57%

H3M-SSMoEs: Hypergraph-based Multimodal Learning with LLM Reasoning and Style-Structured Mixture of Experts

Peilin Tan, Liang Xie, Churan Zhi, Dian Tu, Chuanqi Shi

机构 * University of California, San Diego(加州大学圣迭戈分校) Wuhan University of Technology(武汉科技大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24951 2025-10-30 cs.LG cs.AR cs.NE 57%

Resource-Efficient and Robust Inference of Deep and Bayesian Neural Networks on Embedded and Analog Computing Platforms

Bernhard Klein

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Ph.D. dissertation, Heidelberg University, October 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24749 2025-10-30 cs.SE cs.AI 57%

Beyond Function-Level Search: Repository-Aware Dual-Encoder Code Retrieval with Adversarial Verification

Aofan Liu, Shiyuan Song, Haoxuan Li, Cehao Yang, Yiyan Qi

机构 * International Digital Economy Academy (IDEA)(国际数字经济学院) School of Electronic and Computer Engineering, Peking University(电子与计算机工程学院,北京大学) Shenzhen International Graduate School, Tsinghua University(深圳国际研究生院,清华大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏