arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-31 至 2025-10-31 共收录 13 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 13 篇

2510.26024 2025-10-31 cs.CL cs.AI 81%

Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs

HyoJung Han, Sweta Agrawal, Eleftheria Briakou

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08525 2025-10-31 cs.LG cs.AI 76%

A mathematical certification for positivity conditions in Neural Networks with applications to partial monotonicity and Trustworthy AI

Alejandro Polo-Molina, David Alfaya, Jose Portela

机构 * CDTI

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments 16 pages, 4 figures

Journal ref IEEE Transactions on Neural Networks and Learning Systems, Early Access, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14681 2025-10-31 cs.CL 74%

Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Quality

Yuto Harada, Yusuke Yamauchi, Yusuke Oda, Yohei Oseki, Yusuke Miyao, Yu Takagi

机构 * NII LLMC(日本信息处理学会大语言模型中心) The University of Tokyo(东京大学) NAIST(日本科学技术大学) Nagoya Institute of Technology(名古屋技术大学)

专题命中 安全评测 :alignment(title);分类 cs.CL

Comments Accepted to EMNLP 2025 (Main Conference). Models and evaluation results available at: https://github.com/llm-jp/massive-sft

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25908 2025-10-31 cs.AI 70%

SciTrust 2.0: A Comprehensive Framework for Evaluating Trustworthiness of Large Language Models in Scientific Applications

Emily Herron, Junqi Yin, Feiyi Wang

机构 * Oak Ridge National Laboratory(橡树岭国家实验室)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

Comments Preprint Submitted to ACM Transactions on AI for Science (TAIS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02927 2025-10-31 eess.SY cs.AI cs.LG cs.SY 62%

Multivariate Physics-Informed Convolutional Autoencoder for Anomaly Detection in Power Distribution Systems with High Penetration of DERs

Mehdi Jabbari Zideh, Sarika Khushalani Solanki

机构 * Lane Department of Computer Science and Electrical Engineering, West Virginia University(计算机科学与电气工程系,西弗吉尼亚大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Journal ref Sustainable Energy, Grids and Networks, Vol. 44, December 2025, 102022

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26402 2025-10-31 cs.AI cs.LG 62%

Autograder+: A Multi-Faceted AI Framework for Rich Pedagogical Feedback in Programming Education

Vikrant Sahu, Gagan Raj Gupta, Raghav Borikar, Nitin Mane

机构 * Indian Institute of Technology(印度理工学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26037 2025-10-31 cs.CR cs.AI cs.CL 62%

SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning

Kaiwen Zhou, Ahmed Elgohary, A S M Iftekhar, Amin Saied

机构 * Microsoft Responsible AI Research(微软负责任人工智能研究) University of California(加州大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21497 2025-10-31 cs.CV cs.AI cs.CL cs.MA 62%

Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers

Wei Pang, Kevin Qinghong Lin, Xiangru Jian, Xi He, Philip Torr

机构 * University of Waterloo(滑铁卢大学) University of Oxford(牛津大学) Vector Institute(向量研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Project Page: https://github.com/Paper2Poster/Paper2Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26052 2025-10-31 cs.CV cs.AI 57%

Dynamic VLM-Guided Negative Prompting for Diffusion Models

Hoyeon Chang, Seungjin Kim, Yoonseok Choi

机构 * KAIST(韩国科学技术院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: The First Workshop on Generative and Protective AI for Content Creation

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05747 2025-10-31 cs.CL 57%

SEA-LION: Southeast Asian Languages in One Network

Raymond Ng, Thanh Ngan Nguyen, Yuli Huang, Ngee Chia Tai, Wai Yi Leong, Wei Qi Leong, Xianbin Yong, Jian Gang Ngui, Yosephine Susanto, Nicholas Cheng, Hamsawardhini Rengarajan, Peerat Limkonchotiwat, Adithya Venkatadri Hulagadri, Kok Wai Teng, Yeo Yeow Tong, Bryan Siow, Wei Yi Teo, Wayne Lau, Choon Meng Tan, Brandon Ong, Zhi Hao Ong, Jann Railey Montalan, Adwin Chan, Sajeban Antonyrex, Ren Lee, Esther Choa, David Ong Tat-Wee, Bing Jie Darius Liu, William Chandra Tjhi, Erik Cambria, Leslie Teo

机构 * AI Singapore National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted at IJCNLP-AACL 2025 (Main Track). We released our model at https://huggingface.co/collections/aisingapore/sea-lionv3-672589a39cdadd6a5b199581

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04852 2025-10-31 cs.CV cs.LG 57%

CAUSAL3D: A Comprehensive Benchmark for Causal Learning from Visual Data

Disheng Liu, Yiran Qiao, Wuche Liu, Yiren Lu, Yunlai Zhou, Tuo Liang, Yu Yin, Jing Ma

机构 * Case Western Reserve University(凯斯西储大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Datasets link: https://huggingface.co/datasets/LLDDSS/Causal3D_Dataset

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06771 2025-10-31 cs.CV 50%

D-HUMOR: Dark Humor Understanding via Multimodal Open-ended Reasoning -- A Benchmark Dataset and Method

Sai Kartheek Reddy Kasu, Mohammad Zia Ur Rehman, Shahid Shafi Dar, Rishi Bharat Junghare, Dhanvin Sanjay Namboodiri, Nagendra Kumar

机构 * Indian Institute of Information Technology Dharwad, India(印度达拉瓦德信息科技学院) Indian Institute of Technology Indore, India(印度印度理工学院) Malaviya National Institute of Technology Jaipur, India(马拉维亚国家理工学院)

专题命中 安全评测 :alignment(abstract)

Comments Accepted at IEEE International Conference on Data Mining (ICDM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25402 2025-10-31 cs.IR cs.CE 50%

Towards Automated Quality Assurance of Patent Specifications: A Multi-Dimensional LLM Framework

Yuqian Chai, Chaochao Wang, Weilei Wang

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏