arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-14 至 2025-11-14 共收录 11 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 11 篇

2502.08045 2025-11-14 cs.CL cs.AI cs.CY 82%

Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs

Mohsinul Kabir, Ajwad Abrar, Sophia Ananiadou

机构 * Department of Computer Science, National Center for Text Mining, The University of Manchester(计算机科学系,文本挖掘国家中心,曼彻斯特大学) Department of Computer Science and Engineering, Islamic University of Technology(计算机科学与工程系,伊斯兰技术大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted at EMNLP 2025 (Main)

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09964 2025-11-14 cs.SE cs.AI cs.PL 79%

EnvTrace: Simulation-Based Semantic Evaluation of LLM Code via Execution Trace Alignment -- Demonstrated at Synchrotron Beamlines

Noah van der Vleuten, Anthony Flores, Shray Mathur, Max Rakitin, Thomas Hopkins, Kevin G. Yager, Esther H. R. Tsai

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07871 2025-11-14 cs.CL 79%

AlignSurvey: A Comprehensive Benchmark for Human Preferences Alignment in Social Surveys

Chenxi Lin, Weikang Yuan, Zhuoren Jiang, Biao Huang, Ruitao Zhang, Jianan Ge, Yueqian Xu, Jianxing Yu

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15434 2025-11-14 cs.CV cs.LG 79%

Semantic4Safety: Causal Insights from Zero-shot Street View Imagery Segmentation for Urban Road Safety

Huan Chen, Ting Han, Siyu Chen, Zhihao Guo, Yiping Chen, Meiliu Wu

机构 * School of Geospatial Engineering and Science, Sun Yat-sen University(地理空间工程与科学学院,中山大学) School of Geographical and Earth Sciences, University of Glasgow(地理与地球科学学院,格拉斯哥大学) School of Economics and Management, Shanxi University(经济学与管理学院,山西大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.LG

Comments 11 pages, 10 figures, The 8th ACM SIGSPATIAL International Workshop on AI for Geographic Knowledge Discovery (GeoAI '25), November 3--6, 2025, Minneapolis, MN, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09855 2025-11-14 cs.LG 74%

Unlearning Imperative: Securing Trustworthy and Responsible LLMs through Engineered Forgetting

James Jin Kang, Dang Bui, Thanh Pham, Huo-Chong Ling

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments 14 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09748 2025-11-14 cs.CL cs.AI 62%

How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation

Muskaan Chopra, Lorenz Sparrenberg, Sarthak Khanna, Rafet Sifa

机构 * Fraunhofer IAIS - Department of Media Engineering(弗劳恩霍夫人工智能研究所-媒体工程部门) University of Bonn - Department of Computer Science(波恩大学-计算机科学系) Lamarr Institute for Machine Learning(拉马尔人工智能与机器学习研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted in IEEE BigData 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10203 2025-11-14 cs.CV cs.AI cs.RO 57%

VISTA: A Vision and Intent-Aware Social Attention Framework for Multi-Agent Trajectory Prediction

Stephane Da Silva Martins, Emanuel Aldea, Sylvie Le Hégarat-Mascle

机构 * SATIE - CNRS UMR 8029 Paris-Saclay University, France(巴黎-萨克雷大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Paper accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07127 2025-11-14 cs.LG 57%

REACT-LLM: A Benchmark for Evaluating LLM Integration with Causal Features in Clinical Prognostic Tasks

Linna Wang, Zhixuan You, Qihui Zhang, Jiunan Wen, Ji Shi, Yimin Chen, Yusen Wang, Fanqi Ding, Ziliang Feng, Li Lu

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03553 2025-11-14 cs.CL 57%

CCD-Bench: Probing Cultural Conflict in Large Language Model Decision-Making

Hasibur Rahman, Hanan Salam

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05796 2025-11-14 cs.CV cs.AI 57%

Dual-Mode Deep Anomaly Detection for Medical Manufacturing: Structural Similarity and Feature Distance

Julio Zanon Diaz, Georgios Siogkas, Peter Corcoran

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 12 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09742 2025-11-14 cs.CV cs.AI 57%

Feature Quality and Adaptability of Medical Foundation Models: A Comparative Evaluation for Radiographic Classification and Segmentation

Frank Li, Theo Dapamede, Mohammadreza Chavoshi, Young Seok Jeon, Bardia Khosravi, Abdulhameed Dere, Beatrice Brown-Mulry, Rohan Satya Isaac, Aawez Mansuri, Chiratidzo Sanyika, Janice Newsome, Saptarshi Purkayastha, Imon Banerjee, Hari Trivedi, Judy Gichoya

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 7 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏