arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-04 至 2025-11-04 共收录 16 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 16 篇

2506.03659 2025-11-04 cs.CL 85%

Trustworthy Medical Question Answering: An Evaluation-Centric Survey

Yinuo Wang, Baiyang Wang, Robert E. Mercer, Frank Rudzicz, Sudipta Singha Roy, Pengjie Ren, Zhumin Chen, Xindi Wang

机构 * Shandong University(山东大学) University of Western Ontario(西方大学) Dalhousie University(达尔豪斯大学) University of Toronto(多伦多大学)

专题命中 安全评测 :trustworthy(title,abstract);alignment(abstract);safety(abstract);分类 cs.CL

Comments accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07170 2025-11-04 cs.LG cs.AI 82%

Trustworthy AI Must Account for Interactions

Jesse C. Cresswell

机构 * Jesse C. Cresswell(独立研究者)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG;alignment(comments)

Comments Presented at the ICLR 2025 Workshop on Bidirectional Human-AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03665 2025-11-04 cs.LG cs.AI 76%

A DbC Inspired Neurosymbolic Layer for Trustworthy Agent Design

Claudiu Leoveanu-Condrei

机构 * ExtensityAI

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments 4 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00843 2025-11-04 cs.HC 75%

Portal UX Agent -- A Plug-and-Play Engine for Rendering UIs from Natural Language Specifications

Xinsong Li, Ning Jiang, Jay Selvaraj

专题命中 安全评测 :alignment(abstract);safety(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26781 2025-11-04 cs.CV 71%

ChartAB: A Benchmark for Chart Grounding & Dense Alignment

Aniruddh Bansal, Davit Soselia, Dang Nguyen, Tianyi Zhou

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 安全评测 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10707 2025-11-04 cs.LG cs.AI 62%

ConTextTab: A Semantics-Aware Tabular In-Context Learner

Marco Spinaci, Marek Polewczyk, Maximilian Schambach, Sam Thelin

机构 * SAP France(SAP法国分公司) SAP SE(SAP德国分公司)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted as spotlight at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14622 2025-11-04 cs.CR cs.AI cs.LG 62%

Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection

Yihao Guo, Haocheng Bian, Liutong Zhou, Ze Wang, Zhaoyi Zhang, Francois Kawala, Milan Dean, Ian Fischer, Yuantao Peng, Noyan Tokgozoglu, Ivan Barrientos, Riyaaz Shaik, Rachel Li, Chandru Venkataraman, Reza Shifteh Far, Moses Pawar, Venkat Sundaranatha, Michael Xu, Frank Chu

机构 * Apple(苹果公司) Cohere(Cohere公司) DeepMind(深度思维公司) Meta MongoDB(MongoDB公司)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01226 2025-11-04 cs.LG 57%

WindMiL: Equivariant Graph Learning for Wind Loading Prediction

Themistoklis Vargiemezis, Charilaos Kanatsoulis, Catherine Gorlé

机构 * Department of Civil & Environmental Engineering, Stanford, CA, USA(土木与环境工程系,斯坦福大学) Department of Computer Science, Stanford, CA, USA(计算机科学系,斯坦福大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00122 2025-11-04 cs.AI 57%

Engineering.ai: A Platform for Teams of AI Engineers in Computational Design

Ran Xu, Yupeng Qi, Jingsen Feng, Xu Chu

机构 * Faculty for Aerospace Engineering and Geodesy, University of Stuttgart, Stuttgart, Germany(航空航天工程与大地测量学系,斯图加特大学) Cluster of Excellence SimTech, University of Stuttgart, Stuttgart, Germany(卓越中心SimTech,斯图加特大学) Faculty of Environment, Science and Economy, University of Exeter, Exeter EX4 4QF, United Kingdom(环境、科学与经济学院,埃克塞特大学) University of Stuttgart, Stuttgart, Germany(斯图加特大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17267 2025-11-04 cs.CL 57%

GreekBarBench: A Challenging Benchmark for Free-Text Legal Reasoning and Citations

Odysseas S. Chlapanis, Dimitrios Galanis, Nikolaos Aletras, Ion Androutsopoulos

机构 * Department of Informatics, Athens University of Economics and Business(信息学院,雅典经济与商业大学) Archimedes, Athena Research Center(阿提卡研究中心-阿基米德) Athena Research Center(阿提卡研究中心) University of Sheffield(谢菲尔德大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 19 pages, 17 figures, accepted in EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01341 2025-11-04 cs.CL 57%

AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding

Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang, Suyuchen Wang, Chao Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu, Nicolas Chapados, Yoshua Bengio, Enamul Hoque, Christopher Pal, Issam H. Laradji, David Vazquez, Perouz Taslakian, Spandana Gella, Sai Rajeswar

机构 * ServiceNow York University(约克大学) Mila – Quebec AI Institute(魁北克人工智能研究院) École de Technologie Supérieure(魁北克高等技术学院) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) University of Waterloo(滑铁卢大学) CIFAR AI Chair(CIFAR人工智能 chair) Polytechnique Montréal(蒙特利尔理工学院) University of British Columbia(不列颠哥伦比亚大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01757 2025-11-04 cs.SE 50%

Towards LLM-Powered Task-Aware Retrieval of Scientific Workflows for Galaxy

Shamse Tasnim Cynthia, Banani Roy

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05887 2025-11-04 eess.AS 50%

Aligning Speech to Languages to Enhance Code-switching Speech Recognition

Hexin Liu, Xiangyu Zhang, Haoyang Zhang, Leibny Paola Garcia, Andy W. H. Khong, Eng Siong Chng, Shinji Watanabe

专题命中 安全评测 :alignment(abstract)

Comments Accepted to IEEE Trans. Audio Speech Lang. Process., copyright has been transferred to IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00659 2025-11-04 eess.SY cs.SY 50%

Unveiling Uniform Shifted Power Law in Stochastic Human and Autonomous Driving Behavior

Wang Chen, Heye Huang, Ke Ma, Hangyu Li, Shixiao Liang, Hang Zhou, Xiaopeng Li

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00613 2025-11-04 cs.CV 50%

CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World

Yating Yu, Congqi Cao, Zhaoying Wang, Weihua Meng, Jie Li, Yuxin Li, Zihao Wei, Zhongpei Shen, Jiajun Zhang

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14918 2025-11-04 cs.CV 50%

Semantic-Aware Representation Learning via Conditional Transport for Multi-Label Image Classification

Ren-Dong Xie, Zhi-Fen He, Bo Li, Bin Liu, Jin-Yan Hu

机构 * School of Mathematics and Information Science(数学与信息科学学院) Key Laboratory of Jiangxi Province for Image Processing and Pattern Recognition(江西省图像处理与模式识别重点实验室)

专题命中 安全评测 :alignment(abstract)

Comments The paper is under consideration at Pattern Recognition Letters

详情

展开后加载摘要…

URL PDF HTML 收藏