arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-20 至 2025-11-20 共收录 12 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 12 篇

2509.01418 2025-11-20 cs.CL 79%

On the Alignment of Large Language Models with Global Human Opinion

Yang Liu, Masahiro Kaneko, Chenhui Chu

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments 28 pages, 26 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04759 2025-11-20 cs.AI 79%

Driving with Regulation: Trustworthy and Interpretable Decision-Making for Autonomous Driving with Retrieval-Augmented Reasoning

Tianhui Cai, Yifan Liu, Zewei Zhou, Haoxuan Ma, Seth Z. Zhao, Zhiwen Wu, Xu Han, Zhiyu Huang, Jiaqi Ma

专题命中 安全评测 :trustworthy(title);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15206 2025-11-20 cs.CR cs.IT math.IT 78%

Trustworthy GenAI over 6G: Integrated Applications and Security Frameworks

Bui Duc Son, Trinh Van Chien, Dong In Kim

专题命中 安全评测 :trustworthy(title,abstract)

Comments 8 pages, 5 figures. Submitted for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14805 2025-11-20 cs.SE cs.AI 70%

Towards Continuous Assurance with Formal Verification and Assurance Cases

Dhaminda B. Abeywickrama, Michael Fisher, Frederic Wheeler, Louise Dennis

机构 * Department of Computer Science, The University of Manchester(曼彻斯特大学计算机科学系) Regulatory Support Directorate, Amentum(Amentum监管支持部门)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14010 2025-11-20 cs.CL cs.AI 62%

Knowledge-Grounded Agentic Large Language Models for Multi-Hazard Understanding from Reconnaissance Reports

Chenchen Kuai, Zihao Li, Braden Rosen, Stephanie Paal, Navid Jafari, Jean-Louis Briaud, Yunlong Zhang, Youssef M. A. Hashash, Yang Zhou

机构 * organization= Department One , addressline= Address One , city= City One , postcode= 00000 , state= State One , country= Country One organization= Department Two , addressline= Address Two , city= City Two , postcode= 22222 , state= State Two , country= Country Two organization= Zachry Department of Civil \& Environmental Engineering, Texas A\&M University , addressline= 3136 TAMU , city= College Station , postcode= 77843 , state= TX , country= USA organization= Department of Engineering Technology Industrial Distribution, Texas A\&M University , city= College Station , postcode= 77843 , state= TX , country= USA organization= Department of Civil Environmental Engineering, University of Illinois Urbana-Champaign , city= Urbana , postcode= 61801 , state= IL , country= USA

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments 17 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14767 2025-11-20 cs.IR cs.AI cs.CY 62%

An LLM-Powered Agent for Real-Time Analysis of the Vietnamese IT Job Market

Minh-Thuan Nguyen, Thien Vo-Thanh, Thai-Duy Dinh, Xuan-Quang Phan, Tan-Ha Mai, Lam-Son Lê

机构 * Computer Science Department Vietnamese-German University, Vietnam(越南德意志大学计算机科学系) Business Administration Department FPT University, Vietnam(越南FPT大学商学院) IT Operation Department Mantu Group, Vietnam(越南Mantu集团IT运营部) CSIE Department National Taiwan University, Taiwan(台湾国立台湾大学CSIE系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments Accepted at ACOMPA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17979 2025-11-20 cs.AI cs.CL 62%

Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities

Weixiang Zhao, Xingyu Sui, Jiahe Guo, Yulin Hu, Yang Deng, Yanyan Zhao, Xuda Zhi, Yongbo Huang, Hao He, Wanxiang Che, Ting Liu, Bing Qin

专题命中 安全评测 :harmlessness(abstract);分类 cs.CL、cs.AI

Comments To appear at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15203 2025-11-20 cs.CR cs.AI 57%

Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks

Zimo Ji, Xunguang Wang, Zongjie Li, Pingchuan Ma, Yudong Gao, Daoyuan Wu, Xincheng Yan, Tian Tian, Shuai Wang

机构 * The Hong Kong University of Science and Technology(香港科技大学) Zhejiang University of Technology(浙江工业大学) Lingnan University(岭南大学) School of Cyber Science and Engineering, Southeast University(东南大学计算机科学与工程学院) ZTE Corporation(中兴通讯有限公司)

专题命中 安全评测 :prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14903 2025-11-20 cs.LG cs.SE 57%

It's LIT! Reliability-Optimized LLMs with Inspectable Tools

Ruixin Zhang, Jon Donnelly, Zhicheng Guo, Ghazal Khalighinejad, Haiyang Huang, Alina Jade Barnett, Cynthia Rudin

机构 * Department of Computer Science(计算机科学系) Duke University(杜克大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Accepted to the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop on Multi-Turn Interactions in Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14439 2025-11-20 cs.CL 57%

MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents

Jinru Ding, Lu Lu, Chao Ding, Mouxiao Bian, Jiayuan Chen, Wenrao Pang, Ruiyao Chen, Xinwei Peng, Renjie Lu, Sijie Ren, Guanxu Zhu, Xiaoqin Wu, Zhiqiang Liu, Rongzhao Zhang, Luyi Jiang, Bing Han, Yunqiu Wang, Jie Xu

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12527 2025-11-20 cs.LG stat.ML 57%

Selective Risk Certification for LLM Outputs via Information-Lift Statistics: PAC-Bayes, Robustness, and Skeleton Design

Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma

机构 * Department of Computer Science Iowa State University(计算机科学系爱荷华州立大学) Department of Civil, Construction and Environmental Engineering Iowa State University(土木、建设与环境工程系爱荷华州立大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15308 2025-11-20 cs.CV 50%

Text2Loc++: Generalizing 3D Point Cloud Localization from Natural Language

Yan Xia, Letian Shi, Yilin Di, Joao F. Henriques, Daniel Cremers

机构 * School of Artificial Intelligence and Data Science, University of Science and Technology of China(人工智能与数据科学学院,中国科学技术大学) Technical University of Munich(慕尼黑技术大学) Visual Geometry Group, University of Oxford(牛津大学视觉几何组)

专题命中 安全评测 :alignment(abstract)

Comments This paper builds upon and extends our earlier conference paper Text2Loc presented at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏