arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9400 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9400 篇

2512.04118 2025-12-05 cs.CY cs.LG 81%

Patient Safety Risks from AI Scribes: Signals from End-User Feedback

人工智能记录员对患者安全的风险:来自终端用户反馈的信号

Jessica Dai, Anwen Huang, Catherine Nasrallah, Rhiannon Croci, Hossein Soleimani, Sarah J. Pollet, Julia Adler-Milstein, Sara G. Murray, Jinoos Yazdany, Irene Y. Chen

机构 * University of California Berkeley(加州大学伯克利分校) University of California San Francisco(加州大学旧金山分校)

专题命中 安全评测 :safety(title,abstract);分类 cs.CY、cs.LG

AI总结 本研究通过混合方法分析了AI记录员在医疗反馈中引发的患者安全风险,发现转录错误可能导致药物和治疗方面的风险,需进一步研究以明确风险程度。

Comments ML4H Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01786 2025-12-02 cs.AI cs.LG 81%

Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems

谁评判法官?LLM陪审团即需:构建可信的LLM评估系统

Xiaochuan Li, Ke Wang, Girija Gouda, Shubham Choudhary, Yaqun Wang, Linwei Hu, Joel Vaughan, Freddy Lecue

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出LLM陪审团即需,通过动态学习框架提升LLM评估的可靠性和可扩展性,实验表明其在关键决策中的相关性优于传统方法。

Comments 66 pages, 22 figures, 37 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14469 2025-12-02 cs.CL cs.AI 81%

Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations

大语言模型在混合语言扰动下的归因安全故障

Somnath Banerjee, Pratyush Chatterjee, Shanu Kumar, Sayan Layek, Parag Agrawal, Rima Hazra, Animesh Mukherjee

机构 * Microsoft Corporation(微软公司)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

AI总结 研究发现大语言模型在混合语言扰动下存在归因安全故障,提出显著性漂移归因框架并提出翻译基恢复策略以提升安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18008 2025-12-02 cs.CY cs.CL 81%

GermanPartiesQA: Benchmarking Commercial Large Language Models and AI Companions for Political Alignment and Sycophancy

GermanPartiesQA: 对商业大语言模型和AI伴侣在政治倾向和趋炎附势方面的基准测试

Jan Batzner, Volker Stocker, Stefan Schmid, Gjergji Kasneci

机构 * TUM(慕尼黑大学) TUB(慕尼黑大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.CY

AI总结 GermanPartiesQA研究了六个商业LLMs在政治倾向和趋炎附势方面的表现,揭示了LLMs在生成事实性政党立场上的局限性以及模型特定的意识形态倾向模式。

Comments Published at AAAI/ACM AIES 2025. Presented at NeurIPS 2025 Workshop on LLM Evaluation and the International Monetary Fund's 12th Statistical Forum. GermanPartiesQA Benchmark under https://github.com/janbatzner/germanpartiesqa

Journal ref Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 8(1), 2025, pp. 330-342

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21737 2025-12-01 cs.CL cs.AI 81%

Polarity-Aware Probing for Quantifying Latent Alignment in Language Models

考虑极性的人探针用于量化语言模型中的潜在对齐

Sabrina Sadiekh, Elena Ericheva, Chirag Agarwal

机构 * University of Virginia(弗吉尼亚大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 提出考虑极性的CCS方法,用于量化语言模型潜在对齐的鲁棒性,揭示模型内部表示的一致性差异。

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21363 2025-11-27 cs.LG cs.AI 81%

The Directed Prediction Change - Efficient and Trustworthy Fidelity Assessment for Local Feature Attribution Methods

定向预测变化 - 用于局部特征归因方法高效且可信的保真度评估

Kevin Iselborn, David Dembinsky, Adriano Lucieri, Andreas Dengel

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出DPC度量标准,用于高效且可信地评估局部特征归因方法的保真度,通过改进预测变化度量并消除随机性,实现快速且确定性的评估。

Comments 13 pages, 10 figures, 5 tables, accepted at AAAI SECURE-AI4H workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20726 2025-11-27 cs.LG cs.AI 81%

Learning from Risk: LLM-Guided Generation of Safety-Critical Scenarios with Prior Knowledge

从风险学习:利用先验知识的LLM引导的安全关键场景生成

Yuhang Wang, Heye Huang, Zhenhua Xu, Kailai Sun, Baoshen Guo, Jinhua Zhao

机构 * Chinese Academy of Sciences, China(中国科学院) Department of Urban Studies and Planning, Massachusetts Institute of Technology, USA(麻省理工学院城市研究与规划系) School of Vehicle and Mobility, Tsinghua University, China(清华大学车辆与移动系统学院)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出利用LLM和CVAE生成安全关键场景的方法,通过知识驱动优化提升自动驾驶系统在罕见高风险事件下的鲁棒性与可控性。

Comments 24 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01710 2025-11-11 cs.CL cs.LG 81%

CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications

Raviraj Joshi, Rakesh Paul, Kanishk Singla, Anusha Kamath, Michael Evans, Katherine Luna, Shaona Ghosh, Utkarsh Vaidya, Eileen Long, Sanjay Singh Chauhan, Niranjan Wartikar

机构 * NVIDIA(英伟达)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07416 2025-11-07 cs.CV cs.CL cs.LG 81%

RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability

Jonggwon Park, Byungmu Yoon, Soobum Kim, Kyoyun Choi

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26024 2025-10-31 cs.CL cs.AI 81%

Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs

HyoJung Han, Sweta Agrawal, Eleftheria Briakou

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22095 2025-10-28 cs.AI cs.CL 81%

Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies

Yankai Chen, Xinni Zhang, Yifei Zhang, Yangning Li, Henry Peng Zou, Chunyu Miao, Weizhi Zhang, Xue Liu, Philip S. Yu

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) MBZUAI(马克斯·普朗克人工智能研究所) McGill University(麦吉尔大学) The Chinese University of Hong Kong(香港中文大学) Nanyang Technological University(南洋理工大学) Tsinghua University(清华大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by NeurIPS'25 Position Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25271 2025-10-24 cs.AI cs.CV cs.LG cs.MA 81%

RADAR: A Risk-Aware Dynamic Multi-Agent Framework for LLM Safety Evaluation via Role-Specialized Collaboration

Xiuyuan Chen, Jian Zhao, Yuchen Yuan, Tianle Zhang, Huilin Zhou, Zheng Zhu, Ping Hu, Linghe Kong, Chi Zhang, Weiran Huang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信) School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) University of Science and Technology of China(中国科学技术大学) GigaAI School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04150 2025-10-22 cs.CL cs.AI 81%

Temporal Alignment of LLMs through Cycle Encoding for Long-Range Time Representations

Xue Han, Qian Hu, Yitong Wang, Wenchun Gao, Lianlian Zhang, Qing Wang, Lijun Mei, Chao Deng, Junlan Feng

机构 * JIUTIAN Team China Mobile Research Institute(中国移动研究院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09710 2025-10-16 cs.CL cs.AI 81%

SeCon-RAG: A Two-Stage Semantic Filtering and Conflict-Free Framework for Trustworthy RAG

Xiaonan Si, Meilin Zhu, Simeng Qin, Lijia Yu, Lijun Zhang, Shuaitong Liu, Xinfeng Li, Ranjie Duan, Yang Liu, Xiaojun Jia

机构 * Institute of Software Chinese Academy of Sciences Beijing China(中国科学院软件研究所) Key Laboratory of System Software (Chinese Academy of Sciences) and State Key Laboratory of Computer Science, Institute of Software, Chinese Academy of Sciences, Beijing, China(中国科学院系统软件重点实验室和计算机科学国家重点实验室) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学) Northeast University China(东北大学) Institute of Ai For industries Nanjing China(人工智能产业研究院) Southwest University China(西南大学) Nanyang Technological University Singapore(新加坡南洋理工大学) Alibaba China(阿里巴巴(中国))

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17244 2025-10-16 cs.CL cs.AI 81%

ReasoningShield: Safety Detection over Reasoning Traces of Large Reasoning Models

Changyi Li, Jiayi Wang, Xudong Pan, Geng Hong, Min Yang

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08983 2025-10-13 cs.SE cs.AI cs.LG 81%

Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations

David N. Palacio, Daniel Rodriguez-Cardenas, Alejandro Velasco, Dipin Khati, Kevin Moran, Denys Poshyvanyk

机构 * University of Central Florida(佛罗里达中央大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments Under Review to appear in ACM Transactions on Software Engineering and Methodology (TOSEM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10663 2025-10-08 q-bio.NC cs.AI cs.CV cs.LG 81%

Optimal Transport for Brain-Image Alignment: Unveiling Redundancy and Synergy in Neural Information Processing

Yang Xiao, Wang Lu, Jie Ji, Ruimeng Ye, Gen Li, Xiaolong Ma, Bo Hui

机构 * University of Tulsa(图拉大学) Tsinghua University(清华大学) Clemson University(克莱姆森大学) The University of Arizona(亚利桑那大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 14pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01606 2025-10-03 cs.IR cs.AI cs.CL 81%

Bridging Collaborative Filtering and Large Language Models with Dynamic Alignment, Multimodal Fusion and Evidence-grounded Explanations

Bo Ma, LuYao Liu, Simon Lau, Chandler Yuan, and XueY Cui, Rosie Zhang

机构 * Department of Software \& Microelectronics, Peking University, Beijing, China Economic Law School, China University of Political Science Financial Media, Peking University, ChangSha, China

专题命中 安全评测 :alignment(title);trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01231 2025-10-03 cs.CL cs.AI stat.ML 81%

Trustworthy Summarization via Uncertainty Quantification and Risk Awareness in Large Language Models

Shuaidong Pan, Di Wu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15463 2025-10-01 cs.HC cs.AI cs.CL 81%

Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?

Hua Shen, Nicholas Clark, Tanushree Mitra

机构 * University of Washington(华盛顿大学) NYU Shanghai(纽约大学上海校区) New York University(纽约大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 Main Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13742 2025-09-23 cs.LG cs.AI math.OC 81%

Search-Optimized Quantization in Biomedical Ontology Alignment

Oussama Bouaggad, Natalia Grabar

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments Accepted for publication in Frontiers in Artificial Intelligence - Medicine and Public Health (Original Research)

Journal ref Front. Artif. Intell. 8:1662984 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16394 2025-09-23 cs.CL cs.AI cs.HC 81%

Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans

Deuksin Kwon, Kaleen Shrestha, Bin Han, Elena Hayoung Lee, Gale Lucas

机构 * University of Southern California(南加州大学) USC for Institute of Creative Technologies(南加州大学创意技术研究所)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15932 2025-09-22 cs.LG cs.AI cs.IT math.IT stat.ML 81%

The Alignment Bottleneck

Wenjun Cao

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08380 2025-09-18 cs.AI cs.LG 81%

Co-Investigator AI: The Rise of Agentic AI for Smarter, Trustworthy AML Compliance Narratives

Prathamesh Vasudeo Naik, Naresh Kumar Dintakurthi, Zhanghao Hu, Yue Wang, Robby Qiu

专题命中 安全评测 :trustworthy(title);alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01166 2025-09-03 cs.CL cs.AI 81%

Enhancing Large Language Model for Knowledge Graph Completion via Structure-Aware Alignment-Tuning

Yu Liu, Yanan Cao, Xixun Lin, Yanmin Shang, Shi Wang, Shirui Pan

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Griffith University(格里菲斯大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025, Main, Long Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02531 2025-08-28 cs.CY cs.CL 81%

Towards New Benchmark for AI Alignment & Sentiment Analysis in Socially Important Issues: A Comparative Study of Human and LLMs in the Context of AGI

Ljubisa Bojic, Dylan Seychell, Milan Cabarkapa

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.CY

Comments 34 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15192 2025-08-22 cs.AI cs.CL 81%

LLM4Sweat: A Trustworthy Large Language Model for Hyperhidrosis Support

Wenjie Lin, Jin Wei-Kocsis

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22146 2025-08-22 cs.CV cs.AI cs.CL q-bio.NC 81%

Flexible Tool Selection through Low-dimensional Attribute Alignment of Vision and Language

Guangfu Hao, Haojie Wen, Liangxuan Guo, Yang Chen, Yanchao Bi, Shan Yu

机构 * Laboratory of Brain Atlas and Brain-inspired Intelligence, Institute of Automation Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所脑图谱与类脑智能实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS)(中国科学院大学人工智能学院) School of Systems Science, Beijing Normal University(北京师范大学系统科学学院) School of Psychological and Cognitive Sciences & Beijing Key Laboratory of Behavior and Mental Health, Peking University(北京大学心理与认知科学学院) IDG/McGovern Institute for Brain Research, Peking University(北京大学IDG/ McGovern脑科学研究院) Institute for Artificial Intelligence & Key Laboratory of Machine Perception (Ministry of Education), Peking University(北京大学人工智能研究所) School of Future Technology, University of Chinese Academy of Sciences (UCAS)(中国科学院大学未来技术学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14735 2025-08-21 cs.CL cs.AI 81%

Evaluating Multilingual and Code-Switched Alignment in LLMs via Synthetic Natural Language Inference

Samir Abdaljalil, Erchin Serpedin, Khalid Qaraqe, Hasan Kurban

机构 * Texas A\&M University, College Station, TX., USA(德克萨斯大学) Hamad Bin Khalifa University, Doha, Qatar(哈马德·本·卡伊夫大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11247 2025-08-19 cs.AI cs.LG cs.RO 81%

LD-Scene: LLM-Guided Diffusion for Controllable Generation of Adversarial Safety-Critical Driving Scenarios

Mingxing Peng, Yuting Xie, Xusen Guo, Ruoyu Yao, Hai Yang, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏