arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-13 至 2025-11-13 共收录 13 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 13 篇

2511.04962 2025-11-13 cs.CL cs.AI 73%

Too Good to be Bad: On the Failure of LLMs to Role-Play Villains

Zihao Yi, Qingxuan Jiang, Ruotian Ma, Xingyu Chen, Qu Yang, Mengru Wang, Fanghua Ye, Ying Shen, Zhaopeng Tu, Xiaolong Li, Linus

机构 * Tencent(腾讯)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08702 2025-11-13 cs.LG cs.AI cs.CR cs.CY 67%

FAIRPLAI: A Human-in-the-Loop Approach to Fair and Private Machine Learning

David Sanchez, Holly Lopez, Michelle Buraczyk, Anantaa Kotal

机构 * Dept. of Computer Science, The University of Texas at El Paso(得克萨斯大学埃尔帕索分校计算机科学系) Dept. of Mathematics, Mountain View High School(山景高中数学系) Dept. of Mathematics, El Paso Independent School District(埃尔帕索独立学区数学系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09067 2025-11-13 cs.CL cs.AI 62%

MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique

Gailun Zeng, Ziyang Luo, Hongzhan Lin, Yuchen Tian, Kaixin Li, Ziyang Gong, Jianxiong Guo, Jing Ma

机构 * Hong Kong Baptist University(香港 Baptist 大学) Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist 大学) National University of Singapore(新加坡国立大学) Beijing Normal University(北京师范大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments 28 pages, 14 figures, 19 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09363 2025-11-13 cs.AI cs.GT cs.LG 62%

ElicitationGPT: Text Elicitation Mechanisms via Language Models

Yifan Wu, Jason Hartline

机构 * Microsoft Research(微软研究院) Northwestern University(西北大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09443 2025-11-13 cs.CV cs.AI 57%

BronchOpt : Vision-Based Pose Optimization with Fine-Tuned Foundation Models for Accurate Bronchoscopy Navigation

Hongchao Shu, Roger D. Soberanis-Mukul, Jiru Xu, Hao Ding, Morgan Ringel, Mali Shen, Saif Iftekar Sayed, Hedyeh Rafii-Tari, Mathias Unberath

机构 * Johnson & Johnson MedTech(强生医疗科技)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03133 2025-11-13 cs.CL 57%

ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models

Boyang Xue, Qi Zhu, Rui Wang, Sheng Wang, Hongru Wang, Minda Hu, Fei Mi, Yasheng Wang, Lifeng Shang, Qun Liu, Kam-Fai Wong

机构 * The Chinese University of Hong Kong(香港中文大学) Huawei Noah’s Ark Lab(华为诺亚实验室) The University of Hong Kong(香港大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23015 2025-11-13 cs.CL 57%

Detecting Stealthy Backdoor Samples based on Intra-class Distance for Large Language Models

Jinwen Chen, Hainan Zhang, Fei Sun, Qinnan Zhang, Sijia Wen, Ziwei Wang, Zhiming Zheng

机构 * Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(未来区块链与隐私计算先进创新中心) Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments EMNLP2025Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16774 2025-11-13 cs.CL 57%

IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models

Yiming Gao, Bin Wang, Chengwei Wei, Shuo Sun, AiTi Aw

机构 * Nanyang Technological University (NTU)(南洋理工大学) MiroMind(米罗Mind) Institute for Infocomm Research (I 2 R)(信息与通信研究院) A*STAR(科技研究局)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Link: https://github.com/AudioLLMs/AudioBench/tree/main/IFEval-Audio

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06227 2025-11-13 cs.CL 57%

LExT: Towards Evaluating Trustworthiness of Natural Language Explanations

Krithi Shailya, Shreya Rajpal, Gokul S Krishnan, Balaraman Ravindran

机构 * Centre for Responsible AI, IIT Madras(责任人工智能中心,IIT马德拉斯)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08649 2025-11-13 q-bio.QM cs.AI 57%

Bio AI Agent: A Multi-Agent Artificial Intelligence System for Autonomous CAR-T Cell Therapy Development with Integrated Target Discovery, Toxicity Prediction, and Rational Molecular Design

Yi Ni, Liwei Zhu, Shuai Li

机构 * Bio LIMS INC(Bio LIMS公司)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 12 pages, 0 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15153 2025-11-13 cs.CL 57%

Evaluating Deep Unlearning in Large Language Models

Ruihan Wu, Chhavi Yadav, Russ Salakhutdinov, Kamalika Chaudhuri

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02557 2025-11-13 eess.IV cs.CV 50%

RL-U$^2$Net: A Dual-Branch UNet with Reinforcement Learning-Assisted Multimodal Feature Fusion for Accurate 3D Whole-Heart Segmentation

Jierui Qu, Jianchun Zhao

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15526 2025-11-13 eess.IV cs.CV 50%

Multi-scale Cascaded Foundation Model for Whole-body Organs-at-risk Segmentation

Rui Hao, Dayu Tan, Qiankun Li, Chunhou Zheng, Weimin Zhong, Zhigang Zeng

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学) Institute of Artificial Intelligence, Huazhong University of Science and Technology(人工智能研究院,华中科技大学) Hubei Key Laboratory of Brain-Inspired Intelligent Systems, Huazhong University of Science and Technology(湖北省脑启发智能系统重点实验室,华中科技大学) Key Laboratory of Image Processing and Intelligent Control (Huazhong University of Science and Technology), Ministry of Education(图像处理与智能控制重点实验室(华中科技大学),教育部) Key Laboratory of Intelligent Computing and Signal Processing, Ministry of Education, Anhui University(智能计算与信号处理重点实验室(安徽大学),教育部) College of Computing and Data Science (CCDS), Nanyang Technological University(计算与数据科学学院(CCDS),南洋理工大学) East China University of Science and Technology(东华大学)

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏