arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-22 至 2025-08-22 共收录 41 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 20 篇

2508.14931 2025-08-22 eess.IV cs.GR 50%

Pixels Under Pressure: Exploring Fine-Tuning Paradigms for Foundation Models in High-Resolution Medical Imaging

Zahra TehraniNasab, Amar Kumar, Tal Arbel

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12633 2025-08-22 eess.SY cs.SY 50%

DCT-MARL: A Dynamic Communication Topology-Based MARL Algorithm for Connected Vehicle Platoon Control

Yaqi Xu, Yan Shi, Jin Tian, Fanzeng Xia, Tongxin Li, Shanzhi Chen, Yuming Ge

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18246 2025-08-22 cs.CV 50%

Referring Expression Instance Retrieval and A Strong End-to-End Baseline

Xiangzhao Hao, Kuan Zhu, Hongyu Guo, Haiyun Guo, Ning Jiang, Quan Lu, Ming Tang, Jinqiao Wang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Mashang Consumer Finance Co, Ltd(马商消费金融有限公司)

专题命中 安全评测 :alignment(abstract)

Comments ACMMM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13067 2025-08-22 cond-mat.mtrl-sci physics.chem-ph 50%

Integrating Density Functional Theory with Deep Neural Networks for Accurate Voltage Prediction in Alkali-Metal-Ion Battery Materials

Sk Mujaffar Hossain, Namitha Anna Koshi, Seung-Cheol Lee, G. P Das, Satadeep Bhattacharjee

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 2 篇

2502.15860 2025-08-22 cs.CL cs.AI cs.LG 67%

Synthetic vs. Gold: The Role of LLM Generated Labels and Data in Cyberbullying Detection

Arefeh Kazemi, Sri Balaaji Natarajan Kalaivendan, Joachim Wagner, Hamza Qadeer, Kanishk Verma, Brian Davis

机构 * School of Computing, ADAPT Centre, Dublin City University, Dublin, Ireland(计算学院、ADAPT中心、都柏林城市大学、都柏林、爱尔兰)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09222 2025-08-22 cs.CY 61%

Democratic AI is Possible. The Democracy Levels Framework Shows How It Might Work

Aviv Ovadya, Kyle Redman, Luke Thorburn, Quan Ze Chen, Oliver Smith, Flynn Devine, Andrew Konya, Smitha Milli, Manon Revel, K. J. Kevin Feng, Amy X. Zhang, Bilva Chandra, Michiel A. Bakker, Atoosa Kasirzadeh

专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.CY

Comments 31 pages. Accepted to the position paper track at ICML 2025. A previous version was presented at the Pluralistic Alignment Workshop at NeurIPS 2024. For ongoing work, see: https://democracylevels.org

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 5 篇

2507.06992 2025-08-22 cs.CV cs.AI 74%

MCA-RG: Enhancing LLMs with Medical Concept Alignment for Radiology Report Generation

Qilong Xing, Zikai Song, Youjia Zhang, Na Feng, Junqing Yu, Wei Yang

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 其他安全 :alignment(title);分类 cs.AI

Comments MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15262 2025-08-22 cs.IR cs.AI 57%

M-$LLM^3$REC: A Motivation-Aware User-Item Interaction Framework for Enhancing Recommendation Accuracy with LLMs

Lining Chen, Qingwen Zeng, Huaming Chen

机构 * The University of Sydney(悉尼大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 10pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15346 2025-08-22 q-bio.QM cs.LG 57%

Drug-Target Interaction/Affinity Prediction: Deep Learning Models and Advances Review

Ali Vefghi, Zahed Rahmati, Mohammad Akbari

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments 64 pages, 7 figures, 10 tables

Journal ref Journal of Computers in Biology and Medicine Volume 196, Part A, September 2025, 110438

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15207 2025-08-22 cs.CV 50%

Adversarial Agent Behavior Learning in Autonomous Driving Using Deep Reinforcement Learning

Arjun Srinivasan, Anubhav Paras, Aniket Bera

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13560 2025-08-22 cs.CV 50%

DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary Lookup

Zhen Qu, Xian Tao, Xinyi Gong, ShiChen Qu, Xiaopei Zhang, Xingang Wang, Fei Shen, Zhengtao Zhang, Mukesh Prasad, Guiguang Ding

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Casivision Longmen Laboratory(龙门实验室) HDU UTS UCLA(加州大学洛杉矶分校) Tsinghua University(清华大学)

专题命中 其他安全 :alignment(abstract)

Comments Accepted by ICCV 2025, Project: https://github.com/xiaozhen228/DictAS

详情

展开后加载摘要…

URL PDF HTML 收藏