arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-28 至 2025-07-28 共收录 10 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 10 篇

2412.13666 2025-07-28 cs.CL cs.AI cs.CY 75%

Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation

Aneta Zugecova, Dominik Macko, Ivan Srba, Robert Moro, Jakub Kopal, Katarina Marcincinova, Matus Mesarcik

机构 * Kempelen Institute of Intelligent Technologies(凯普勒智能技术研究所) University of Copenhagen(哥本哈根大学) Comenius University in Bratislava(布拉迪斯拉瓦科เมนius大学)

专题命中 安全评测 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.CY

Comments ACL 2025 main

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Volume 1: Long Papers)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19132 2025-07-28 cs.AI cs.CL cs.CV cs.HC 62%

OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?

Xuetian Chen, Yinghao Chen, Xinfeng Yuan, Zhuo Peng, Lu Chen, Yuekeng Li, Zhoujia Zhang, Yingqian Huang, Leyan Huang, Jiaqing Liang, Tianbao Xie, Zhiyong Wu, Qiushi Sun, Biqing Qi, Bowen Zhou

机构 * Fudan University(复旦大学) Shanghai AI Lab(上海人工智能实验室) Tsinghua University(清华大学) The University of Hong Kong(香港大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18918 2025-07-28 cs.CL cs.AI 62%

Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders

Richmond Sin Jing Xuan, Jalil Huseynov, Yang Zhang

机构 * National University of Singapore(新加坡国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19455 2025-07-28 cs.LG 57%

Forest-Guided Clustering -- Shedding Light into the Random Forest Black Box

Lisa Barros de Andrade e Sousa, Gregor Miller, Ronan Le Gleut, Dominik Thalmeier, Helena Pelin, Marie Piraud

机构 * Helmholtz AI(海德堡人工智能研究所) Helmholtz Munich(海德堡慕尼黑)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19174 2025-07-28 cs.LG 57%

Automatic Cough Analysis for Non-Small Cell Lung Cancer Detection

Chiara Giangregorio, Cristina Maria Licciardello, Vanja Miskovic, Leonardo Provenzano, Alessandra Laura Giulia Pedrocchi, Andra Diana Dumitrascu, Arsela Prelaj, Marina Chiara Garassino, Emilia Ambrosini, Simona Ferrante

机构 * Department of Electronics, Information and Bioengineering, Politecnico di Milano(电子、信息与生物工程学院,米兰理工学院) Fondazione IRCCS Istituto Nazionale dei Tumori di Milano(米兰国家肿瘤研究所) Department of Medicine, Section of Hematology/Oncology, University of Chicago(医学学院,血液学/肿瘤学部门,芝加哥大学) LEARNLab, IRCCS Istituto Neurologico Carlo Besta(LEARN实验室,卡尔·贝斯塔神经病学研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Emilia Ambrosini and Simona Ferrante equally contributed to the work

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18667 2025-07-28 cs.CV cs.AI 57%

Gen-AI Police Sketches with Stable Diffusion

Nicholas Fidalgo, Aaron Contreras, Katherine Harvey, Johnny Ni

机构 * Harvard College(哈佛学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01482 2025-07-28 cs.AI 57%

Understanding LLM Scientific Reasoning through Promptings and Model's Explanation on the Answers

Alice Rueda, Mohammed S. Hassan, Argyrios Perivolaris, Bazen G. Teferra, Reza Samavi, Sirisha Rambhatla, Yuqi Wu, Yanbo Zhang, Bo Cao, Divya Sharma, Sridhar Krishnan, Venkat Bhat

机构 * University of Toronto Department of Psychiatry(多伦多大学精神病学系) Toronto Metropolitan University(多伦多 Metropolitan 大学) St. Michael’s Hospital, Unity Health Toronto(圣米歇尔医院,统一健康多伦多) Department of Electrical, Computer, and Biomedical Engineering(电气、计算机和生物医学工程系) University of Waterloo(滑铁卢大学) Department of Management Science and Engineering(管理科学与工程系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17543 2025-07-28 cs.HC 50%

Anticipate, Simulate, Reason (ASR): A Comprehensive Generative AI Framework for Combating Messaging Scams

Xue Wen Tan, Kenneth See, Stanley Kok

专题命中 安全评测 :safety(abstract)

Comments arXiv admin note: text overlap with arXiv:2412.13528

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22531 2025-07-28 cs.CV 50%

Preserve Anything: Controllable Image Synthesis with Object Preservation

Prasen Kumar Sharma, Neeraj Matiyali, Siddharth Srivastava, Gaurav Sharma

专题命中 安全评测 :alignment(abstract)

Comments Accepted at ICCV 2025 (main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10568 2025-07-28 cs.CV 50%

AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark

Aruna Gauba, Irene Pi, Yunze Man, Ziqi Pang, Vikram S. Adve, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Rice University(Rice大学) Carnegie Mellon University(卡内基梅隆大学) AIFARMS Center for Digital Agriculture at UIUC(伊利诺伊大学厄巴纳-香槟分校数字农业中心)

专题命中 安全评测 :trustworthy(abstract)

Comments Project Website: https://agmmu.github.io/ Huggingface: https://huggingface.co/datasets/AgMMU/AgMMU_v1/

详情

展开后加载摘要…

URL PDF HTML 收藏