arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-28 至 2025-07-28 共收录 19 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 2 篇

2502.05773 2025-07-28 cs.LG cs.AI stat.ML 84%

PIPA: Preference Alignment as Prior-Informed Statistical Estimation

Junbo Li, Zhangyang Wang, Qiang Liu

机构 * The University of Texas at Austin, US(德克萨斯大学奥斯汀分校)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18802 2025-07-28 cs.HC cs.AI 83%

DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition

Danqing Shi, Furui Cheng, Tino Weinkauf, Antti Oulasvirta, Mennatallah El-Assady

机构 * Aalto University(阿alto大学) ETH Zürich(苏黎世联邦理工学院) KTH Royal Institute of Technology(皇家理工学院)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 越狱攻击 2 篇

2507.19227 2025-07-28 cs.CL 83%

Jailbreaking Large Language Diffusion Models: Revealing Hidden Safety Flaws in Diffusion-Based Text Generation

Yuanhe Zhang, Fangzhou Xie, Zhenhong Zhou, Zherui Li, Hao Chen, Kun Wang, Yufei Guo

专题命中 越狱攻击 :safety(title,abstract);jailbreak(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18656 2025-07-28 cs.CV cs.LG 57%

ShrinkBox: Backdoor Attack on Object Detection to Disrupt Collision Avoidance in Machine Learning-based Advanced Driver Assistance Systems

Muhammad Zaeem Shahzad, Muhammad Abdullah Hanif, Bassem Ouni, Muhammad Shafique

机构 * eBRAIN Lab, New York University Abu Dhabi (NYUAD), UAE(eBRAIN实验室,纽约大学阿布扎赫尔分校(NYUAD),阿联酋) AI and Digital Science Research Center, Technology Innovation Institute (TII), Abu Dhabi, UAE(人工智能与数字科学研究中心,技术创新研究所(TII),阿布扎赫尔,阿联酋)

专题命中 越狱攻击 :safety(abstract);分类 cs.LG

Comments 8 pages, 8 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 隐私与版权 1 篇

2502.20719 2025-07-28 cs.LG cs.AI 62%

Generating Clinically Realistic EHR Data via a Hierarchy- and Semantics-Guided Transformer

Guanglin Zhou, Sebastiano Barbieri

机构 * University of Queensland(昆士兰大学)

专题命中 隐私与版权 :alignment(abstract);分类 cs.AI、cs.LG

Comments The camera ready version for ECAI-2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 安全评测 10 篇

2412.13666 2025-07-28 cs.CL cs.AI cs.CY 75%

Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation

Aneta Zugecova, Dominik Macko, Ivan Srba, Robert Moro, Jakub Kopal, Katarina Marcincinova, Matus Mesarcik

机构 * Kempelen Institute of Intelligent Technologies(凯普勒智能技术研究所) University of Copenhagen(哥本哈根大学) Comenius University in Bratislava(布拉迪斯拉瓦科เมนius大学)

专题命中 安全评测 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.CY

Comments ACL 2025 main

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Volume 1: Long Papers)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19132 2025-07-28 cs.AI cs.CL cs.CV cs.HC 62%

OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?

Xuetian Chen, Yinghao Chen, Xinfeng Yuan, Zhuo Peng, Lu Chen, Yuekeng Li, Zhoujia Zhang, Yingqian Huang, Leyan Huang, Jiaqing Liang, Tianbao Xie, Zhiyong Wu, Qiushi Sun, Biqing Qi, Bowen Zhou

机构 * Fudan University(复旦大学) Shanghai AI Lab(上海人工智能实验室) Tsinghua University(清华大学) The University of Hong Kong(香港大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18918 2025-07-28 cs.CL cs.AI 62%

Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders

Richmond Sin Jing Xuan, Jalil Huseynov, Yang Zhang

机构 * National University of Singapore(新加坡国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19455 2025-07-28 cs.LG 57%

Forest-Guided Clustering -- Shedding Light into the Random Forest Black Box

Lisa Barros de Andrade e Sousa, Gregor Miller, Ronan Le Gleut, Dominik Thalmeier, Helena Pelin, Marie Piraud

机构 * Helmholtz AI(海德堡人工智能研究所) Helmholtz Munich(海德堡慕尼黑)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19174 2025-07-28 cs.LG 57%

Automatic Cough Analysis for Non-Small Cell Lung Cancer Detection

Chiara Giangregorio, Cristina Maria Licciardello, Vanja Miskovic, Leonardo Provenzano, Alessandra Laura Giulia Pedrocchi, Andra Diana Dumitrascu, Arsela Prelaj, Marina Chiara Garassino, Emilia Ambrosini, Simona Ferrante

机构 * Department of Electronics, Information and Bioengineering, Politecnico di Milano(电子、信息与生物工程学院,米兰理工学院) Fondazione IRCCS Istituto Nazionale dei Tumori di Milano(米兰国家肿瘤研究所) Department of Medicine, Section of Hematology/Oncology, University of Chicago(医学学院,血液学/肿瘤学部门,芝加哥大学) LEARNLab, IRCCS Istituto Neurologico Carlo Besta(LEARN实验室,卡尔·贝斯塔神经病学研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Emilia Ambrosini and Simona Ferrante equally contributed to the work

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18667 2025-07-28 cs.CV cs.AI 57%

Gen-AI Police Sketches with Stable Diffusion

Nicholas Fidalgo, Aaron Contreras, Katherine Harvey, Johnny Ni

机构 * Harvard College(哈佛学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01482 2025-07-28 cs.AI 57%

Understanding LLM Scientific Reasoning through Promptings and Model's Explanation on the Answers

Alice Rueda, Mohammed S. Hassan, Argyrios Perivolaris, Bazen G. Teferra, Reza Samavi, Sirisha Rambhatla, Yuqi Wu, Yanbo Zhang, Bo Cao, Divya Sharma, Sridhar Krishnan, Venkat Bhat

机构 * University of Toronto Department of Psychiatry(多伦多大学精神病学系) Toronto Metropolitan University(多伦多 Metropolitan 大学) St. Michael’s Hospital, Unity Health Toronto(圣米歇尔医院,统一健康多伦多) Department of Electrical, Computer, and Biomedical Engineering(电气、计算机和生物医学工程系) University of Waterloo(滑铁卢大学) Department of Management Science and Engineering(管理科学与工程系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17543 2025-07-28 cs.HC 50%

Anticipate, Simulate, Reason (ASR): A Comprehensive Generative AI Framework for Combating Messaging Scams

Xue Wen Tan, Kenneth See, Stanley Kok

专题命中 安全评测 :safety(abstract)

Comments arXiv admin note: text overlap with arXiv:2412.13528

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22531 2025-07-28 cs.CV 50%

Preserve Anything: Controllable Image Synthesis with Object Preservation

Prasen Kumar Sharma, Neeraj Matiyali, Siddharth Srivastava, Gaurav Sharma

专题命中 安全评测 :alignment(abstract)

Comments Accepted at ICCV 2025 (main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10568 2025-07-28 cs.CV 50%

AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark

Aruna Gauba, Irene Pi, Yunze Man, Ziqi Pang, Vikram S. Adve, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Rice University(Rice大学) Carnegie Mellon University(卡内基梅隆大学) AIFARMS Center for Digital Agriculture at UIUC(伊利诺伊大学厄巴纳-香槟分校数字农业中心)

专题命中 安全评测 :trustworthy(abstract)

Comments Project Website: https://agmmu.github.io/ Huggingface: https://huggingface.co/datasets/AgMMU/AgMMU_v1/

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 其他安全 4 篇

2507.18631 2025-07-28 cs.CR 88%

Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment

Hao Li, Lijun Li, Zhenghao Lu, Xianyi Wei, Rui Li, Jing Shao, Lei Sha

专题命中 其他安全 :alignment(title,abstract);safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07003 2025-07-28 cs.AI cs.LG cs.RO 62%

RACER: Rational Artificial Intelligence Car-following-model Enhanced by Reality

Tianyi Li, Alexander Halatsis, Raphael Stern

机构 * Department of Civil Engineering(土木工程系) Department of Aerospace Engineering(航空航天工程系) Department of Civil, Environmental, and Geo-Engineering(土木、环境与地球工程系)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Journal ref "RACER: Rational Artificial Intelligence Car-Following-Model Enhanced by Reality," in IEEE Transactions on Intelligent Transportation Systems,

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07919 2025-07-28 cs.CL q-bio.BM 57%

Advancing biomolecular understanding and design following human instructions

Xiang Zhuang, Keyan Ding, Tianwen Lyu, Yinuo Jiang, Xiaotong Li, Zhuoyi Xiang, Zeyuan Wang, Ming Qin, Kehua Feng, Jike Wang, Qiang Zhang, Huajun Chen

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Journal ref Nature Machine Intelligence volume 7, pages1154-1167 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18750 2025-07-28 cs.MM cs.SD eess.AS 50%

CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation

Hyunwoo Oh, SeungJu Cha, Kwanyoung Lee, Si-Woo Kim, Dong-Jin Kim

机构 * Hanyang University(翰阳大学)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏