arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-31 至 2025-07-31 共收录 8 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 8 篇

2403.16591 2025-07-31 cs.LG cs.AI cs.CR 81%

Bridging Privacy and Robustness for Trustworthy Machine Learning

Xiaojin Zhang, Wei Chen

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22389 2025-07-31 cs.RO cs.SY eess.SY 78%

Safety Evaluation of Motion Plans Using Trajectory Predictors as Forward Reachable Set Estimators

Kaustav Chakraborty, Zeyuan Feng, Sushant Veer, Apoorva Sharma, Wenhao Ding, Sever Topan, Boris Ivanovic, Marco Pavone, Somil Bansal

机构 * Department of Electrical Engineering, University of Southern California(电气工程系,美国南加州大学) Department of Aeronautics and Astronautics, Stanford University(航空与宇航系,斯坦福大学) NVIDIA Research(NVIDIA研究)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21919 2025-07-31 cs.CL cs.AI cs.CY 67%

Training language models to be warm and empathetic makes them less reliable and more sycophantic

Lujain Ibrahim, Franziska Sofia Hafner, Luc Rocher

机构 * Oxford Internet Institute(牛津互联网研究所) University of Oxford(牛津大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22576 2025-07-31 cs.CV cs.AI cs.LG 62%

COOkeD: Ensemble-based OOD detection in the era of zero-shot CLIP

Galadrielle Humblot-Renaux, Gianni Franchi, Sergio Escalera, Thomas B. Moeslund

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments accepted at ICCVW'25 - Systematic Trust in AI Models: Ensuring Fairness, Reliability, Explainability, and Accountability in Machine Learning Frameworks

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01282 2025-07-31 cs.CL 57%

Prompt-Reverse Inconsistency: LLM Self-Inconsistency Beyond Generative Randomness and Prompt Paraphrasing

Jihyun Janice Ahn, Wenpeng Yin

机构 * Department of Computer Science & Engineering(计算机科学与工程系) The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments accepted in COLM2025, 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18044 2025-07-31 cs.LG 57%

The Geometry of Queries: Query-Based Innovations in Retrieval-Augmented Generation for Healthcare QA

Eric Yang, Jonathan Amar, Jong Ha Lee, Bhawesh Kumar, Yugang Jia

机构 * Verily Life Sciences(Verily 生物科技)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22076 2025-07-31 cs.LG 57%

Test-time Prompt Refinement for Text-to-Image Models

Mohammad Abdul Hafeez Khan, Yash Jain, Siddhartha Bhattacharyya, Vibhav Vineet

机构 * Florida Institute of Technology(佛罗里达理工学院) Microsoft Research(微软研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Accepted to ICCV 2025, MARS2 Workshop. Total 14 pages, 12 figures and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22100 2025-07-31 cs.CV 50%

Trade-offs in Image Generation: How Do Different Dimensions Interact?

Sicheng Zhang, Binzhu Xie, Zhonghao Yan, Yuli Zhang, Donghao Zhou, Xiaofei Chen, Shi Qiu, Jiaqi Liu, Guoyang Xie, Zhichao Lu

机构 * Khalifa University(卡利法大学) The Chinese University of Hong Kong(香港中文大学) Queen Mary University of London(伦敦大学玛丽女王学院) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) City University of Hong Kong(香港城市大学)

专题命中 安全评测 :alignment(abstract)

Comments Accepted in ICCV 2025, Codebase: https://github.com/fesvhtr/TRIG

详情

展开后加载摘要…

URL PDF HTML 收藏