arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-12 至 2025-08-12 共收录 16 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 16 篇

2508.07390 2025-08-12 cs.HC cs.AI 79%

Urbanite: A Dataflow-Based Framework for Human-AI Interactive Alignment in Urban Visual Analytics

Gustavo Moreira, Leonardo Ferreira, Carolina Veiga, Maryam Hosseini, Fabio Miranda

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments Accepted at IEEE VIS 2025. Urbanite is available at https://urbantk.org/urbanite

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07031 2025-08-12 eess.IV cs.AI cs.CV 79%

Trustworthy Medical Imaging with Large Language Models: A Study of Hallucinations Across Modalities

Anindya Bijoy Das, Shahnewaz Karim Sakib, Shibbir Ahmed

机构 * The University of Akron(阿克隆大学) University of Tennessee at Chattanooga(田纳西大学查塔努加分校) Texas State University(德克萨斯州立大学)

专题命中 安全评测 :trustworthy(title);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06885 2025-08-12 cs.LG 79%

Conformal Prediction and Trustworthy AI

Anthony Bellotti, Xindi Zhao

机构 * School of Computer Science, University of Nottingham Ningbo China(诺丁汉大学宁波校区计算机科学学院)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments Preprint for an essay to be published in The Importance of Being Learnable (Enhancing the Learnability and Reliability of Machine Learning Algorithms) Essays Dedicated to Alexander Gammerman on His 80th Birthday, LNCS Springer Nature Switzerland AG ed. Nguyen K.A. and Luo Z

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07560 2025-08-12 cs.RO cs.CV 78%

Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey

Yan Gong, Naibang Wang, Jianli Lu, Xinyu Zhang, Yongsheng Gao, Jie Zhao, Zifan Huang, Haozhi Bai, Nanxin Zeng, Nayu Su, Lei Yang, Ziying Song, Xiaoxi Hu, Xinmin Jiang, Xiaojuan Zhang, Susanto Rahardja

机构 * State Key Laboratory of Robotics and System(机器人系统国家重点实验室) Harbin Institute of Technology(哈尔滨工业大学) State Key Laboratory of Intelligent Green Vehicle and Mobility(智能绿色车辆与移动性国家重点实验室) Tsinghua University(清华大学) the School of Mechanical and Aerospace Engineering(机械与航空航天工程学院) Nanyang Technological University(南洋理工大学) Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence(北京交通数据挖掘与具身智能重点实验室) Beijing Jiaotong University(北京交通大学) the Institute for Infocomm Research(信息通信研究所) A*STAR the Engineering Cluster(工程集群) the Singapore Institute of Technology(新加坡理工学院)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07308 2025-08-12 cs.CL cs.AI cs.IR cs.LG 67%

HealthBranches: Synthesizing Clinically-Grounded Question Answering Datasets via Decision Pathways

Cristian Cosentino, Annamaria Defilippo, Marco Dossena, Christopher Irwin, Sara Joubbi, Pietro Liò

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07668 2025-08-12 cs.LG cs.AI 62%

AIS-LLM: A Unified Framework for Maritime Trajectory Prediction, Anomaly Detection, and Collision Risk Assessment with Explainable Forecasting

Hyobin Park, Jinwook Jung, Minseok Seo, Hyunsoo Choi, Deukjae Cho, Sekil Park, Dong-Geol Choi

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07221 2025-08-12 cs.LG cs.AI cs.MA stat.AP stat.ME 62%

LLM-based Agents for Automated Confounder Discovery and Subgroup Analysis in Causal Inference

Po-Han Lee, Yu-Cheng Lin, Chan-Tung Ku, Chan Hsu, Pei-Cing Huang, Ping-Hsun Wu, Yihuang Kang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07163 2025-08-12 cs.RO cs.AI cs.NE 57%

Integrating Neurosymbolic AI in Advanced Air Mobility: A Comprehensive Survey

Kamal Acharya, Iman Sharifi, Mehul Lad, Liang Sun, Houbing Song

机构 * Department of Information Systems, University of Maryland, Baltimore County(信息系统系,马里兰大学巴尔的摩分校) Department of Mechanical and Aerospace Engineering, The George Washington University(机械与航空航天工程系,乔治华盛顿大学) Department of Mechanical Engineering, Baylor University(机械工程系,贝勒大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 9 pages, 4 figures, IJCAI-2025 (accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17632 2025-08-12 cs.AI cs.CV cs.MM 57%

D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal Guidance

Renyang Liu, Ziyu Lyu, Wei Zhou, See-Kiong Ng

机构 * School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学信息科学与技术学院(深圳校区)) Institute of Data Science, National University of Singapore(新加坡国立大学数据科学研究所) College of Modern Engineering and the Engineering Research Center of Cyberspace, Yunnan University(云南大学现代工程学院及空天信息工程研究中心)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11569 2025-08-12 eess.IV cs.AI cs.CV 57%

Are Vision Foundation Models Ready for Out-of-the-Box Medical Image Registration?

Hanxue Gu, Yaqian Chen, Nicholas Konz, Qihang Li, Maciej A. Mazurowski

机构 * Department of Electrical and Computer Engineering, Duke University(电子工程与计算机科学系,杜克大学) Department of Biostatistics and Bioinformatics, Duke University(生物统计学与生物信息学系,杜克大学) Departments of Biostatistics and Bioinformatics, Radiology, Electrical and Computer Engineering, and Computer Science, Duke University(生物统计学与生物信息学系、放射学、电子工程与计算机科学系,杜克大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 3 figures, 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17590 2025-08-12 cs.CV cs.AI cs.RO 57%

DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving

Mihir Godbole, Xiangbo Gao, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯A&M大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 19 pages, 5 figures, Preprint under review. Code available at: https://github.com/taco-group/DRAMA-X

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07596 2025-08-12 cs.CV 50%

From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users

Shahroz Tariq, Simon S. Woo, Priyanka Singh, Irena Irmalasari, Saakshi Gupta, Dev Gupta

机构 * Sungkyunkwan University, S. Korea(顺天大学) University of Queensland, Australia(昆士兰大学)

专题命中 安全评测 :trustworthy(abstract)

Comments 11 pages, 3 tables, 5 figures, accepted for publicaiton in the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07401 2025-08-12 cs.CV 50%

LET-US: Long Event-Text Understanding of Scenes

Rui Chen, Xingyu Chen, Shaoan Wang, Shihan Kong, Junzhi Yu

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07256 2025-08-12 cs.HC 50%

Exploring Micro Accidents and Driver Responses in Automated Driving: Insights from Real-world Videos

Wei Xiang, Chuyue Zhang, Jie Yan

专题命中 安全评测 :safety(abstract)

Comments 31 pages, 5 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06916 2025-08-12 cs.CV 50%

Talk2Image: A Multi-Agent System for Multi-Turn Image Generation and Editing

Shichao Ma, Yunhe Guo, Jiahao Su, Qihe Huang, Zhengyang Zhou, Yang Wang

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06768 2025-08-12 cs.CV cs.GR 50%

DiffUS: Differentiable Ultrasound Rendering from Volumetric Imaging

Noe Bertramo, Gabriel Duguey, Vivek Gopalakrishnan

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 安全评测 :alignment(abstract)

Comments 10 pages, accepted to MICCAI ASMUS 25

详情

展开后加载摘要…

URL PDF HTML 收藏