arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-12 至 2025-08-12 共收录 77 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 16 篇

2508.07308 2025-08-12 cs.CL cs.AI cs.IR cs.LG 67%

HealthBranches: Synthesizing Clinically-Grounded Question Answering Datasets via Decision Pathways

Cristian Cosentino, Annamaria Defilippo, Marco Dossena, Christopher Irwin, Sara Joubbi, Pietro Liò

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07668 2025-08-12 cs.LG cs.AI 62%

AIS-LLM: A Unified Framework for Maritime Trajectory Prediction, Anomaly Detection, and Collision Risk Assessment with Explainable Forecasting

Hyobin Park, Jinwook Jung, Minseok Seo, Hyunsoo Choi, Deukjae Cho, Sekil Park, Dong-Geol Choi

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07221 2025-08-12 cs.LG cs.AI cs.MA stat.AP stat.ME 62%

LLM-based Agents for Automated Confounder Discovery and Subgroup Analysis in Causal Inference

Po-Han Lee, Yu-Cheng Lin, Chan-Tung Ku, Chan Hsu, Pei-Cing Huang, Ping-Hsun Wu, Yihuang Kang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07163 2025-08-12 cs.RO cs.AI cs.NE 57%

Integrating Neurosymbolic AI in Advanced Air Mobility: A Comprehensive Survey

Kamal Acharya, Iman Sharifi, Mehul Lad, Liang Sun, Houbing Song

机构 * Department of Information Systems, University of Maryland, Baltimore County(信息系统系,马里兰大学巴尔的摩分校) Department of Mechanical and Aerospace Engineering, The George Washington University(机械与航空航天工程系,乔治华盛顿大学) Department of Mechanical Engineering, Baylor University(机械工程系,贝勒大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 9 pages, 4 figures, IJCAI-2025 (accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17632 2025-08-12 cs.AI cs.CV cs.MM 57%

D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal Guidance

Renyang Liu, Ziyu Lyu, Wei Zhou, See-Kiong Ng

机构 * School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学信息科学与技术学院(深圳校区)) Institute of Data Science, National University of Singapore(新加坡国立大学数据科学研究所) College of Modern Engineering and the Engineering Research Center of Cyberspace, Yunnan University(云南大学现代工程学院及空天信息工程研究中心)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11569 2025-08-12 eess.IV cs.AI cs.CV 57%

Are Vision Foundation Models Ready for Out-of-the-Box Medical Image Registration?

Hanxue Gu, Yaqian Chen, Nicholas Konz, Qihang Li, Maciej A. Mazurowski

机构 * Department of Electrical and Computer Engineering, Duke University(电子工程与计算机科学系,杜克大学) Department of Biostatistics and Bioinformatics, Duke University(生物统计学与生物信息学系,杜克大学) Departments of Biostatistics and Bioinformatics, Radiology, Electrical and Computer Engineering, and Computer Science, Duke University(生物统计学与生物信息学系、放射学、电子工程与计算机科学系,杜克大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 3 figures, 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17590 2025-08-12 cs.CV cs.AI cs.RO 57%

DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving

Mihir Godbole, Xiangbo Gao, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯A&M大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 19 pages, 5 figures, Preprint under review. Code available at: https://github.com/taco-group/DRAMA-X

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07596 2025-08-12 cs.CV 50%

From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users

Shahroz Tariq, Simon S. Woo, Priyanka Singh, Irena Irmalasari, Saakshi Gupta, Dev Gupta

机构 * Sungkyunkwan University, S. Korea(顺天大学) University of Queensland, Australia(昆士兰大学)

专题命中 安全评测 :trustworthy(abstract)

Comments 11 pages, 3 tables, 5 figures, accepted for publicaiton in the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07401 2025-08-12 cs.CV 50%

LET-US: Long Event-Text Understanding of Scenes

Rui Chen, Xingyu Chen, Shaoan Wang, Shihan Kong, Junzhi Yu

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07256 2025-08-12 cs.HC 50%

Exploring Micro Accidents and Driver Responses in Automated Driving: Insights from Real-world Videos

Wei Xiang, Chuyue Zhang, Jie Yan

专题命中 安全评测 :safety(abstract)

Comments 31 pages, 5 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06916 2025-08-12 cs.CV 50%

Talk2Image: A Multi-Agent System for Multi-Turn Image Generation and Editing

Shichao Ma, Yunhe Guo, Jiahao Su, Qihe Huang, Zhengyang Zhou, Yang Wang

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06768 2025-08-12 cs.CV cs.GR 50%

DiffUS: Differentiable Ultrasound Rendering from Volumetric Imaging

Noe Bertramo, Gabriel Duguey, Vivek Gopalakrishnan

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 安全评测 :alignment(abstract)

Comments 10 pages, accepted to MICCAI ASMUS 25

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 6 篇

2508.06849 2025-08-12 cs.CY cs.AI cs.HC 73%

Towards Experience-Centered AI: A Framework for Integrating Lived Experience in Design and Development

Sanjana Gautam, Mohit Chandra, Ankolika De, Tatiana Chakravorti, Girik Malik, Munmun De Choudhury

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07284 2025-08-12 cs.CL cs.AI cs.CY 67%

"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas

Junchen Ding, Penghao Jiang, Zihao Xu, Ziqi Ding, Yichen Zhu, Jiaojiao Jiang, Yuekang Li

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16663 2025-08-12 cs.LG cs.AI cs.CY cs.LO cs.SE 67%

Runtime Monitoring and Enforcement of Conditional Fairness in Generative AIs

Chih-Hong Cheng, Changshun Wu, Xingyu Zhao, Saddek Bensalem, Harald Ruess

机构 * Chalmers University of Technology, Sweden Carl von Ossietzky Universität Oldenburg, Germany Universit\'e Grenoble Alpes, France University of Warwick, United Kingdom CSX-AI, France SRI International, United States

专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07673 2025-08-12 cs.AI cs.LG 62%

Ethics2vec: aligning automatic agents and human preferences

Gianluca Bontempi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07111 2025-08-12 cs.CL cs.AI 62%

Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution

Falaah Arif Khan, Nivedha Sivakumar, Yinong Oliver Wang, Katherine Metcalf, Cezanne Camacho, Barry-John Theobald, Luca Zappella, Nicholas Apostoloff

机构 * Apple(苹果公司)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16170 2025-08-12 cs.AI 57%

Learning How to Vote with Principles: Axiomatic Insights Into the Collective Decisions of Neural Networks

Levin Hornischer, Zoi Terzopoulou

机构 * Munich Center for Mathematical Philosophy, LMU Munich Munich Germany GATE, CNRS, Universit\'e Jean Monnet, Universit\'e Lumiere Lyon 2 Saint-Etienne France Munich Center for Mathematical Philosophy, LMU Munich GATE, CNRS, Universit\'e Jean Monnet, Universit\'e Lumiere Lyon 2

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 44 pages, 21 figures, 14 tables. Updated and published version

Journal ref Journal of Artificial Intelligence Research 83, Article 25 (August 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 29 篇

2508.08131 2025-08-12 cs.CL cs.AI 81%

Optimal Transport Regularization for Speech Text Alignment in Spoken Language Models

Wenze Xu, Chun Wang, Jiazhen Yu, Sheng Chen, Liang Gao, Weihong Deng

机构 * Mashang Consumer Finance Co., Ltd.(Mashang消费金融有限公司) The University of Sydney(悉尼大学) Macau University of Science and Technology(澳门科学技术大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments To be presented at ACPR 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06592 2025-08-12 cs.CY cs.AI 81%

Towards Integrated Alignment

Ben Y. Reis, William La Cava

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06895 2025-08-12 cs.CV cs.AI 79%

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models

Jianting Tang, Yubo Wang, Haoyu Cao, Linli Xu

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21298 2025-08-12 cs.SD cs.AI cs.CL cs.LG cs.MM eess.AS 67%

Exploring Adapter Design Tradeoffs for Low Resource Music Generation

Atharva Mehta, Shivam Chauhan, Monojit Choudhury

机构 * Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14701 2025-08-12 cs.CL cs.AI cs.HC cs.LG q-bio.NC 67%

COMPASS: Computational Mapping of Patient-Therapist Alliance Strategies with Language Modeling

Baihan Lin, Djallel Bouneffouf, Yulia Landa, Rachel Jespersen, Cheryl Corcoran, Guillermo Cecchi

机构 * Department of Artificial Intelligence and Human Health, Icahn School of Medicine at Mount Sinai(人工智能与人类健康系,伊坎医学院 Mount Sinai 分校) Department of Psychiatry, Icahn School of Medicine at Mount Sinai(精神病学系,伊坎医学院 Mount Sinai 分校) Department of Neuroscience, Icahn School of Medicine at Mount Sinai(神经科学系,伊坎医学院 Mount Sinai 分校) Berkman Klein Center for Internet & Society, Harvard University(互联网与社会研究中心,哈佛大学) IBM Research, T.J. Watson Research Center(IBM 研究,T.J. Watson 研究中心) Mental Illness Research, Education and Clinical Center, James J. Peters VA Medical Center(精神疾病研究、教育与临床中心,James J. Peters VA 医疗中心)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Translational Psychiatry, in press. This work extends our research series in computational psychiatry (e.g auto annotation in arXiv:2204.05522, topic extraction in arXiv:2204.10189, and diagnosis in arXiv:2210.15603) with the introduction of LLMs to complete the full cycle of interpreting and understanding psychotherapy strategies as a comprehensive analytical framework

Journal ref Transl Psychiatry 15, 166 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07484 2025-08-12 cs.CL cs.AI 62%

ALOPE: Adaptive Layer Optimization for Translation Quality Estimation using Large Language Models

Archchana Sindhujan, Shenbin Qian, Chan Chi Chun Matthew, Constantin Orasan, Diptesh Kanojia

机构 * Institute for People-Centred AI and Centre for Translation Studies, School of Computer Science and Electronic Engineering, University of Surrey(以人为本的人工智能研究所和翻译研究中心,计算机科学与电子工程学院,萨里大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to COLM 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22919 2025-08-12 cs.CL cs.AI 62%

A novel language model for predicting serious adverse event results in clinical trials from their prospective registrations

Qixuan Hu, Xumou Zhang, Jinman Kim, Florence Bourgeois, Adam G. Dunn

机构 * School of Computer Science, Faculty of Engineering, University of Sydney(悉尼大学计算机科学学院、工程学院) Computational Health Informatics Program, Boston Children’s Hospital(波士顿儿童医院计算健康信息学项目) Harvard-MIT Center for Regulatory Science and Department of Pediatrics, Harvard Medical School(哈佛-麻省理工监管科学中心和哈佛医学院儿科部门) Sydney School of Public Health, Faculty of Medicine and Health, University of Sydney(悉尼大学公共卫生学院、医学与健康学院)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments 12 pages, 4 figures. Updated to include Table 2, Supplementary Table 1, and an additional baseline random forest model

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21931 2025-08-12 cs.IR cs.AI cs.CL cs.MA 62%

ARAG: Agentic Retrieval Augmented Generation for Personalized Recommendation

Reza Yousefi Maragheh, Pratheek Vadla, Priyank Gupta, Kai Zhao, Aysenur Inan, Kehui Yao, Jianpeng Xu, Praveen Kanumala, Jason Cho, Sushant Kumar

机构 * Walmart Global Tech(沃尔玛全球科技)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07480 2025-08-12 eess.SP cs.AI cs.LG 62%

EEG-Language Pretraining for Highly Label-Efficient Clinical Phenotyping

Sam Gijsen, Kerstin Ritter

机构 * Charité – Universitätsmedizin Berlin, Department of Psychiatry and Psychotherapy, Berlin, Germany(柏林查理医院医学大学精神病与心理治疗系) Hertie Institute for AI in Brain Health, University of Tübingen, Germany(图宾根大学健康人工智能研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01196 2025-08-12 cs.CL cs.AI 62%

$μ$KE: Matryoshka Unstructured Knowledge Editing of Large Language Models

Zian Su, Ziyang Huang, Kaiyuan Zhang, Xiangyu Zhang

机构 * Purdue University(普渡大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments COLM 2025. The first two authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20758 2025-08-12 stat.AP cs.AI cs.CL 62%

Collective Reasoning Among LLMs: A Framework for Answer Validation Without Ground Truth

Seyed Pouyan Mousavi Davoudi, Amin Gholami Davodi, Alireza Amiri-Margavi, Alireza Shafiee Fard, Mahdi Jafari

机构 * Independent Researcher in AI and Statistics(人工智能与统计学独立研究者) Shahrood University of Technology(沙霍罗德大学) University of Pittsburgh(匹兹堡大学) Duquesne University(杜克森大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 6pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17797 2025-08-12 cs.AI cs.GT cs.LG cs.MA 62%

Observation Interference in Partially Observable Assistance Games

Scott Emmons, Caspar Oesterheld, Vincent Conitzer, Stuart Russell

机构 * Center for Human-Compatible AI, University of California, Berkeley(人类兼容人工智能中心,加州大学伯克利分校) Foundations of Cooperative AI Lab, Carnegie Mellon University(协作人工智能实验室,卡内基梅隆大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏