arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2503.12435 2025-03-18 cs.IT cs.LG math.IT 57%

XAI-Driven Client Selection for Federated Learning in Scalable 6G Network Slicing

Martino Chiarani, Swastika Roy, Christos Verikoukis, Fabrizio Granelli

机构 * University of Trento(特伦托大学) Iquadrat Informatica S.L.(伊夸德拉特信息有限公司) University of Patras(帕特雷大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 8 pages, 7 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12687 2025-03-18 cs.AI 57%

AI Agents: Evolution, Architecture, and Real-World Applications

Naveen Krishnan

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 52 pages, 4 figures, comprehensive survey and analysis of AI agent evolution, architecture, evaluation frameworks, and applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12018 2025-03-18 cs.CV cs.AI 57%

Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art

Zhe Jin, Tat-Seng Chua

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10702 2025-03-17 cs.CL cs.IR 57%

ClaimTrust: Propagation Trust Scoring for RAG Systems

Hangkai Qian, Bo Li, Qichen Wang

机构 * UIUC(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments 6 pages, 2 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10382 2025-03-14 cs.LG 57%

Subgroup Performance Analysis in Hidden Stratifications

Alceu Bissoto, Trung-Dung Hoang, Tim Flühmann, Susu Sun, Christian F. Baumgartner, Lisa M. Koch

机构 * Inselspital, Bern University Hospital, University of Bern(伯尔尼大学医院因塞尔spital(伯尔尼大学医院)) Diabetes Center Berne(伯尔尼糖尿病中心) Cluster of Excellence: Machine Learning - New Perspectives for Science, University of Tübingen(蒂宾根大学卓越集群:机器学习——科学新视角) Faculty of Health Sciences and Medicine, University of Lucerne(卢塞恩大学健康科学与医学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13163 2025-03-14 cs.LG 57%

Unlocking Historical Clinical Trial Data with ALIGN: A Compositional Large Language Model System for Medical Coding

Nabeel Seedat, Caterina Tozzi, Andrea Hita Ardiaca, Mihaela van der Schaar, James Weatherall, Adam Taylor

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00263 2025-03-14 cs.CV cs.AI 57%

Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation

Kun Yuan, Vinkle Srivastav, Nassir Navab, Nicolas Padoy

机构 * University of Strasbourg(斯特拉斯堡大学) CNRS(法国国家科学研究中心) INSERM(法国国家健康与医学研究院) IHU Strasbourg(斯特拉斯堡大学医院研究所) Technische Universität München(慕尼黑工业大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted at the 38th Conference on Neural Information Processing Systems (NeurIPS 2024 Spolight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09964 2025-03-14 cs.CR cs.CL 57%

ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Content

Bhavik Chandna, Mariam Aboujenane, Usman Naseem

机构 * UC San Diego(加州大学圣地亚哥分校) Euromed University of Fez(非斯欧罗梅德大学) Macquarie University(麦考瑞大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09219 2025-03-13 cs.CL 57%

Rethinking Prompt-based Debiasing in Large Language Models

Xinyi Yang, Runzhe Zhan, Derek F. Wong, Shu Yang, Junchao Wu, Lidia S. Chao

机构 * University of Macau(澳门大学) KAUST(阿卜杜拉国王科技大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17202 2025-03-13 cs.SD cs.CL eess.AS 57%

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators

Chen Chen, Yuchen Hu, Siyin Wang, Helin Wang, Zhehuai Chen, Chao Zhang, Chao-Han Huck Yang, Eng Siong Chng

机构 * Nanyang Technological University(南洋理工大学) NVIDIA(英伟达) Tsinghua University(清华大学) Johns Hopkins University(约翰斯·霍普金斯大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15604 2025-03-13 cs.HC cs.AI 57%

Investigating Use Cases of AI-Powered Scene Description Applications for Blind and Low Vision People

Ricardo Gonzalez, Jazmin Collins, Shiri Azenkot, Cynthia Bennett

机构 * Cornell Tech(康奈尔科技学院) Google(谷歌公司)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 21 pages, 18 figures, 5 tables, main track CHI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08683 2025-03-12 cs.CV cs.AI cs.MA 57%

CoLMDriver: LLM-based Negotiation Benefits Cooperative Autonomous Driving

Changxing Liu, Genjia Liu, Zijun Wang, Jinchang Yang, Siheng Chen

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07078 2025-03-11 cs.CL eess.AS 57%

Linguistic Knowledge Transfer Learning for Speech Enhancement

Kuo-Hsuan Hung, Xugang Lu, Szu-Wei Fu, Huan-Hsin Tseng, Hsin-Yi Lin, Chii-Wann Lin, Yu Tsao

机构 * National Taiwan University(台湾大学) National Institute of Information and Communications Technology(日本信息通信研究机构) NVIDIA(英伟达) Brookhaven National Laboratory(布鲁克海文国家实验室) Seton Hall University(薛顿贺尔大学) Academia Sinica(中央研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07003 2025-03-11 cs.CL 57%

Large Language Models Often Say One Thing and Do Another

Ruoxi Xu, Hongyu Lin, Xianpei Han, Jia Zheng, Weixiang Zhou, Le Sun, Yingfei Sun

机构 * University of Chinese Academy of Sciences(中国科学院大学) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) State Key Laboratory of Computer Science, Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所计算机科学国家重点实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Published on ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14833 2025-03-11 cs.LG 57%

Probabilistic Robustness in Deep Learning: A Concise yet Comprehensive Guide

Xingyu Zhao

机构 * WMG, University of Warwick(华威大学WMG学院)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments This is a preprint of the following chapter: X. Zhao, Probabilistic Robustness in Deep Learning: A Concise yet Comprehensive Guide, published in the book Adversarial Example Detection and Mitigation Using Machine Learning, edited by Ehsan Nowroozi, Rahim Taheri, Lucas Cordeiro, 2025, Springer Nature. The final authenticated version will available online soon

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00178 2025-03-11 cs.CV cs.AI eess.IV 57%

Clinical Evaluation of Medical Image Synthesis: A Case Study in Wireless Capsule Endoscopy

Panagiota Gatoula, Dimitrios E. Diamantis, Anastasios Koulaouzidis, Cristina Carretero, Stefania Chetcuti-Zammit, Pablo Cortegoso Valdivia, Begoña González-Suárez, Alessandro Mussetto, John Plevris, Alexander Robertson, Bruno Rosa, Ervin Toth, Dimitris K. Iakovidis

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments This work has been submitted for possible journal publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12827 2025-03-11 cs.CL 57%

An Evaluation Benchmark for Adverse Drug Event Prediction from Clinical Trial Results

Anthony Yazdani, Alban Bornet, Philipp Khlebnikov, Boya Zhang, Hossein Rouhizadeh, Poorya Amini, Douglas Teodoro

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05899 2025-03-11 cs.HC cs.AI 57%

Towards Understanding the Use of MLLM-Enabled Applications for Visual Interpretation by Blind and Low Vision People

Ricardo E. Gonzalez Penuela, Ruiying Hu, Sharon Lin, Tanisha Shende, Shiri Azenkot

机构 * Cornell University(康奈尔大学) Cornell Tech(康奈尔科技学院) Oberlin College(欧柏林学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 8 pages, 1 figure, 4 tables, to appear at CHI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05789 2025-03-11 cs.LG 57%

EXALT: EXplainable ALgorithmic Tools for Optimization Problems

Zuzanna Bączek, Michał Bizoń, Aneta Pawelec, Piotr Sankowski

机构 * University of Warsaw(华沙大学) Ideas-NCBR Lodz University of Technology(罗兹工业大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05775 2025-03-11 cs.LG stat.ML 57%

Evaluation of Missing Data Imputation for Time Series Without Ground Truth

Rania Farjallah, Bassant Selim, Brigitte Jaumard, Samr Ali, Georges Kaddoum

机构 * École de Technologie Supérieure(高等技术学院) Concordia University(康考迪亚大学) Ericsson(爱立信) Lebanese American University(黎巴嫩美国大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Accepted for publication in IEEE ICC 2025 (International Conference on Communications). The paper consists of 6 pages including references and contains 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00203 2025-03-11 cs.CL 57%

Llamarine: Open-source Maritime Industry-specific Large Language Model

William Nguyen, An Phan, Konobu Kimura, Hitoshi Maeno, Mika Tanaka, Quynh Le, William Poucher, Christopher Nguyen

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05242 2025-03-10 cs.CL 57%

MM-StoryAgent: Immersive Narrated Storybook Video Generation with a Multi-Agent Paradigm across Text, Image and Audio

Xuenan Xu, Jiahao Mei, Chenliang Li, Yuning Wu, Ming Yan, Shaopeng Lai, Ji Zhang, Mengyue Wu

机构 * Shanghai Jiao Tong University(上海交通大学) Alibaba Group(阿里巴巴集团) East China Normal University(华东师范大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14677 2025-03-10 cs.CL 57%

Are AI Detectors Good Enough? A Survey on Quality of Datasets With Machine-Generated Texts

German Gritsai, Anastasia Voznyuk, Andrey Grabovoy, Yury Chekhovich

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments Presented at Preventing and Detecting LLM Misinformation (PDLM) at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00965 2025-03-10 cs.RO cs.AI 57%

HBTP: Heuristic Behavior Tree Planning with Large Language Model Reasoning

Yishuai Cai, Xinglin Chen, Yunxin Mao, Minglong Li, Shaowu Yang, Wenjing Yang, Ji Wang

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09452 2025-03-10 cs.CY 57%

Keep the Future Human: Why and How We Should Close the Gates to AGI and Superintelligence, and What We Should Build Instead

Anthony Aguirre

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY

Comments 62 pages, 2 figures, 3 appendices. This is a total rewrite and major expansion of a previous version entitled "Close the Gates."

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13898 2025-03-07 cs.LG 57%

Cross-Modal Prototype based Multimodal Federated Learning under Severely Missing Modality

Huy Q. Le, Chu Myaet Thwal, Yu Qiao, Ye Lin Tun, Minh N. H. Nguyen, Eui-Nam Huh, Choong Seon Hong

机构 * Kyung Hee University(庆熙大学) The University of Danang—Vietnam-Korea University of Information and Communication Technology(岘港大学—越韩信息与通信技术大学) Digital Science and Technology Institute(数字科学技术研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 14 pages, 8 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02675 2025-03-05 cs.CV cs.AI 57%

State of play and future directions in industrial computer vision AI standards

Artemis Stefanidou, Panagiotis Radoglou-Grammatikis, Vasileios Argyriou, Panagiotis Sarigiannidis, Iraklis Varlamis, Georgios Th. Papadopoulos

机构 * Harokopio University of Athens(雅典哈罗科皮奥大学) K3Y Ltd.(K3Y有限公司) Kingston University(金斯顿大学) University of Western Macedonia(西马其顿大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02476 2025-03-05 cs.CV cs.AI 57%

BioD2C: A Dual-level Semantic Consistency Constraint Framework for Biomedical VQA

Zhengyang Ji, Shang Gao, Li Liu, Yifan Jia, Yutao Yue

机构 * Shandong University(山东大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Institute of Deep Perception Technology, JITRI(JITRI深度感知技术研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09913 2025-03-05 cs.AI cs.HC 57%

AutoS$^2$earch: Unlocking the Reasoning Potential of Large Models for Web-based Source Search

Zhengqiu Zhu, Yatai Ji, Jiaheng Huang, Yong Zhao, Sihang Qiu, Rusheng Ju

机构 * National University of Defense Technology(国防科技大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01262 2025-03-05 cs.CL cs.IR 57%

RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework

Kunlun Zhu, Yifan Luo, Dingling Xu, Yukun Yan, Zhenghao Liu, Shi Yu, Ruobing Wang, Shuo Wang, Yishan Li, Nan Zhang, Xu Han, Zhiyuan Liu, Maosong Sun

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments https://github.com/OpenBMB/RAGEval

详情

展开后加载摘要…

URL PDF HTML 收藏