arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2502.06193 2025-04-22 cs.SE cs.AI 57%

Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering

Ruiqi Wang, Jiyu Guo, Cuiyun Gao, Guodong Fan, Chun Yong Chong, Xin Xia

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Monash University Malaysia(墨尔本大学马来西亚分校) Zhejiang University(浙江大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted by ISSTA 2025: https://conf.researchr.org/details/issta-2025/issta-2025-papers/85/Can-LLMs-replace-Human-Evaluators-An-Empirical-Study-of-LLM-as-a-Judge-in-Software-E

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14395 2025-04-22 cs.CV cs.AI cs.MA 57%

Hydra: An Agentic Reasoning Approach for Enhancing Adversarial Robustness and Mitigating Hallucinations in Vision-Language Models

Chung-En, Yu, Hsuan-Chih, Chen, Brian Jalaian, Nathaniel D. Bastian

机构 * University of West Florida(西弗吉尼亚大学) New York University(纽约大学) United States Military Academy(美国军事学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14280 2025-04-22 cs.CV cs.LG 57%

CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey

Jindong Li, Yongguang Li, Yali Fu, Jiahong Liu, Yixin Liu, Menglin Yang, Irwin King

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Jilin University(吉林大学) Griffith University(格里菲斯大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13877 2025-04-22 cs.HC cs.AI 57%

New care pathways for supporting transitional care from hospitals to home using AI and personalized digital assistance

Ionut Anghel, Tudor Cioara, Roberta Bevilacqua, Federico Barbarossa, Terje Grimstad, Riitta Hellman, Arnor Solberg, Lars Thomas Boye, Ovidiu Anchidin, Ancuta Nemes, Camilla Gabrielsen

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments submitted to journal (under review)

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.07711 2025-04-22 cs.CY 57%

Explainability Auditing for Intelligent Systems: A Rationale for Multi-Disciplinary Perspectives

Markus Langer, Kevin Baum, Kathrin Hartmann, Stefan Hessel, Timo Speith, Jonas Wahl

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY

Comments Accepted at the First International Workshop on Requirements Engineering for Explainable Systems (RE4ES) co-located with the 29th IEEE International Requirements Engineering Conference (RE'21)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10353 2025-04-21 cs.CV cs.CR cs.LG 57%

Robust image classification with multi-modal large language models

Francesco Villani, Igor Maljkovic, Dario Lazzaro, Angelo Sotgiu, Antonio Emanuele Cinà, Fabio Roli

机构 * University of Genoa(热那亚大学) Sapienza University of Rome(罗马第一大学(萨皮恩扎大学)) University of Cagliari(卡利亚里大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Paper accepted at Pattern Recognition Letters journal Keywords: adversarial examples, rejection defense, multimodal-informed systems, machine learning security

Journal ref Pattern Recognition Letters 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12722 2025-04-18 cs.IR cs.AI 57%

SimUSER: Simulating User Behavior with Large Language Models for Recommender System Evaluation

Nicolas Bougie, Narimasa Watanabe

机构 * Woven by Toyota(丰田编织)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12422 2025-04-18 cs.HC cs.AI 57%

Mitigating LLM Hallucinations with Knowledge Graphs: A Case Study

Harry Li, Gabriel Appleby, Kenneth Alperin, Steven R Gomez, Ashley Suh

机构 * MIT Lincoln Laboratory(麻省理工学院林肯实验室) National Renewable Energy Laboratory(国家可再生能源实验室)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Presented at the Human-centered Explainable AI Workshop (HCXAI) @ CHI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11543 2025-04-18 cs.AI 57%

REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

Divyansh Garg, Shaun VanWeelden, Diego Caples, Andis Draguns, Nikil Ravi, Pranav Putta, Naman Garg, Tomas Abraham, Michael Lara, Federico Lopez, James Liu, Atharva Gundawar, Prannay Hebbar, Youngchul Joo, Jindong Gu, Charles London, Christian Schroeder de Witt, Sumeet Motwani

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments The websites, framework, and leaderboard are available at https://realevals.xyz and https://github.com/agi-inc/REAL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17385 2025-04-18 cs.CL cs.CV 57%

Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Zheyuan Zhang, Fengyuan Hu, Jayjun Lee, Freda Shi, Parisa Kordjamshidi, Joyce Chai, Ziqiao Ma

机构 * University of Michigan(密歇根大学) University of Waterloo(滑铁卢大学) Vector Institute(向量研究所) Michigan State University(密歇根州立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted to ICLR 2025 (Oral) | Project page: https://spatial-comfort.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11990 2025-04-17 cs.LG cs.CR 57%

Secure Transfer Learning: Training Clean Models Against Backdoor in (Both) Pre-trained Encoders and Downstream Datasets

Yechao Zhang, Yuxuan Zhou, Tianyu Li, Minghui Li, Shengshan Hu, Wei Luo, Leo Yu Zhang

机构 * Huazhong University of Science and Technology(华中科技大学) Deakin University(迪肯大学) Griffith University(格里菲斯大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments To appear at IEEE Symposium on Security and Privacy 2025, 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10757 2025-04-16 cs.CV cs.LG cs.RO 57%

ReasonDrive: Efficient Visual Question Answering for Autonomous Vehicles with Reasoning-Enhanced Small Vision-Language Models

Amirhosein Chahe, Lifeng Zhou

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10615 2025-04-16 cs.CL 57%

Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models

Thilo Hagendorff, Sarah Fabi

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18682 2025-04-16 cs.HC cs.AI 57%

AI Mismatches: Identifying Potential Algorithmic Harms Before AI Development

Devansh Saxena, Ji-Youn Jung, Jodi Forlizzi, Kenneth Holstein, John Zimmerman

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Carnegie Mellon University(卡内基梅隆大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments CHI Conference on Human Factors in Computing Systems (CHI '25), April 26-May 1, 2025, Yokohama, Japan

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09277 2025-04-15 cs.IR cs.AI 57%

SynthTRIPs: A Knowledge-Grounded Framework for Benchmark Query Generation for Personalized Tourism Recommenders

Ashmi Banerjee, Adithi Satish, Fitri Nur Aisyah, Wolfgang Wörndl, Yashar Deldjoo

机构 * Technical University of Munich(慕尼黑工业大学) Polytechnic University of Bari(巴里理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted for publication at SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04569 2025-04-15 cs.CL 57%

KnowsLM: A framework for evaluation of small language models for knowledge augmentation and humanised conversations

Chitranshu Harbola, Anupam Purwar

机构 * AIGuruKul Foundation(AIGuruKul基金会)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06370 2025-04-15 cs.SE cs.AI 57%

Towards a Probabilistic Framework for Analyzing and Improving LLM-Enabled Software

Juan Manuel Baldonado, Flavia Bonomo-Braberman, Víctor Adrián Braberman

机构 * Universidad de Buenos Aires(布宜诺斯艾利斯大学) CONICET(阿根廷国家科学技术研究委员会)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Journal ref ICST Workshops 2025, Naples, Italy: SAFE-ML 2025, 418--422

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18449 2025-04-15 eess.IV cs.CV cs.LG 57%

Towards A Generalizable Pathology Foundation Model via Unified Knowledge Distillation

Jiabo Ma, Zhengrui Guo, Fengtao Zhou, Yihui Wang, Yingxue Xu, Jinbang Li, Fang Yan, Yu Cai, Zhengjie Zhu, Cheng Jin, Yi Lin, Xinrui Jiang, Chenglong Zhao, Danyi Li, Anjia Han, Zhenhui Li, Ronald Cheong Kin Chan, Jiguang Wang, Peng Fei, Kwang-Ting Cheng, Shaoting Zhang, Li Liang, Hao Chen

机构 * The Hong Kong University of Science and Technology(香港科技大学) Southern Medical University(南方医科大学) Guangdong Provincial Key Laboratory of Molecular Tumor Pathology(广东省分子肿瘤病理学重点实验室) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The First Affiliated Hospital of Shandong First Medical University and Shandong Provincial Qianfoshan Hospital(山东第一医科大学第一附属医院暨山东省千佛山医院) Sun Yat-sen University(中山大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments update

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02191 2025-04-15 cs.LG cs.CV stat.ME stat.ML 57%

Evaluating AI systems under uncertain ground truth: a case study in dermatology

David Stutz, Ali Taylan Cemgil, Abhijit Guha Roy, Tatiana Matejovicova, Melih Barsbey, Patricia Strachan, Mike Schaekermann, Jan Freyberg, Rajeev Rikhye, Beverly Freeman, Javier Perez Matos, Umesh Telang, Dale R. Webster, Yuan Liu, Greg S. Corrado, Yossi Matias, Pushmeet Kohli, Yun Liu, Arnaud Doucet, Alan Karthikesalingam

机构 * Google DeepMind(谷歌DeepMind) Google(谷歌) Bogazici University(博加齐奇大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08786 2025-04-15 cs.IR cs.AI 57%

AdaptRec: A Self-Adaptive Framework for Sequential Recommendations with Large Language Models

Tong Zhang

机构 * University of Technology Sydney(悉尼科技大学) The Hong Kong Polytechnic University(香港理工大学) University of Science and Technology of China(中国科学技术大学) The Education University of Hong Kong(香港教育大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08583 2025-04-14 astro-ph.IM cs.LG 57%

AstroLLaVA: towards the unification of astronomical data and natural language

Sharaf Zaman, Michael J. Smith, Pranav Khetarpal, Rishabh Chakrabarty, Michele Ginolfi, Marc Huertas-Company, Maja Jabłońska, Sandor Kruk, Matthieu Le Lain, Sergio José Rodríguez Méndez, Dimitrios Tanoglidis

机构 * UniverseTBD Indian Institute of Technology Delhi(印度理工学院德里分校) Intelligent Internet Inc.(Intelligent Internet公司) University of Florence(佛罗伦萨大学) Instituto de Astrofísica de Canarias (IAC)(加那利群岛天体物理学研究所) Observatoire de Paris(巴黎天文台) PSL University(巴黎文理研究大学) Université Paris-Cité(巴黎城市大学) ANU RSAA(澳大利亚国立大学天体物理学与天体生物学研究学院) European Space Agency(欧洲空间局) IRISA(IRISA研究所) Université Bretagne Sud(南布列塔尼大学) ANU School of Computing(澳大利亚国立大学计算机学院)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 8 pages, 3 figures, accepted to SCI-FM@ICLR 2025. Code at https://w3id.org/UniverseTBD/AstroLLaVA

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08269 2025-04-14 cs.CV cs.CL 57%

VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering

Qi Zhi Lim, Chin Poo Lee, Kian Ming Lim, Kalaiarasi Sonai Muthu Anbananthen

机构 * Multimedia University(多媒体大学) University of Nottingham Ningbo China(宁波诺丁汉大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07596 2025-04-14 cs.AI 57%

Boosting Universal LLM Reward Design through Heuristic Reward Observation Space Evolution

Zen Kit Heng, Zimeng Zhao, Tianhao Wu, Yuanfei Wang, Mingdong Wu, Yangang Wang, Hao Dong

机构 * Center on Frontiers of Computing Studies, School of Computer Science, Peking University(北京大学计算机学院计算前沿研究中心) PKU-Agibot Lab, School of Computer Science, Peking University(北京大学计算机学院PKU-Agibot实验室) National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理国家重点实验室) School of Automation, Southeast University(东南大学自动化学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07113 2025-04-11 cs.CL cs.DB 57%

How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities

Aly M. Kassem, Bernhard Schölkopf, Zhijing Jin

机构 * MPI for Intelligent Systems(智能系统MPI研究所) University of Toronto(多伦多大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06785 2025-04-10 cs.CV cs.AI 57%

Zero-Shot Image-Based Large Language Model Approach to Road Pavement Monitoring

Shuoshuo Xu, Kai Zhao, James Loney, Zili Li, Andrea Visentin

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05259 2025-04-08 cs.AI cs.CR 57%

How to evaluate control measures for LLM agents? A trajectory from today to superintelligence

Tomek Korbak, Mikita Balesni, Buck Shlegeris, Geoffrey Irving

机构 * UK AI Security Institute(英国人工智能安全研究院) Apollo Research(阿波罗研究院) Redwood Research(红杉研究院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21157 2025-04-08 cs.LG 57%

Real-Time Evaluation Models for RAG: Who Detects Hallucinations Best?

Ashish Sardana

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 11 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04608 2025-04-08 cs.AI cs.SY eess.SY 57%

AI in a vat: Fundamental limits of efficient world modelling for agent sandboxing and interpretability

Fernando Rosas, Alexander Boyd, Manuel Baltieri

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 38 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03478 2025-04-07 cs.LG cs.CV 57%

Probabilistic Machine Learning for Noisy Labels in Earth Observation

Spyros Kondylatos, Nikolaos Ioannis Bountos, Ioannis Prapas, Angelos Zavras, Gustau Camps-Valls, Ioannis Papoutsis

机构 * Orion Lab(猎户座实验室) National Observatory of Athens(雅典国家天文台) National Technical University of Athens(雅典国立技术大学) Image Processing Laboratory (IPL), Universitat de València(瓦伦西亚大学图像处理实验室) Harokopio University of Athens(哈罗科皮奥大学) Archimedes, Athena Research Center(雅典娜研究中心阿基米德分部)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02178 2025-04-07 cs.CL 57%

Subasa - Adapting Language Models for Low-resourced Offensive Language Detection in Sinhala

Shanilka Haturusinghe, Tharindu Cyril Weerasooriya, Marcos Zampieri, Christopher M. Homan, S. R. Liyanage

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments Accepted to appear at NAACL SRW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏