arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9378 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9378 篇

2408.05284 2025-06-17 cs.AI cs.LG 62%

Can a Bayesian Oracle Prevent Harm from an Agent?

Yoshua Bengio, Michael K. Cohen, Nikolay Malkin, Matt MacDermott, Damiano Fornasiere, Pietro Greiner, Younesse Kaddar

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted at UAI 2025 (Uncertainty in Artificial Intelligence). 20 pages, 2 figures. Code available at: https://github.com/saifh-github/conservative-bayesian-public

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11515 2025-06-16 cs.CV cs.CL cs.LG 62%

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs

Xiao Xu, Libo Qin, Wanxiang Che, Min-Yen Kan

机构 * Research Center for Social Computing and Information Retrieval, Harbin Institute of Technology(社会计算与信息检索研究中心,哈尔滨工业大学) School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学) School of Computing, National University of Singapore(计算学院,新加坡国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted by IEEE Transactions on Circuits and Systems for Video Technology (TCSVT). June 2025. DOI: https://doi.org/10.1109/TCSVT.2025.3578266

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10885 2025-06-13 cs.CL cs.AI 62%

Slimming Down LLMs Without Losing Their Minds

Qingda, Mai

机构 * University of Waterloo(滑铁卢大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06104 2025-06-12 cs.LG cs.AI 62%

Gradient Aligned Regression via Pairwise Losses

Dixian Zhu, Tianbao Yang, Livnat Jerby

机构 * Department of Genetics, Stanford University, CA, USA(遗传学系,斯坦福大学,加州,美国) Department of Computer Science, Texas A\&M University, TX, USA(计算机科学系,德克萨斯A&M大学,德克萨斯州,美国)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments ICML 2025; 23 pages, 12 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09160 2025-06-11 cs.AI cs.LG cs.RO 62%

Innate-Values-driven Reinforcement Learning based Cognitive Modeling

Qin Yang

机构 * Intelligent Social Systems and Swarm Robotics Lab (IS$^3$R)(智能社会系统与群体机器人实验室) Computer Science and Information Systems Department(计算机科学与信息系统系) Bradley University(布拉德利大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments The paper had been accepted by the 2025 IEEE Conference on Cognitive and Computational Aspects of Situation Management (CogSIMA). arXiv admin note: text overlap with arXiv:2401.05572

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06528 2025-06-10 cs.CL cs.AI cs.HC 62%

Epistemic Integrity in Large Language Models

Bijean Ghafouri, Shahrad Mohammadzadeh, James Zhou, Pratheeksha Nair, Jacob-Junqi Tian, Hikaru Tsujimura, Mayank Goel, Sukanya Krishna, Reihaneh Rabbany, Jean-François Godbout, Kellin Pelrine

机构 * University of Southern California(南加州大学) McGill University(麦吉尔大学) UC Berkeley(加州大学伯克利分校) Vector Institute(向量研究所) Cardiff University(卡迪夫大学) University College London(伦敦大学学院) IIIT Hyderabad(海得拉巴印度理工学院) Harvard University(哈佛大学) Université de Montréal(蒙特利尔大学) Mila

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06737 2025-06-10 cs.CL cs.AI 62%

C-PATH: Conversational Patient Assistance and Triage in Healthcare System

Qi Shi, Qiwei Han, Cláudia Soares

机构 * School of Business and Economics(商业与经济学院) School of Science and Technology(科学与技术学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted in IEEE ICDH 2025, 10 pages, 8 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06020 2025-06-09 cs.CL cs.AI 62%

When to Trust Context: Self-Reflective Debates for Context Reliability

Zeqi Zhou, Fang Wu, Shayan Talaei, Haokai Zhao, Cheng Meixin, Tinson Xu, Amin Saberi, Yejin Choi

机构 * Brown University(布朗大学) Stanford University(斯坦福大学) University of New South Wales(新南威尔士大学) Xi’an University of Electronic Science and Technology(西安电子科技大学) University of Chicago(芝加哥大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04474 2025-06-06 cs.LG cs.AI 62%

Classifying Dental Care Providers Through Machine Learning with Features Ranking

Mohammad Subhi Al-Batah, Mowafaq Salem Alzboon, Muhyeeddin Alqaraleh, Mohammed Hasan Abu-Arqoub, Rashiq Rafiq Marie

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Journal ref Data and Metadata. 2025; 4:755

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01782 2025-06-03 cs.CY cs.AI cs.SY eess.SY 62%

Systematic Hazard Analysis for Frontier AI using STPA

Simon Mylius

机构 * Simon Mylius

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY

Comments 29 pages, 5 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15865 2025-06-03 q-fin.GN cs.AI cs.CL 62%

Standard Benchmarks Fail -- Auditing LLM Agents in Finance Must Prioritize Risk

Zichen Chen, Jiaao Chen, Jianda Chen, Misha Sra

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments 46 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01206 2025-06-03 cs.CL cs.AI 62%

Mamba Drafters for Speculative Decoding

Daewon Choi, Seunghyuk Oh, Saket Dingliwal, Jihoon Tack, Kyuyoung Kim, Woomin Song, Seojin Kim, Insu Han, Jinwoo Shin, Aram Galstyan, Shubham Katiyar, Sravan Babu Bodapati

机构 * KAIST(韩国科学技术院) Amazon AGI(亚马逊人工智能实验室) Seoul National University(首尔国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14301 2025-06-03 cs.CL cs.AI 62%

SEA-HELM: Southeast Asian Holistic Evaluation of Language Models

Yosephine Susanto, Adithya Venkatadri Hulagadri, Jann Railey Montalan, Jian Gang Ngui, Xian Bin Yong, Weiqi Leong, Hamsawardhini Rengarajan, Peerat Limkonchotiwat, Yifan Mai, William Chandra Tjhi

机构 * AI Singapore(AI新加坡) National University of Singapore(国立新加坡大学) Center for Research on Foundation Models (CRFM)(基础模型研究中心) Stanford University(斯坦福大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06560 2025-06-03 cs.CL cs.CY 62%

Position: It's Time to Act on the Risk of Efficient Personalized Text Generation

Eugenia Iofinova, Andrej Jovanovic, Dan Alistarh

机构 * Institute of Science and Technology Austria(奥地利科学与技术研究所)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00113 2025-06-03 cs.CL cs.AI 62%

Wait, that's not an option: LLMs Robustness with Incorrect Multiple-Choice Options

Gracjan Góral, Emilia Wiśnios, Piotr Sankowski, Paweł Budzianowski

机构 * University of Warsaw(华沙大学) Institute of Mathematics, Polish Academy of Sciences(波兰科学院数学研究所) MIM Solutions(MIM解决方案) K-Scale Labs(K-Scale实验室) IDEAS NCBR IDEAS Research Institute(IDEAS研究学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted for ACL 2025 Main Conference and NeurIPS 2024 FM-EduAssess Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00047 2025-06-03 cs.CY cs.AI cs.CE 62%

Risks of AI-driven product development and strategies for their mitigation

Jan Göpfert, Jann M. Weinand, Patrick Kuckertz, Noah Pflugradt, Jochen Linßen

机构 * Institute of Climate and Energy Systems, Jülich Systems Analysis (ICE-2), Forschungszentrum Jülich(气候与能源系统研究所,朱利奇系统分析(ICE-2),朱利奇研究中心) Chair for Fuel Cells, Faculty of Mechanical Engineering, RWTH Aachen University(燃料电池主任,机械工程学院,亚琛工业大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14956 2025-06-03 cs.CL cs.AI cs.IR 62%

ExPerT: Effective and Explainable Evaluation of Personalized Long-Form Text Generation

Alireza Salemi, Julian Killingback, Hamed Zamani

机构 * Center for Intelligent Information Retrieval(智能信息检索中心) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12583 2025-06-02 cs.RO cs.AI cs.LG 62%

A Comprehensive Survey on Physical Risk Control in the Era of Foundation Model-enabled Robotics

Takeshi Kojima, Yaonan Zhu, Yusuke Iwasawa, Toshinori Kitamura, Gang Yan, Shu Morikuni, Ryosuke Takanami, Alfredo Solano, Tatsuya Matsushima, Akiko Murakami, Yutaka Matsuo

机构 * The University of Tokyo(东京大学) Japan AI Safety Institute(日本人工智能安全研究所)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to IJCAI 2025 Survey Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09429 2025-06-02 cs.LG cs.CL cs.CV 62%

Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models

Kening Zheng, Junkai Chen, Yibo Yan, Xin Zou, Xuming Hu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Guangxi Zhuang Autonomous Region Big Data Research Institute(广西壮族自治区大数据研究院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

Comments Accepted by Findings of ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14279 2025-05-30 cs.CL cs.AI 62%

YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering

Jennifer D'Souza, Hamed Babaei Giglou, Quentin Münch

机构 * TIB Leibniz Information Centre for Science and Technology(蒂宾根科学与技术信息中心) Leibniz Universität Hannover(汉诺威莱布尼茨大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 4 figures, Accepted as a Long Paper at the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12941 2025-05-30 cs.CL cs.LG 62%

HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Language Model

Haiyang Guo, Fanhu Zeng, Ziwei Xiang, Fei Zhu, Da-Han Wang, Xu-Yao Zhang, Cheng-Lin Liu

机构 * School of Advanced Interdisciplinary Sciences, UCAS(UCAS先进交叉科学学院) MAIS, CASIA(CASIA人工智能研究所) School of Artificial Intelligence, UCAS(UCAS人工智能学院) Centre for Artificial Intelligence and Robotics, HKISI-CAS(HKISI-CAS人工智能与机器人中心) FKLPRIU, School of Computer and Information Engineering, Xiamen University of Technology(厦门理工学院计算机与信息工程学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments ACL 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22572 2025-05-29 cs.CL cs.AI 62%

Fusion Steering: Prompt-Specific Activation Control

Waldemar Chang, Alhassan Yasin

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 14 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22093 2025-05-29 cs.CY cs.AI cs.HC 62%

From Coders to Critics: Empowering Students through Peer Assessment in the Age of AI Copilots

Santiago Berrezueta-Guzman, Stephan Krusche, Stefan Wagner

机构 * Technical University of Munich Heilbronn, Germany(慕尼黑技术大学(海因斯贝格分校)) Technical University of Munich Munich, Germany(慕尼黑技术大学(慕尼黑分校))

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments This is the authors' preprint version of a paper accepted at the 11th International Symposium on Educational Technology, to be held in July 2025, in Bangkok, Thailand. The final published version will be available via IEEE Xplore Library

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14510 2025-05-29 cs.AI cs.LG 62%

BACON: A fully explainable AI model with graded logic for decision making problems

Haishi Bai, Jozo Dujmovic, Jianwu Wang

机构 * Department of Information Systems University of Maryland, Baltimore County (UMBC)(信息系统系大学马里兰大学巴尔的摩县(UMBC)) Department of Computer Science San Francisco State University(计算机科学系旧金山州立大学) Department of Information Systems UMBC(信息系统系大学马里兰大学巴尔的摩县(UMBC))

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21191 2025-05-28 cs.CL cs.LG 62%

Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities

Junyan Zhang, Yubo Gao, Yibo Yan, Jungang Li, Zhaorui Hou, Sicheng Tao, Shuliang Liu, Song Dai, Yonghua Hei, Junzhuo Li, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21154 2025-05-28 cs.MA cs.AI cs.CY 62%

GGBond: Growing Graph-Based AI-Agent Society for Socially-Aware Recommender Simulation

Hailin Zhong, Hanlin Wang, Yujun Ye, Meiyi Zhang, Shengxin Zhu

机构 * Faculty of Science and Technology, Beijing Normal-Hong Kong Baptist University(科技学院,北京师范大学-香港 Baptist大学) Research Centers for Mathematics, Advanced Institute of Natural Sciences, Beijing Normal University(数学研究中心,北京师范大学自然科学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20312 2025-05-28 cs.CY cs.AI cs.MA 62%

Let's Get You Hired: A Job Seeker's Perspective on Multi-Agent Recruitment Systems for Explaining Hiring Decisions

Aditya Bhattacharya, Katrien Verbert

机构 * KU Leuven(根特大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments Pre-print version only. Please check the published version for any reference or citation

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05374 2025-05-28 cs.LG cs.CL 62%

Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond

Chongyu Fan, Jinghan Jia, Yihua Zhang, Anil Ramakrishna, Mingyi Hong, Sijia Liu

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.LG

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19568 2025-05-27 cs.AI cs.LG 62%

MSD-LLM: Predicting Ship Detention in Port State Control Inspections with Large Language Model

Jiongchao Jin, Xiuju Fu, Xiaowei Gao, Tao Cheng, Ran Yan

机构 * Institute of High Performance Computing(高性能计算研究所) Agency for Science, Technology and Research(科技研究局) Department of Earth Science and Engineering(地球科学与工程系) Imperial College London(伦敦帝国理工学院) SpaceTimeLab, Department of Civil, Environmental and Geomatic Engineering(空间时间实验室,土木、环境与测绘工程系) University College London(伦敦大学学院) School of Civil and Environmental Engineering(土木与环境工程学院) Nanyang Technological University(南洋理工大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18622 2025-05-27 cs.LG cs.AI stat.ML 62%

Trust, or Don't Predict: Introducing the CWSA Family for Confidence-Aware Model Evaluation

Kourosh Shahnazari, Seyed Moein Ayyoubzadeh, Mohammadali Keshtparvar, Pegah Ghaffari

机构 * Amirkabir University of Technology(阿米尔卡比尔技术大学) Semnan University(塞姆南大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏