arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9400 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9400 篇

2511.00606 2025-11-05 cs.CL 74%

SpecDiff-2: Scaling Diffusion Drafter Alignment For Faster Speculative Decoding

Jameson Sandler, Jacob K. Christopher, Thomas Hartvigsen, Ferdinando Fioretto

专题命中 安全评测 :alignment(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02148 2025-11-05 cs.LG 74%

CFL: On the Use of Characteristic Function Loss for Domain Alignment in Machine Learning

Abdullah Almansour, Ozan Tonguz

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 安全评测 :alignment(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14681 2025-10-31 cs.CL 74%

Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Quality

Yuto Harada, Yusuke Yamauchi, Yusuke Oda, Yohei Oseki, Yusuke Miyao, Yu Takagi

机构 * NII LLMC(日本信息处理学会大语言模型中心) The University of Tokyo(东京大学) NAIST(日本科学技术大学) Nagoya Institute of Technology(名古屋技术大学)

专题命中 安全评测 :alignment(title);分类 cs.CL

Comments Accepted to EMNLP 2025 (Main Conference). Models and evaluation results available at: https://github.com/llm-jp/massive-sft

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01128 2025-10-29 cs.CV cs.AI 74%

RipVIS: Rip Currents Video Instance Segmentation Benchmark for Beach Monitoring and Safety

Andrei Dumitriu, Florin Tatui, Florin Miron, Aakash Ralhan, Radu Tudor Ionescu, Radu Timofte

机构 * Computer Vision Lab, CAIDAS & IFI, University of Würzburg, Germany(计算机视觉实验室,CAIDAS与IFI,乌尔姆大学,德国) University of Bucharest, Romania(布加勒斯特大学,罗马尼亚)

专题命中 安全评测 :safety(title);分类 cs.AI

Comments Accepted at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03815 2025-10-07 eess.SY cs.LG cs.SY eess.SP 74%

A Trustworthy Industrial Fault Diagnosis Architecture Integrating Probabilistic Models and Large Language Models

Yue wu

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments 1tables,6 figs,11pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14744 2025-10-07 cs.LG 74%

Beyond the Single-Best Model: Rashomon Partial Dependence Profile for Trustworthy Explanations in AutoML

Mustafa Cavus, Jan N. van Rijn, Przemysław Biecek

机构 * Department of Statistics, Eskisehir Technical University, Turkiye(埃斯基谢普大学统计系) Leiden Institute of Advanced Computer Science, Leiden University, the Netherlands(莱顿大学高级计算机科学研究所) Faculty of Mathematics and Information Science, Warsaw University of Technology, Poland(华沙理工大学数学与信息科学学院) Informatics and Mechanics, University of Warsaw, Faculty of Mathematics, Poland(华沙大学信息技术与力学系)

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments Accepted at 28th International Conference on Discovery Science 2025

Journal ref In: Džeroski, S., Levatić, J., Pio, G., Simidjievski, N. (eds) Discovery Science. DS 2025. Lecture Notes in Computer Science, vol 16090. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17559 2025-09-23 cs.CL 74%

Specification-Aware Machine Translation and Evaluation for Purpose Alignment

Yoko Kayano, Saku Sugawara

机构 * The Graduate University for Advanced Studies (SOKENDAI)(高级研究大学(SOKENDAI)) National Institute of Informatics(信息研究所)

专题命中 安全评测 :alignment(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17514 2025-09-18 cs.AI 74%

TAI Scan Tool: A RAG-Based Tool With Minimalistic Input for Trustworthy AI Self-Assessment

Athanasios Davvetas, Xenia Ziouvelou, Ypatia Dami, Alexios Kaponis, Konstantina Giouvanopoulou, Michael Papademas

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments 9 pages, 1 figure, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08912 2025-09-12 cs.CY cs.HC 74%

Towards Trustworthy AI: Characterizing User-Reported Risks across LLMs "In the Wild"

Lingyao Li, Renkai Ma, Zhaoqian Xue, Junjie Xiong

专题命中 安全评测 :trustworthy(title);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18827 2025-09-09 cs.SE cs.AI 74%

Test It Before You Trust It: Applying Software Testing for Trustworthy In-context Learning

Teeradaj Racharak, Chaiyong Ragkhitwetsagul, Chommakorn Sontesadisai, Thanwadee Sunetnanta

机构 * Advanced Institute of So-Go-Chi (Convergence Knowledge) Informatics(融合知识研究院) Tohoku University(东北大学) Japan Advanced Institute of Science and Technology(日本先进科学研究院) Faculty of Information and Communication Technology(信息与通信技术学院) Mahidol University(玛希敦大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Journal ref Natural Language Processing and Information Systems (NLDB 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00673 2025-08-04 cs.CL 74%

MELAC: Massive Evaluation of Large Language Models with Alignment of Culture in Persian Language

Farhan Farsi, Farnaz Aghababaloo, Shahriar Shariati Motlagh, Parsa Ghofrani, MohammadAli SadraeiJavaheri, Shayan Bali, Amirhossein Shabani, Farbod Bijary, Ghazal Zamaninejad, AmirMohammad Salehoof, Saeedeh Momtazi

机构 * Amirkabir University of Technology(阿姆irkabir技术大学) Part AI Research Center(Part人工智能研究中心) University of Mazandaran(马赞德兰大学) King’s College London(伦敦国王学院)

专题命中 安全评测 :alignment(title);分类 cs.CL

Comments Preprint. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12384 2025-07-17 cs.LG cs.ET 74%

Trustworthy Tree-based Machine Learning by $MoS_2$ Flash-based Analog CAM with Inherent Soft Boundaries

Bo Wen, Guoyun Gao, Zhicheng Xu, Ruibin Mao, Xiaojuan Qi, X. Sharon Hu, Xunzhao Yin, Can Li

机构 * Department of Electrical and Electronic Engineering, The University of Hong Kong, Hong Kong SAR, China(香港大学电子与电气工程系) College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou, China(浙江大学信息科学与电子工程学院) Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, IN, USA(Notre Dame 大学计算机科学与工程系) Center for Advanced Semiconductor and Integrated Circuit, The University of Hong Kong, Hong Kong SAR, China(香港大学先进半导体与集成电路中心)

专题命中 安全评测 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15821 2025-06-23 cs.GR cs.AI cs.CV eess.IV 74%

VEIGAR: View-consistent Explicit Inpainting and Geometry Alignment for 3D object Removal

Pham Khai Nguyen Do, Bao Nguyen Tran, Nam Nguyen, Duc Dung Nguyen

机构 * AITech Lab(AITech实验室) Computer Science and Engineering Faculty(计算机科学与工程学院) Ho Chi Minh City University of Technology(胡志明市技术大学) VNUHCM

专题命中 安全评测 :alignment(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02621 2025-06-18 cs.NE cs.AI 74%

LLMs Help Alleviate the Cross-Subject Variability in Brain Signal and Language Alignment

Yifei Liu, Hengwei Ye, Shuhang Li

专题命中 安全评测 :alignment(title);分类 cs.AI

Comments The result is no longer believeable. Teaching force issue exists in the infer time of LLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17934 2025-06-06 cs.HC cs.CL cs.CR 74%

Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents

Chaoran Chen, Zhiping Zhang, Ibrahim Khalilov, Bingcan Guo, Simret A Gebreegziabher, Yanfang Ye, Ziang Xiao, Yaxing Yao, Tianshi Li, Toby Jia-Jun Li

机构 * University of Notre Dame(圣约翰大学) Northeastern University(东北大学) Virginia Tech(弗吉尼亚理工大学) University of Washington(华盛顿大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 安全评测 :trustworthy(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23813 2025-06-02 cs.CR cs.AI 74%

DP-RTFL: Differentially Private Resilient Temporal Federated Learning for Trustworthy AI in Regulated Industries

Abhijit Talluri

机构 * RTFL Project Contributor(RTFL项目贡献者)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments 6 pages (IEEE conference format), 10 figures. Source code available at https://github.com/abhitall/federated-credit-risk-rtfl.git

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21115 2025-05-28 cs.CL 74%

Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA

Sergey Pletenev, Maria Marina, Nikolay Ivanov, Daria Galimzianova, Nikita Krayko, Mikhail Salnikov, Vasily Konovalov, Alexander Panchenko, Viktor Moskvoretskii

专题命中 安全评测 :trustworthy(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20882 2025-05-28 cs.LG cs.SI 74%

Fedivertex: a Graph Dataset based on Decentralized Social Networks for Trustworthy Machine Learning

Marc Damie, Edwige Cyffers

机构 * University of Twente(代尔夫特理工大学) Inria(法国国家信息与自动化技术研究所) Institute of Science and Technology Austria(奥地利科学与技术研究所)

专题命中 安全评测 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11066 2025-05-21 cs.CL 74%

CARMA: Enhanced Compositionality in LLMs via Advanced Regularisation and Mutual Information Alignment

Nura Aljaafari, Danilo S. Carvalho, André Freitas

机构 * Department of Computer Science, University of Manchester(曼彻斯特大学计算机科学系) Idiap Research Institute(Idiap研究机构) National Biomarker Centre, CRUK-MI, Univ. of Manchester(国家生物标记中心,CRUK-MI,曼彻斯特大学)

专题命中 安全评测 :alignment(title);分类 cs.CL

Comments 19 pages, 8 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09083 2025-05-15 econ.GN cs.CL q-fin.EC 74%

Ornithologist: Towards Trustworthy "Reasoning" about Central Bank Communications

Dominic Zaun Eu Jones

机构 * Reserve Bank of Australia(澳大利亚储备银行)

专题命中 安全评测 :trustworthy(title);分类 cs.CL

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02843 2025-05-07 eess.IV cs.AI cs.CV physics.med-ph 74%

Physical foundations for trustworthy medical imaging: a review for artificial intelligence researchers

Miriam Cobo, David Corral Fontecha, Wilson Silva, Lara Lloret Iglesias

机构 * IFCA.unican.es(IFCA大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments 17 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15496 2025-04-29 cs.LG 74%

Verification and Validation for Trustworthy Scientific Machine Learning

John D. Jakeman, Lorena A. Barba, Joaquim R. R. A. Martins, Thomas O'Leary-Roseberry

机构 * Sandia National Laboratories(桑迪亚国家实验室) The George Washington University(乔治·华盛顿大学) University of Michigan(密歇根大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 安全评测 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13399 2025-04-21 cs.CV cs.AI 74%

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety

Shashank Shriram, Srinivasa Perisetla, Aryan Keskar, Harsha Krishnaswamy, Tonko Emil Westerhof Bossen, Andreas Møgelmose, Ross Greer

机构 * Machine Intelligence, Interaction, and Imagination (Mi 3 ) Laboratory(机器智能、交互与想象(Mi 3)实验室) University of California, Merced(加州大学默塞德分校) Aalborg Universitet(奥胡斯大学)

专题命中 安全评测 :safety(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17139 2025-04-18 cs.AI 74%

Trustworthy XAI and Application

MD Abdullah Al Nasim, A. S. M Anas Ferdous, Abdur Rashid, Fatema Tuj Johura Soshi, Parag Biswas, Angona Biswas, Kishor Datta Gupta

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20472 2025-03-27 cs.CV cs.AI 74%

From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment

Yucheng Suo, Fan Ma, Linchao Zhu, Tianyi Wang, Fengyun Rao, Yi Yang

专题命中 安全评测 :alignment(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00519 2025-02-05 cs.SE cs.LG 74%

CoDocBench: A Dataset for Code-Documentation Alignment in Software Maintenance

Kunal Pai, Premkumar Devanbu, Toufique Ahmed

专题命中 安全评测 :alignment(title);分类 cs.LG

Comments Accepted at the 2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR) - Data and Tool Showcase Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11909 2025-01-22 cs.AI 74%

Bridging the Communication Gap: Evaluating AI Labeling Practices for Trustworthy AI Development

Raphael Fischer, Magdalena Wischnewski, Alexander van der Staay, Katharina Poitz, Christian Janiesch, Thomas Liebig

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10379 2025-01-22 cs.CY 74%

What Information Should Be Shared with Whom "Before and During Training"?

Haydn Belfield

专题命中 安全评测 :safety(abstract,comments);AI safety(abstract,comments);分类 cs.CY

Comments To be published in the proceedings of the 2024 Conference on Frontier AI Safety Commitments. 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09218 2025-01-17 q-bio.QM cs.AI 74%

Interpretable Droplet Digital PCR Assay for Trustworthy Molecular Diagnostics

Yuanyuan Wei, Yucheng Wu, Fuyang Qu, Yao Mu, Yi-Ping Ho, Ho-Pui Ho, Wu Yuan, Mingkun Xu

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20674 2024-12-31 cs.DC cs.CR cs.LG 74%

Blockchain-Empowered Cyber-Secure Federated Learning for Trustworthy Edge Computing

Ervin Moore, Ahmed Imteaj, Md Zarif Hossain, Shabnam Rezapour, M. Hadi Amini

专题命中 安全评测 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏