arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9346 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9346 篇

2509.00806 2025-09-03 cs.CL cs.AI cs.LG 67%

CaresAI at BioCreative IX Track 1 -- LLM for Biomedical QA

Reem Abdel-Salam, Mary Adewunmi, Modinat A. Abayomi

机构 * Cairo University, Faculty of Engineering, Computer Engineering Department(开罗大学工程学院计算机工程系) Menzies School of Health Research, Charles Darwin University, NT, Australia(梅恩兹健康研究学院,查尔斯达尔文大学,澳大利亚NT) Department of Biology, Boston College, Massachusetts, USA(生物学系,波士顿学院,马萨诸塞州,美国) Peoples' Friendship University of Russia (RUDN University)(俄罗斯人民友谊大学(RUDN大学)) Joint Institute for Nuclear Research(联合核研究中心) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) University of Skövde(斯德哥尔摩大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Proceedings of the BioCreative IX Challenge and Workshop (BC9): Large Language Models for Clinical and Biomedical NLP at the International Joint Conference on Artificial Intelligence (IJCAI), Montreal, Canada, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21271 2025-09-01 cs.RO cs.CV 67%

Mini Autonomous Car Driving based on 3D Convolutional Neural Networks

Pablo Moraes, Monica Rodriguez, Kristofer S. Kappel, Hiago Sodre, Santiago Fernandez, Igor Nunes, Bruna Guterres, Ricardo Grando

机构 * Technological University of Uruguay, UTEC, Uruguay(乌拉圭技术大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14377 2025-08-26 cs.CL cs.AI cs.CY 67%

ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities

Wenhan Dong, Zhen Sun, Yuemeng Zhao, Zifan Peng, Jun Wu, Jingyi Zheng, Yule Liu, Xinlei He, Yu Wang, Ruiming Wang, Xinyi Huang, Lei Mo

机构 * School of Psychology, South China Normal University(南方科技大学心理学院) Information Hub, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)信息中心) School of AI, Guangzhou University(广州大学人工智能学院) College of Cyber Security, Jinan University(济南大学网络安全学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15797 2025-08-25 cs.CL cs.AI cs.LG 67%

Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks

Nouar AlDahoul, Yasir Zaki

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15418 2025-08-22 cs.CL cs.AI cs.LG cs.MM cs.SD 67%

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model

Yirong Sun, Yizhong Geng, Peidong Wei, Yanjun Chen, Jinghan Yang, Rongfei Chen, Wei Zhang, Xiaoyu Shen

机构 * Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT(宁波空间智能与数字衍生关键实验室,数字孪生研究院,EIT) Logic Intelligence Technology(逻辑智能技术) BUPT(北京邮电大学) Xiamen University(厦门大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15085 2025-08-22 cs.CL cs.AI cs.IR cs.LG 67%

LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text

MohamamdJavad Ardestani, Ehsan Kamalloo, Davood Rafiei

机构 * Department of Computing Science University of Alberta(计算科学系阿尔伯塔大学) ServiceNow Research(ServiceNow研究)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10028 2025-08-15 cs.CL cs.AI cs.HC cs.LG 67%

PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs

Xiao Fu, Hossein A. Rahmani, Bin Wu, Jerome Ramos, Emine Yilmaz, Aldo Lipani

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07308 2025-08-12 cs.CL cs.AI cs.IR cs.LG 67%

HealthBranches: Synthesizing Clinically-Grounded Question Answering Datasets via Decision Pathways

Cristian Cosentino, Annamaria Defilippo, Marco Dossena, Christopher Irwin, Sara Joubbi, Pietro Liò

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13677 2025-08-11 cs.CL cs.AI cs.LG 67%

Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results

Andrea Santilli, Adam Golinski, Michael Kirchhof, Federico Danieli, Arno Blaas, Miao Xiong, Luca Zappella, Sinead Williamson

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at ACL 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04350 2025-08-07 cs.CL cs.AI cs.CV cs.LG cs.MA 67%

Chain of Questions: Guiding Multimodal Curiosity in Language Models

Nima Iji, Kia Dashtipour

机构 * Edinburgh Napier University(爱丁堡纳皮尔大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23535 2025-08-01 cs.LG cs.AI cs.CY 67%

Transparent AI: The Case for Interpretability and Explainability

Dhanesh Ramachandram, Himanshu Joshi, Judy Zhu, Dhari Gandhi, Lucas Hartman, Ananya Raval

机构 * Vector Institute for Artificial Intelligence(向量人工智能研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21919 2025-07-31 cs.CL cs.AI cs.CY 67%

Training language models to be warm and empathetic makes them less reliable and more sycophantic

Lujain Ibrahim, Franziska Sofia Hafner, Luc Rocher

机构 * Oxford Internet Institute(牛津互联网研究所) University of Oxford(牛津大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22034 2025-07-30 cs.AI cs.CL cs.LG 67%

UserBench: An Interactive Gym Environment for User-Centric Agents

Cheng Qian, Zuxin Liu, Akshara Prabhakar, Zhiwei Liu, Jianguo Zhang, Haolin Chen, Heng Ji, Weiran Yao, Shelby Heinecke, Silvio Savarese, Caiming Xiong, Huan Wang

机构 * Salesforce AI Research(Salesforce AI研究部) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 25 Pages, 17 Figures, 6 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20526 2025-07-29 cs.AI cs.CL cs.CY 67%

Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition

Andy Zou, Maxwell Lin, Eliot Jones, Micha Nowak, Mateusz Dziemian, Nick Winter, Alexander Grattan, Valent Nathanael, Ayla Croft, Xander Davies, Jai Patel, Robert Kirk, Nate Burnikell, Yarin Gal, Dan Hendrycks, J. Zico Kolter, Matt Fredrikson

专题命中 安全评测 :red teaming(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19598 2025-07-29 cs.CL cs.AI cs.CR cs.LG 67%

MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?

Muntasir Wahed, Xiaona Zhou, Kiet A. Nguyen, Tianjiao Yu, Nirav Diwan, Gang Wang, Dilek Hakkani-Tür, Ismini Lourentzou

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Winner Defender Team at Amazon Nova AI Challenge 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07817 2025-07-16 cs.CL cs.AI cs.LG 67%

On the Effect of Instruction Tuning Loss on Generalization

Anwoy Chatterjee, H S V N S Kowndinya Renduchintala, Sumit Bhatia, Tanmoy Chakraborty

机构 * Dept. of Electrical Engineering(电气工程系) Indian Institute of Technology Delhi(印度理工学院德里) Media and Data Science Research(媒体与数据科学研究) Adobe Inc., India(Adobe公司,印度)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments To appear in Transactions of the Association for Computational Linguistics (TACL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06672 2025-07-15 cs.CY cs.AI cs.LG q-fin.RM 67%

Insuring Uninsurable Risks from AI: Government as Insurer of Last Resort

Cristian Trout

机构 * Independent Researcher, Cambridge Boston Alignment Initiative(独立研究者,剑桥波士顿对齐计划)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Accepted to Generative AI and Law Workshop at the International Conference on Machine Learning (ICML 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06895 2025-07-10 cs.CL cs.AI cs.IR cs.LG 67%

SCoRE: Streamlined Corpus-based Relation Extraction using Multi-Label Contrastive Learning and Bayesian kNN

Luca Mariotti, Veronica Guidetti, Federica Mandreoli

机构 * Department of Physical, Computer and Mathematical Sciences - University of Modena and Reggio Emilia(物理、计算机和数学科学系 - 模纳和雷吉奥艾米利亚大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20589 2025-07-09 cs.CR cs.AI cs.CL cs.LG 67%

LLMs Have Rhythm: Fingerprinting Large Language Models Using Inter-Token Times and Network Traffic Analysis

Saeif Alhazbi, Ahmed Mohamed Hussain, Gabriele Oligeri, Panos Papadimitratos

机构 * College of Science and Engineering (CSE), Hamad Bin Khalifa University (HBKU)(哈马德·本·卡伊夫大学科学与工程学院) Networked Systems Security Group, KTH Royal Institute of Technology -- Stockholm, Sweden(瑞典皇家理工学院网络系统安全组)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18045 2025-06-24 cs.CY cs.AI cs.CL 67%

The Democratic Paradox in Large Language Models' Underestimation of Press Freedom

I. Loaiza, R. Vestrelli, A. Fronzetti Colladon, R. Rigobon

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16831 2025-06-23 cs.SE 67%

Accountability of Robust and Reliable AI-Enabled Systems: A Preliminary Study and Roadmap

Filippo Scaramuzza, Damian A. Tamburri, Willem-Jan van den Heuvel

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

Comments To be published in https://link.springer.com/book/9789819672370

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15293 2025-06-19 cs.HC cs.RO 67%

Designing Intent: A Multimodal Framework for Human-Robot Cooperation in Industrial Workspaces

Francesco Chiossi, Julian Rasch, Robin Welsch, Albrecht Schmidt, Florian Michahelles

机构 * Aalto University(阿alto大学) TU Wien(维也纳技术大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

Comments 9 pages

Journal ref The Future of Human-Robot Synergy in Interactive Environments: The Role of Robots at the Workplace @ CHIWork 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15100 2025-06-19 cs.CR 67%

International Security Applications of Flexible Hardware-Enabled Guarantees

Onni Aarne, James Petrie

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10581 2025-06-19 q-bio.QM cs.AI cs.CL cs.LG eess.SP 67%

Large Language Model-informed ECG Dual Attention Network for Heart Failure Risk Prediction

Chen Chen, Lei Li, Marcel Beetz, Abhirup Banerjee, Ramneek Gupta, Vicente Grau

机构 * Institute of Biomedical Engineering, Department of Engineering Science, University of Oxford(生物医学工程研究所,工程科学系,牛津大学) Imperial College London(伦敦帝国学院) University of Sheffield(谢菲尔德大学) Novo Nordisk Research Centre Oxford (NNRCO)(牛津诺和硕研究中心(NNRCO)) Royal Society(皇家学会)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Under journal revision

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08210 2025-06-17 cs.CV cs.AI cs.CL cs.LG 67%

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation

Andrew Z. Wang, Songwei Ge, Tero Karras, Ming-Yu Liu, Yogesh Balaji

机构 * University of Maryland(马里兰大学) NVIDIA(英伟达)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments CVPR 2025

Journal ref Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), 2025, pp. 28575-28585

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12266 2025-06-17 cs.CL cs.AI cs.HC cs.LG 67%

The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs

Avinash Baidya, Kamalika Das, Xiang Gao

机构 * Intuit AI Research(Intuit AI研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ACL 2025; 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19794 2025-06-12 cs.CV 67%

MVTamperBench: Evaluating Robustness of Vision-Language Models

Amit Agarwal, Srikant Panda, Angeline Charles, Bhargava Kumar, Hitesh Patel, Priyaranjan Pattnayak, Taki Hasan Rafi, Tejaswini Kumar, Hansa Meghwani, Karan Gupta, Dong-Kyu Chae

机构 * Liverpool John Moores University(利物浦约翰·穆里斯大学) Birla Institute of Technology(比拉理工学院) Christ University(基督大学) Columbia University(哥伦比亚大学) New York University(纽约大学) University of Washington(华盛顿大学) Hanyang University(翰阳大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09420 2025-06-12 cs.AI cs.CL cs.HC cs.LG cs.MA 67%

A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy

Henry Peng Zou, Wei-Chieh Huang, Yaozu Wu, Chunyu Miao, Dongyuan Li, Aiwei Liu, Yue Zhou, Yankai Chen, Weizhi Zhang, Yangning Li, Liancheng Fang, Renhe Jiang, Philip S. Yu

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Tokyo(东京大学) Tsinghua University(清华大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13151 2025-06-10 cs.LG cs.AI cs.CL 67%

MIB: A Mechanistic Interpretability Benchmark

Aaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad, Iván Arcuschin, Adam Belfki, Yik Siu Chan, Jaden Fiotto-Kaufman, Tal Haklay, Michael Hanna, Jing Huang, Rohan Gupta, Yaniv Nikankin, Hadas Orgad, Nikhil Prakash, Anja Reusch, Aruna Sankaranarayanan, Shun Shao, Alessandro Stolfo, Martin Tutek, Amir Zur, David Bau, Yonatan Belinkov

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to ICML 2025. Project website at https://mib-bench.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04089 2025-06-05 cs.LG cs.AI cs.CL cs.RO 67%

AmbiK: Dataset of Ambiguous Tasks in Kitchen Environment

Anastasiia Ivanova, Eva Bakaeva, Zoya Volovikova, Alexey K. Kovalev, Aleksandr I. Panov

机构 * LMU(慕尼黑大学) MIPT(莫斯科 Institute of Physics and Technology) AIRI(空气研究所)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ACL 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏