arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9378 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9378 篇

2412.18947 2025-04-01 cs.CL cs.AI 62%

MedHallBench: A New Benchmark for Assessing Hallucination in Medical Large Language Models

Kaiwen Zuo, Yirui Jiang

专题命中 安全评测 :RLHF(abstract);分类 cs.CL、cs.AI

Comments Published to AAAI-25 Bridge Program

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22074 2025-03-31 cs.CL cs.AI 62%

Penrose Tiled Low-Rank Compression and Section-Wise Q&A Fine-Tuning: A General Framework for Domain-Specific Large Language Model Adaptation

Chuan-Wei Kuo, Siyu Chen, Chenqi Yan, Yu Yang Fredrik Liu

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21356 2025-03-28 cs.LG cs.AI 62%

Investigating the Duality of Interpretability and Explainability in Machine Learning

Moncef Garouani, Josiane Mothe, Ayah Barhrhouj, Julien Aligon

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17386 2025-03-27 cs.CL cs.AI 62%

Context-Aware Semantic Recomposition Mechanism for Large Language Models

Richard Katrix, Quentin Carroway, Rowan Hawkesbury, Matthias Heathfield

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19711 2025-03-26 cs.CL cs.AI cs.HC 62%

Writing as a testbed for open ended agents

Sian Gooding, Lucia Lopez-Rivilla, Edward Grefenstette

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19540 2025-03-26 cs.CL cs.AI 62%

FLEX: A Benchmark for Evaluating Robustness of Fairness in Large Language Models

Dahyun Jung, Seungyoon Lee, Hyeonseok Moon, Chanjun Park, Heuiseok Lim

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to NAACL 2025 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19394 2025-03-26 cs.LG cs.AI 62%

Quantifying Symptom Causality in Clinical Decision Making: An Exploration Using CausaLM

Mehul Shetty, Connor Jordan

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18983 2025-03-26 cs.CY cs.AI 62%

Confronting Catastrophic Risk: The International Obligation to Regulate Artificial Intelligence

Bryan Druzin, Anatole Boute, Michael Ramsden

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY

Journal ref Michigan Journal of International Law 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13999 2025-03-26 cs.CL cs.AI 62%

Framework for Progressive Knowledge Fusion in Large Language Models Through Structured Conceptual Redundancy Analysis

Joseph Sakau, Evander Kozlowski, Roderick Thistledown, Basil Steinberger

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18320 2025-03-25 cs.AI cs.CL 62%

Bridging Writing Manner Gap in Visual Instruction Tuning by Creating LLM-aligned Instructions

Dong Jing, Nanyi Fei, Zhiwu Lu

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17623 2025-03-25 cs.LG cs.AI 62%

Unraveling Pedestrian Fatality Patterns: A Comparative Study with Explainable AI

Methusela Sulle, Judith Mwakalonge, Gurcan Comert, Saidi Siuhi, Nana Kankam Gyimah

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 22 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09606 2025-03-24 cs.CL cs.AI 62%

Large Language Models and Causal Inference in Collaboration: A Survey

Xiaoyu Liu, Paiheng Xu, Junda Wu, Jiaxin Yuan, Yifan Yang, Yuhang Zhou, Fuxiao Liu, Tianrui Guan, Haoliang Wang, Tong Yu, Julian McAuley, Wei Ai, Furong Huang

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments Findings of the Association for Computational Linguistics: NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18908 2025-03-21 cs.LG cs.CL cs.CV 62%

Wolf: Dense Video Captioning with a World Summarization Framework

Boyi Li, Ligeng Zhu, Ran Tian, Shuhan Tan, Yuxiao Chen, Yao Lu, Yin Cui, Sushant Veer, Max Ehrlich, Jonah Philion, Xinshuo Weng, Fuzhao Xue, Linxi Fan, Yuke Zhu, Jan Kautz, Andrew Tao, Ming-Yu Liu, Sanja Fidler, Boris Ivanovic, Trevor Darrell, Jitendra Malik, Song Han, Marco Pavone

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14621 2025-03-20 cs.LG cs.AI 62%

Reducing False Ventricular Tachycardia Alarms in ICU Settings: A Machine Learning Approach

Grace Funmilayo Farayola, Akinyemi Sadeeq Akintola, Oluwole Fagbohun, Chukwuka Michael Oforgu, Bisola Faith Kayode, Christian Chimezie, Temitope Kadri, Abiola Oludotun, Nelson Ogbeide, Mgbame Michael, Adeseye Ifaturoti, Toyese Oloyede

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Preprint, Accepted to the International Conference on Machine Learning Technologies (ICMLT 2025), Helsinki, Finland

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13542 2025-03-19 cs.LG cs.AI 62%

HAR-DoReMi: Optimizing Data Mixture for Self-Supervised Human Activity Recognition Across Heterogeneous IMU Datasets

Lulu Ban, Tao Zhu, Xiangqing Lu, Qi Qiu, Wenyong Han, Shuangjian Li, Liming Chen, Kevin I-Kai Wang, Mingxing Nie, Yaping Wan

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13342 2025-03-18 cs.CL cs.AI 62%

Valid Text-to-SQL Generation with Unification-based DeepStochLog

Ying Jiao, Luc De Raedt, Giuseppe Marra

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Journal ref In International Conference on Neural-Symbolic Learning and Reasoning (pp. 312-330). Cham: Springer Nature Switzerland (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09199 2025-03-18 cs.LG cs.AI q-bio.BM stat.AP 62%

GENEOnet: Statistical analysis supporting explainability and trustworthiness

Giovanni Bocchi, Patrizio Frosini, Alessandra Micheletti, Alessandro Pedretti, Carmen Gratteri, Filippo Lunghini, Andrea Rosario Beccari, Carmine Talarico

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Journal ref Statistics 0 (2025), 1-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13081 2025-03-18 cs.CL cs.AI 62%

A Framework to Assess Multilingual Vulnerabilities of LLMs

Likai Tang, Niruth Bogahawatta, Yasod Ginige, Jiarui Xu, Shixuan Sun, Surangika Ranathunga, Suranga Seneviratne

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12651 2025-03-18 cs.AI cs.CL cs.HC 62%

VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures

Yoo Yeon Sung, Hannah Kim, Dan Zhang

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12085 2025-03-18 cs.AI cs.HC cs.LG 62%

Automating the loop in traffic incident management on highway

Matteo Cercola, Nicola Gatti, Pedro Huertas Leyva, Benedetto Carambia, Simone Formentin

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19801 2025-03-18 cs.SE cs.AI cs.CL 62%

CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells

Atharva Naik, Marcus Alenius, Daniel Fried, Carolyn Rose

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.03277 2025-03-17 cs.AI cs.CL cs.CV 62%

ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding

Zhengzhuo Xu, Bowen Qu, Yiyan Qi, Sinan Du, Chengjin Xu, Chun Yuan, Jian Guo

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10242 2025-03-14 cs.CL cs.AI 62%

MinorBench: A hand-built benchmark for content-based risks for children

Shaun Khoo, Gabriel Chua, Rachel Shong

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05336 2025-03-14 cs.AI cs.LG 62%

Toward an Evaluation Science for Generative AI Systems

Laura Weidinger, Inioluwa Deborah Raji, Hanna Wallach, Margaret Mitchell, Angelina Wang, Olawale Salaudeen, Rishi Bommasani, Deep Ganguli, Sanmi Koyejo, William Isaac

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments First two authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08122 2025-03-12 cs.LG cs.AI cs.CV 62%

Toward Stable World Models: Measuring and Addressing World Instability in Generative Environments

Soonwoo Kwon, Jin-Young Kim, Hyojun Go, Kyungjune Baek

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07784 2025-03-12 cs.LG cs.AI 62%

Joint Explainability-Performance Optimization With Surrogate Models for AI-Driven Edge Services

Foivos Charalampakos, Thomas Tsouparopoulos, Iordanis Koutsopoulos

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07036 2025-03-12 cs.CR cs.AI cs.LG 62%

Automated Consistency Analysis of LLMs

Aditya Patwardhan, Vivek Vaidya, Ashish Kundu

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 12 figures, 3 tables, 3 algorithms, 2024 IEEE 6th International Conference on Trust, Privacy and Security in Intelligent Systems, and Applications (TPS-ISA), Washington, DC, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06302 2025-03-11 cs.NI cs.AI cs.LG 62%

Synergizing AI and Digital Twins for Next-Generation Network Optimization, Forecasting, and Security

Zifan Zhang, Minghong Fang, Dianwei Chen, Xianfeng Yang, Yuchen Liu

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by IEEE Wireless Communications

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05856 2025-03-11 cs.CL cs.AI 62%

This Is Your Doge, If It Please You: Exploring Deception and Robustness in Mixture of LLMs

Lorenz Wolf, Sangwoong Yoon, Ilija Bogunovic

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments 35 pages, 9 figures, 16 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04283 2025-03-07 cs.LG cs.AI 62%

Explainable AI in Time-Sensitive Scenarios: Prefetched Offline Explanation Model

Fabio Michele Russo, Carlo Metta, Anna Monreale, Salvatore Rinzivillo, Fabio Pinelli

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Journal ref Discovery Science 2024, Springer Nature

详情

展开后加载摘要…

URL PDF HTML 收藏