arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9380 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9380 篇

2501.00208 2025-01-03 cs.CL cs.AI 62%

An Empirical Evaluation of Large Language Models on Consumer Health Questions

Moaiz Abrar, Yusuf Sermet, Ibrahim Demir

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00059 2025-01-03 cs.CL cs.AI 62%

Large Language Models for Mathematical Analysis

Ziye Chen, Hao Qi

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19906 2024-12-31 cs.CL cs.AI 62%

Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM

Dong Yuan, Eti Rastogi, Fen Zhao, Sagar Goyal, Gautam Naik, Sree Prasanna Rajagopal

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15504 2024-12-31 cs.CL cs.LG 62%

Multi-View Empowered Structural Graph Wordification for Language Models

Zipeng Liu, Likang Wu, Ming He, Zhong Guan, Hongke Zhao, Nan Feng

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16964 2024-12-25 cs.AI cs.CL 62%

System-2 Mathematical Reasoning via Enriched Instruction Tuning

Huanqia Cai, Yijun Yang, Zhifeng Li

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08297 2024-12-25 cs.LG cs.AI cs.CV cs.DC 62%

Distance-Restricted Explanations: Theoretical Underpinnings & Efficient Implementation

Yacine Izza, Xuanxiang Huang, Antonio Morgado, Jordi Planes, Alexey Ignatiev, Joao Marques-Silva

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16341 2024-12-24 cs.LG cs.CL 62%

A Machine Learning Approach for Emergency Detection in Medical Scenarios Using Large Language Models

Ferit Akaybicen, Aaron Cummings, Lota Iwuagwu, Xinyue Zhang, Modupe Adewuyi

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16176 2024-12-24 cs.SD cs.CL cs.LG eess.AS 62%

Efficient VoIP Communications through LLM-based Real-Time Speech Reconstruction and Call Prioritization for Emergency Services

Danush Venkateshperumal, Rahman Abdul Rafi, Shakil Ahmed, Ashfaq Khokhar

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments 15 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22690 2024-12-24 cs.LG cs.AI 62%

Choice Between Partial Trajectories: Disentangling Goals from Beliefs

Henrik Marklund, Benjamin Van Roy

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16751 2024-12-19 cs.AI cs.CL cs.CV cs.MA 62%

REVECA: Adaptive Planning and Trajectory-based Validation in Cooperative Language Agents using Information Relevance and Relative Proximity

SeungWon Seo, SeongRae Noh, Junhyeok Lee, SooBin Lim, Won Hee Lee, HyeongYeop Kang

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments v2 is the AAAI'25 camera-ready version, including the appendix, which has been enhanced based on the reviewers' comments

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13023 2024-12-18 cs.AI cs.LG 62%

Relational Neurosymbolic Markov Models

Lennert De Smet, Gabriele Venturato, Luc De Raedt, Giuseppe Marra

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18580 2024-12-18 cs.AI cs.LG 62%

Artificial Intelligence in Industry 4.0: A Review of Integration Challenges for Industrial Systems

Alexander Windmann, Philipp Wittenberg, Marvin Schieseck, Oliver Niggemann

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 17 pages, 4 figures, 1 table

Journal ref 2024 IEEE 22nd International Conference on Industrial Informatics (INDIN)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10535 2024-12-17 cs.CL cs.AI 62%

On Adversarial Robustness and Out-of-Distribution Robustness of Large Language Models

April Yang, Jordan Tab, Parth Shah, Paul Kotchavong

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09269 2024-12-13 cs.CL cs.AI 62%

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations

Manav Chaudhary, Harshit Gupta, Savita Bhat, Vasudeva Varma

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at ICON 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07979 2024-12-12 cs.LG cs.AI cs.CV 62%

AmCLR: Unified Augmented Learning for Cross-Modal Representations

Ajay Jagannath, Aayush Upadhyay, Anant Mehta

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 16 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06176 2024-12-10 cs.LG cs.AI 62%

AlphaVerus: Bootstrapping Formally Verified Code Generation through Self-Improving Translation and Treefinement

Pranjal Aggarwal, Bryan Parno, Sean Welleck

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01432 2024-12-10 cs.CL cs.AI 62%

Split and Merge: Aligning Position Biases in LLM-based Evaluators

Zongjie Li, Chaozheng Wang, Pingchuan Ma, Daoyuan Wu, Shuai Wang, Cuiyun Gao, Yang Liu

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by EMNLP 2024. Please cite the conference version of this paper, e.g., "Zongjie Li, Chaozheng Wang, Pingchuan Ma, Daoyuan Wu, Shuai Wang, Cuiyun Gao, and Yang Liu. 2024. Split and Merge: Aligning Position Biases in LLM-based Evaluators. (EMNLP 2024)"

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16391 2024-12-10 cs.CL cs.AI 62%

Human-Calibrated Automated Testing and Validation of Generative Language Models

Agus Sudjianto, Aijun Zhang, Srinivas Neppalli, Tarun Joshi, Michal Malohlava

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03712 2024-12-10 cs.CL cs.LG 62%

A Survey on Medical Large Language Models: Technology, Application, Trustworthiness, and Future Directions

Lei Liu, Xiaoyan Yang, Junchi Lei, Yue Shen, Jian Wang, Peng Wei, Zhixuan Chu, Zhan Qin, Kui Ren

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10980 2024-12-10 physics.chem-ph cs.AI cs.CE cs.LG 62%

ChemReasoner: Heuristic Search over a Large Language Model's Knowledge Space using Quantum-Chemical Feedback

Henry W. Sprueill, Carl Edwards, Khushbu Agarwal, Mariefel V. Olarte, Udishnu Sanyal, Conrad Johnston, Hongbin Liu, Heng Ji, Sutanay Choudhury

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 9 pages, accepted by ICML 2024, final version

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17671 2024-12-10 cs.CL cs.AI q-bio.NC 62%

Contextual Feature Extraction Hierarchies Converge in Large Language Models and the Brain

Gavin Mischler, Yinghao Aaron Li, Stephan Bickel, Ashesh D. Mehta, Nima Mesgarani

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 19 pages, 5 figures and 4 supplementary figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03682 2024-12-06 cs.LG cs.AI cs.AR cs.CV eess.IV 62%

Designing DNNs for a trade-off between robustness and processing performance in embedded devices

Jon Gutiérrez-Zaballa, Koldo Basterretxea, Javier Echanobe

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Journal ref 2024 39th Conference on Design of Circuits and Integrated Systems (DCIS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03630 2024-12-06 cs.CV cs.AI cs.AR cs.LG eess.IV 62%

Evaluating Single Event Upsets in Deep Neural Networks for Semantic Segmentation: an embedded system perspective

Jon Gutiérrez-Zaballa, Koldo Basterretxea, Javier Echanobe

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Journal ref 2024 Journal of Systems Architecture (JSA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00741 2024-12-06 cs.CL cs.AI 62%

ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Junjie Ye, Guanyu Li, Songyang Gao, Caishuang Huang, Yilong Wu, Sixian Li, Xiaoran Fan, Shihan Dou, Tao Ji, Qi Zhang, Tao Gui, Xuanjing Huang

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by COLING 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02875 2024-12-05 cs.LG cs.AI cs.CR 62%

Out-of-Distribution Detection for Neurosymbolic Autonomous Cyber Agents

Ankita Samaddar, Nicholas Potteiger, Xenofon Koutsoukos

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 9 pages, 10 figures, IEEE International Conference on AI in Cybersecurity (ICAIC), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12058 2024-12-05 cs.AI cs.CL 62%

WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions

Seyedali Mohammadi, Edward Raff, Jinendra Malekar, Vedant Palit, Francis Ferraro, Manas Gaur

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted in BlackboxNLP @ EMNLP 2024

Journal ref Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, pages 364-388, November 2024, Miami, Florida, US. Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19668 2024-12-02 cs.CL cs.AI 62%

ChineseWebText 2.0: Large-Scale High-quality Chinese Web Text with Multi-dimensional and fine-grained information

Wanyue Zhang, Ziyong Li, Wen Yang, Chunlin Leng, Yinan Bai, Qianlong Du, Chengqing Zong, Jiajun Zhang

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments ChineseWebTex2.0 dataset is available at https://github.com/CASIA-LM/ChineseWebText-2.0

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13139 2024-11-27 cs.LG cs.AI 62%

Explainable AI for Fair Sepsis Mortality Predictive Model

Chia-Hsuan Chang, Xiaoyang Wang, Christopher C. Yang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted to the 22nd International Conference on Artificial Intelligence in Medicine (AIME'24)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15560 2024-11-27 cs.AI cs.CL 62%

Do LLMs Agree on the Creativity Evaluation of Alternative Uses?

Abdullah Al Rabeyah, Fabrício Góes, Marco Volpe, Talles Medeiros

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 19 pages, 7 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16422 2024-11-26 cs.LG cs.AI eess.SP 62%

Turbofan Engine Remaining Useful Life (RUL) Prediction Based on Bi-Directional Long Short-Term Memory (BLSTM)

Abedin Sherifi

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏