arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2407.01461 2025-07-01 cs.CL 57%

Enhancing the Capability and Robustness of Large Language Models through Reinforcement Learning-Driven Query Refinement

Xiaohua Wang, Zisu Huang, Feiran Zhang, Zhibo Xu, Cenyuan Zhang, Qi Qian, Xiaoqing Zheng, Xuanjing Huang

机构 * School of Computer Science, Fudan University, Shanghai, China(复旦大学计算机科学学院)

专题命中 安全评测 :jailbreak(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21881 2025-06-30 cs.CL 57%

A Dual-Layered Evaluation of Geopolitical and Cultural Bias in LLMs

Sean Kim, Hyuhng Joon Kim

机构 * Seoul National University(首尔国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments This paper is accepted to ACL Student Research Workshop (SRW) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19466 2025-06-30 cs.AI 57%

KunLunBaizeRAG: Reinforcement Learning Driven Inference Performance Leap for Large Language Models

Cheng Li, Jiexiong Liu, Yixuan Chen, Qihang Zhou, KunLun Meta

机构 * KunLun Meta(昆仑元)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11108 2025-06-30 cs.RO cs.AI 57%

Personalized Robotic Object Rearrangement from Scene Context

Kartik Ramachandruni, Sonia Chernova

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted at IEEE ROMAN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21393 2025-06-27 cs.AI 57%

TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding

Junwen Zhang, Pu Chen, Yin Zhang

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 43 pages and 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20600 2025-06-26 cs.AI 57%

CogGen: A Learner-Centered Generative AI Architecture for Intelligent Tutoring with Programming Video

Wengxi Li, Roy Pea, Nick Haber, Hari Subramonyam

机构 * City University of Hong Kong(香港城市大学) Stanford University(斯坦福大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20151 2025-06-26 cs.CV cs.AI 57%

EAR: Erasing Concepts from Unified Autoregressive Models

Haipeng Fan, Shiyuan Zhang, Baohunesitu, Zihang Guo, Huaiwen Zhang

机构 * Inner Mongolia University(内蒙古大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 11 pages, 7 figures, 1 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10803 2025-06-26 cs.LG cs.ET quant-ph 57%

Quantum Kernel Learning for Small Dataset Modeling in Semiconductor Fabrication: Application to Ohmic Contact

Zeheng Wang, Fangzhou Wang, Liang Li, Zirui Wang, Timothy van der Laan, Ross C. C. Leon, Jing-Kai Huang, Muhammad Usman

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Journal version 2.0

Journal ref Adv. Sci. 2025, e06213

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19665 2025-06-25 cs.CV cs.CL 57%

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation

Yuanhe Tian, Lei Mao, Yan Song

机构 * University of Washington(华盛顿大学) Origin Omics University of Science and Technology of China(中国科学技术大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19607 2025-06-25 cs.CL 57%

Correcting Hallucinations in News Summaries: Exploration of Self-Correcting LLM Methods with External Knowledge

Juraj Vladika, Ihsan Soydemir, Florian Matthes

机构 * Technical University of Munich School of Computation, Information and Technology Department of Computer Science(慕尼黑技术大学计算、信息与技术学院计算机科学系)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted to FEVER @ ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19109 2025-06-25 cs.CR cs.AI 57%

Enhancing Security in LLM Applications: A Performance Evaluation of Early Detection Systems

Valerii Gakh, Hayretdin Bahsi

机构 * School of Information Technologies(信息科技学院) Tallinn University of Technology(塔林技术大学) School of Informatics, Computing, and Cyber Systems(信息学、计算与网络安全学院) Northern Arizona University(北亚利桑那大学)

专题命中 安全评测 :prompt injection(abstract);分类 cs.AI

Comments 18 pages, 8 tables, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00466 2025-06-25 cs.LG cs.LO 57%

A General Framework for Property-Driven Machine Learning

Thomas Flinkow, Marco Casadio, Colin Kessler, Rosemary Monahan, Ekaterina Komendantskaya

机构 * Maynooth University(梅诺特大学) Heriot-Watt University(赫瑞-瓦特大学) University of Edinburgh(爱丁堡大学) Edinburgh Centre for Robotics(爱丁堡机器人中心) University of Southampton(南安普顿大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 24 pages, 4 tables, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13816 2025-06-25 cs.CL 57%

Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal Representations

Chenghao Xiao, Hou Pong Chan, Hao Zhang, Mahani Aljunied, Lidong Bing, Noura Al Moubayed, Yu Rong

机构 * DAMO Academy, Alibaba Group(阿里巴巴集团达摩院) Department of Computer Science, Durham University(杜伦大学计算机科学系) Hupan Lab, Hangzhou, China(杭州实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments ACL 2025 main; camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18781 2025-06-24 cs.CL 57%

Existing LLMs Are Not Self-Consistent For Simple Tasks

Zhenru Lin, Jiawen Tao, Yang Yuan, Andrew Chi-Chih Yao

机构 * IIIS, Tsinghua University(清华大学人工智能学院) Shanghai AI Laboratory(上海人工智能实验室) Shanghai Qizhi Institute(上海启智研究院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16301 2025-06-24 cs.CV cs.LG 57%

DiffDesign: Controllable Diffusion with Meta Prior for Efficient Interior Design Generation

Yuxuan Yang, Tao Geng

机构 * Nanjing Forestry University(南京林业大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18538 2025-06-24 cs.AI 57%

A Question Bank to Assess AI Inclusivity: Mapping out the Journey from Diversity Errors to Inclusion Excellence

Rifat Ara Shams, Didar Zowghi, Muneera Bano

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18102 2025-06-24 cs.CL 57%

InspireDebate: Multi-Dimensional Subjective-Objective Evaluation-Guided Reasoning and Optimization for Debating

Fuyu Wang, Jiangtong Li, Kun Zhu, Changjun Jiang

机构 * Key Laboratory of Embedded System and Service Computing, Ministry of Education, Tongji University(嵌入式系统与服务计算重点实验室,教育部,同济大学) School of Computer Science and Technology, Tongji University(计算机科学与技术学院,同济大学)

专题命中 安全评测 :DPO(abstract);分类 cs.CL

Comments 20 pages; Accepted to ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17863 2025-06-24 cs.CL 57%

LLMs for Customized Marketing Content Generation and Evaluation at Scale

Haoran Liu, Amir Tahmasbi, Ehtesham Sam Haque, Purak Jain

机构 * Texas A\&M University(德克萨斯大学) Amazon Inc.(亚马逊公司)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments KDD LLM4ECommerce Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17601 2025-06-24 cs.RO cs.AI 57%

Risk-Guided Diffusion: Toward Deploying Robot Foundation Models in Space, Where Failure Is Not An Option

Rohan Thakker, Adarsh Patnaik, Vince Kurtz, Jonas Frey, Jonathan Becktor, Sangwoo Moon, Rob Royce, Marcel Kaufmann, Georgios Georgakis, Pascal Roth, Joel Burdick, Marco Hutter, Shehryar Khattak

机构 * NASA-JPL, Caltech(美国国家航空航天局喷气推进实验室,加州理工学院) CME, Caltech(加州理工学院计算机工程系) Robotic Systems Lab, ETH Zurich(苏黎世联邦理工学院机器人系统实验室)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Journal ref Robotics Science and Systems 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02000 2025-06-24 cs.CL 57%

NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts

Abhay Gupta, Michael Lu, Kevin Zhu, Sean O'Brien, Vasu Sharma

机构 * Algoverse AI Research(Algoverse AI研究机构) University of California, Berkeley(加州大学伯克利分校) Meta FAIR Lab(Meta FAIR实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03424 2025-06-24 cs.LG 57%

Prediction of the Most Fire-Sensitive Point in Building Structures with Differentiable Agents for Thermal Simulators

Yuan Xinjie, Khalid M. Mosalam

机构 * Shenzhen International Graduate School Tsinghua University(深圳国际研究生院清华大学) Pacific Earthquake Engineering Research (PEER) Center University of California, Berkeley(太平洋地震工程研究中心加州大学伯克利分校) University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments This paper has been accepted by journal Computer-Aided Civil and Infrastructure Engineering

Journal ref Computer-Aided Civil and Infrastructure Engineering: Volume 40, Issue 18. 2025 July

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09174 2025-06-24 cs.CV cs.AI 57%

DART: An Automated End-to-End Object Detection Pipeline with Data Diversification, Open-Vocabulary Bounding Box Annotation, Pseudo-Label Review, and Model Training

Chen Xin, Andreas Hartel, Enkelejda Kasneci

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Corrected minor typos; no changes to results or conclusions

Journal ref Expert Systems with Applications 258 (2024): 125124

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16718 2025-06-23 cs.MA cs.AI 57%

Generalizable Agent Modeling for Agent Collaboration-Competition Adaptation with Multi-Retrieval and Dynamic Generation

Chenxu Wang, Yonggang Jin, Cheng Hu, Youpeng Zhao, Zipeng Dai, Jian Zhao, Shiyu Huang, Liuyu Xiang, Junge Zhang, Zhaofeng He

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Institute of Technology(北京理工大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Tsinghua University(清华大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments This manuscript is under submission to Neurocomputing

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16416 2025-06-23 stat.ML cs.LG 57%

On Continuous Monitoring of Risk Violations under Unknown Shift

Alexander Timans, Rajeev Verma, Eric Nalisnick, Christian A. Naesseth

机构 * UvA-Bosch Delta Lab, University of Amsterdam(UvA-Bosch Delta Lab,阿姆斯特丹大学) University of Amsterdam(阿姆斯特丹大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments AT and RV are joint first authors. Accepted at the Conference on Uncertainty in Artificial Intelligence (UAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13987 2025-06-23 cs.LG 57%

Quantum-Informed Contrastive Learning with Dynamic Mixup Augmentation for Class-Imbalanced Expert Systems

Md Abrar Jahin, Adiba Abid, M. F. Mridha

机构 * organization= Thomas Lord Department of Computer Science, Viterbi School of Engineering, University of Southern California , city= Los Angeles , state= CA , postcode= 90089 , country= USA organization= Department of Industrial Engineering Management, Khulna University of Engineering \& Technology (KUET) , city= Khulna , postcode= 9203 , country= Bangladesh organization= Department of Computer Science, American International University-Bangladesh (AIUB) , city= Dhaka , postcode= 1229 , country= Bangladesh

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06967 2025-06-23 cs.CV cs.AI eess.IV 57%

Dual Thinking and Logical Processing -- Are Multi-modal Large Language Models Closing the Gap with Human Vision ?

Kailas Dayanandan, Nikhil Kumar, Anand Sinha, Brejesh Lall

机构 * Indian Institute of Technology Delhi(印度理工学院德里)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16487 2025-06-23 cs.LG math.OC 57%

On the Robustness of Decision-Focused Learning

Yehya Farhat

机构 * Yehya Farhat

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 17 pages, 45 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15425 2025-06-19 cs.CL 57%

Understanding GUI Agent Localization Biases through Logit Sharpness

Xingjian Tao, Yiwei Wang, Yujun Cai, Zhicheng Yang, Jing Tang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学) University of California, Merced(加州大学梅德福分校) The University of Queensland(昆士兰大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14853 2025-06-19 q-bio.QM cs.LG 57%

DisProtEdit: Exploring Disentangled Representations for Multi-Attribute Protein Editing

Max Ku, Sun Sun, Hongyu Guo, Wenhu Chen

机构 * National Research Council Canada(加拿大国家研究理事会) Cheriton School of Computer Science(计算机科学学院)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Accepted to ICMLW (GenBio) 2025 and ICMLW (FM4LS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14827 2025-06-19 cs.CV cs.AI 57%

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning

Yifeng Gao, Yifan Ding, Hongyu Su, Juncheng Li, Yunhan Zhao, Lin Luo, Zixing Chen, Li Wang, Xin Wang, Yixu Wang, Xingjun Ma, Yu-Gang Jiang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏