arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8057 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8057 篇

2505.13006 2025-05-20 cs.CL 57%

Evaluating the Performance of RAG Methods for Conversational AI in the Airport Domain

Yuyang Li, Philip J. M. Kerbusch, Raimon H. R. Pruim, Tobias Käfer

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Royal Schiphol Group(皇家施比尔集团)

专题命中 其他安全 :safety(abstract);分类 cs.CL

Comments Accepted by NAACL 2025 industry track

Journal ref In Proc. NAACL-HLT 2025 Industry Track, pp. 794-808. Albuquerque, NM, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12888 2025-05-20 cs.CL 57%

GAP: Graph-Assisted Prompts for Dialogue-based Medication Recommendation

Jialun Zhong, Yanzeng Li, Sen Hu, Yang Zhang, Teng Xu, Lei Zou

机构 * Wangxuan Institute of Computer Technology, Peking University, Beijing, China(王轩计算机技术研究所,北京大学,北京,中国) Ant Group(蚂蚁集团)

专题命中 其他安全 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02306 2025-05-20 cs.AI 57%

SafeMate: A Modular RAG-Based Agent for Context-Aware Emergency Guidance

Junfeng Jiao, Jihyung Park, Yiming Xu, Kristen Sussman, Lucy Atkinson

机构 * Urban Information Lab, University of Texas at Austin(德克萨斯大学奥斯汀分校城市信息实验室) Department of Computer Science, University of Texas at Austin(德克萨斯大学奥斯汀分校计算机科学系) Department of Advertising, Texas State University(德克萨斯州立大学广告系) Department of Advertising and Public Relations, University of Texas at Austin(德克萨斯大学奥斯汀分校广告与公共关系系)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02355 2025-05-20 cs.SE cs.AI 57%

CodeGRAG: Bridging the Gap between Natural Language and Programming Language via Graphical Retrieval Augmented Generation

Kounianhua Du, Jizheng Chen, Renting Rui, Huacan Chai, Lingyue Fu, Wei Xia, Yasheng Wang, Ruiming Tang, Yong Yu, Weinan Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Huawei Noah’s Ark Lab Shanghai(华为诺亚实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12500 2025-05-20 cs.AI 57%

MARGE: Improving Math Reasoning for LLMs with Guided Exploration

Jingyue Gao, Runji Lin, Keming Lu, Bowen Yu, Junyang Lin, Jianyu Chen

机构 * Institute for Interdisciplinary Information Sciences, Tsinghua University, Beijing, China(清华大学交叉信息学院) Alibaba Group(阿里巴巴集团) Shanghai Qi Zhi Institute, Shanghai, China(上海启智研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments To appear at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11922 2025-05-20 cs.CL 57%

Enhancing Complex Instruction Following for Large Language Models with Mixture-of-Contexts Fine-tuning

Yuheng Lu, ZiMeng Bai, Caixia Yuan, Huixing Jiang, Xiaojie Wang

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(人工智能学院,北京邮电大学) LI Auto Inc.(LI汽车公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11584 2025-05-20 cs.AI 57%

LLM Agents Are Hypersensitive to Nudges

Manuel Cherep, Pattie Maes, Nikhil Singh

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 33 pages, 28 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08838 2025-05-20 eess.IV cs.AI cs.CV 57%

Ultrasound Report Generation with Multimodal Large Language Models for Standardized Texts

Peixuan Ge, Tongkun Su, Faqin Lv, Baoliang Zhao, Peng Zhang, Chi Hong Wong, Liang Yao, Yu Sun, Zenan Wang, Pak Kin Wong, Ying Hu

机构 * Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) University of Macau(澳门大学) Chinese PLA General Hospital(中国人民解放军总医院) Macau University of Science and Technology(澳门科学技术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16902 2025-05-20 cs.CV cs.AI 57%

Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement

Suchae Jeong, Inseong Choi, Youngsik Yun, Jihie Kim

机构 * Department of Computer Science and Engineering, Dongguk University(计算机科学与工程系,东国大学) Department of Computer Science and Artificial Intelligence, Dongguk University(计算机科学与人工智能系,东国大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 31 pages, 23 figures, Accepted by NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15573 2025-05-19 cs.LG 57%

Task-Specific Data Selection for Instruction Tuning via Monosemantic Neuronal Activations

Da Ma, Gonghu Shang, Zhi Chen, Libo Qin, Yijie Luo, Lei Pan, Shuai Fan, Lu Chen, Kai Yu

机构 * X-LANCE Lab, Department of Computer Science and Engineering(X-LANCE实验室,计算机科学与工程系) MoE Key Lab of Artificial Intelligence, SJTU AI Institute(人工智能MOE实验室,SJTU人工智能研究所) Shanghai Jiao Tong University(上海交通大学) AISpeech Co., Ltd.(AISpeech公司) School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments preprint, (20 pages, 7 figures, 13 tables)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02844 2025-05-19 cs.IR cs.CL 57%

Item-Language Model for Conversational Recommendation

Li Yang, Anushya Subbiah, Hardik Patel, Judith Yue Li, Yanwei Song, Reza Mirghaderi, Vikram Aggarwal, Qifan Wang

机构 * Google Research(谷歌研究)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10003 2025-05-16 cs.LG eess.SP 57%

AI2MMUM: AI-AI Oriented Multi-Modal Universal Model Leveraging Telecom Domain Large Model

Tianyu Jiao, Zhuoran Xiao, Yihang Huang, Chenhui Ye, Yijia Feng, Liyu Cai, Jiang Chang, Fangkun Liu, Yin Xu, Dazhi He, Yunfeng Guan, Wenjun Zhang

机构 * Cooperative Medianet Innovation Center, Shanghai Jiao Tong University(上海交通大学 cooperative medianet innovation center) Nokia Bell Labs(诺基亚贝尔实验室) Institute of Intelligent Communications and Network Security, Chongqing University of Posts and Telecommunications(重庆邮电大学智能通信与网络安全研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18673 2025-05-16 cs.AI cs.HC 57%

MapExplorer: New Content Generation from Low-Dimensional Visualizations

Xingjian Zhang, Ziyang Xiong, Shixuan Liu, Yutong Xie, Tolga Ergen, Dongsub Shim, Hua Xu, Honglak Lee, Qiaozhu Me

机构 * University of Michigan(密歇根大学) LG AI Research(LG人工智能研究) Yale University(耶鲁大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00646 2025-05-16 cs.CL 57%

Phase Diagram of Vision Large Language Models Inference: A Perspective from Interaction across Image and Instruction

Houjing Wei, Yuting Shi, Naoya Inoue

机构 * Japan Advanced Institute of Science and Technology(日本先进科学研究院) RIKEN(日本资源技术研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments 6 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09265 2025-05-15 cs.CV cs.AI 57%

MetaUAS: Universal Anomaly Segmentation with One-Prompt Meta-Learning

Bin-Bin Gao

机构 * Tencent YouTu Lab(腾讯优图实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18044 2025-05-15 cs.MA cs.AI 57%

Cognitive Insights and Stable Coalition Matching for Fostering Multi-Agent Cooperation

Jiaqi Shao, Tianjun Yuan, Tao Lin, Bing Luo

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07851 2025-05-14 eess.IV cs.AI cs.CV cs.RO 57%

Pose Estimation for Intra-cardiac Echocardiography Catheter via AI-Based Anatomical Understanding

Jaeyoung Huh, Ankur Kapoor, Young-Ho Kim

机构 * Digital Technology & Innovation, Siemens Healthineers(数字技术与创新,西门子医疗)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07846 2025-05-14 cs.AI cs.CR 57%

Winning at All Cost: A Small Environment for Eliciting Specification Gaming Behaviors in Large Language Models

Lars Malmqvist

机构 * The Tech Collective(技术集体)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments To be presented at SIMLA@ACNS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07457 2025-05-13 econ.GN cs.AI q-fin.EC 57%

Can Generative AI agents behave like humans? Evidence from laboratory market experiments

R. Maria del Rio-Chanona, Marco Pangallo, Cars Hommes

机构 * Computer Science Department, University College London(伦敦大学学院计算机科学系) Complexity Science Hub Bennett Institute for Public Policy, University of Cambridge(剑桥大学复杂科学枢纽贝内特公共政策研究所) CENTAI Institute(CENTAI研究所) Canadian Economic Analysis Department, Bank of Canada Faculty of Economics and Business, University of Amsterdam(荷兰阿姆斯特丹大学经济与商业学院加拿大经济分析部)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06906 2025-05-13 cs.RO cs.LG cs.SY eess.SY 57%

Realistic Counterfactual Explanations for Machine Learning-Controlled Mobile Robots using 2D LiDAR

Sindre Benjamin Remman, Anastasios M. Lekkas

机构 * Department of Engineering Cybernetics, Norwegian University of Science and Technology (NTNU)(工程 cybernetics 系,挪威科学技术大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments Accepted for publication at the 2025 European Control Conference (ECC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06817 2025-05-13 cs.AI 57%

Control Plane as a Tool: A Scalable Design Pattern for Agentic AI Systems

Sivasathivel Kandasamy

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments 2 Figures and 2 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06295 2025-05-13 cs.LG 57%

Benchmarking Traditional Machine Learning and Deep Learning Models for Fault Detection in Power Transformers

Bhuvan Saravanan, Pasanth Kumar M D, Aarnesh Vengateson

机构 * National Institute of Technology, Tiruchirappalli, TamilNadu, India(印度坦米尔纳德邦特里奇里帕利国家理工学院)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16731 2025-05-12 cs.LG cs.NE 57%

Pretraining with Random Noise for Fast and Robust Learning without Weight Transport

Jeonghwan Cheon, Sang Wan Lee, Se-Bum Paik

机构 * Department of Brain and Cognitive Sciences(脑科学与认知科学系) Graduate School of Data Science(数据科学研究生院) Kim Jaechul Graduate School of AI(金在哲人工智能研究生院)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Journal ref Advances in Neural Information Processing Systems 37, 13748-13768, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04660 2025-05-09 cs.CL cs.CV 57%

AI-Generated Fall Data: Assessing LLMs and Diffusion Model for Wearable Fall Detection

Sana Alamgeer, Yasine Souissi, Anne H. H. Ngu

机构 * Texas State University(德克萨斯州立大学) University of North Carolina(北卡罗来纳大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.10191 2025-05-09 cs.RO cs.LG 57%

Adaptive Meta-Learning for Identification of Rover-Terrain Dynamics

S. Banerjee, J. Harrison, P. M. Furlong, M. Pavone

专题命中 其他安全 :safety(abstract);分类 cs.LG

Journal ref Proc. Int. Symp. on Artificial Intelligence, Robotics and Automation in Space (iSAIRAS), 2020, Paper 5054

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03769 2025-05-08 cs.SI cs.AI cs.IR 57%

The Influence of Text Variation on User Engagement in Cross-Platform Content Sharing

Yibo Hu, Yiqiao Jin, Meng Ye, Ajay Divakaran, Srijan Kumar

机构 * Georgia Institute of Technology(佐治亚理工学院) SRI International(SRI国际)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02116 2025-05-08 cs.CL cs.CV cs.GR cs.HC 57%

Advancements and limitations of LLMs in replicating human color-word associations

Makoto Fukushima, Shusuke Eshita, Hiroshige Fukuhara

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments 20 pages, 7 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12296 2025-05-08 cs.HC cs.AI 57%

Generative Artificial Intelligence-Guided User Studies: An Application for Air Taxi Services

Shengdi Xiao, Jingjing Li, Tatsuki Fushimi, Yoichi Ochiai

机构 * Graduate School of Comprehensive Human, University of Tsukuba(大学综合人类研究生院,茨口大学) Institute of Library, Information and Media Science, University of Tsukuba(图书馆、信息与媒体科学研究所,茨口大学) R&D Center for Digital Nature, University of Tsukuba(数字自然研究开发中心,茨口大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments 39 pages, 6 main figures, 10 appendix figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03380 2025-05-07 cs.CV cs.AI eess.IV 57%

Reinforced Correlation Between Vision and Language for Precise Medical AI Assistant

Haonan Wang, Jiaji Mao, Lehan Wang, Qixiang Zhang, Marawan Elbatel, Yi Qin, Huijun Hu, Baoxun Li, Wenhui Deng, Weifeng Qin, Hongrui Li, Jialin Liang, Jun Shen, Xiaomeng Li

机构 * Department of Electronic and Computer Engineering, HKUST(香港科技大学电子与计算机工程系) Department of Radiology, Guangdong Provincial Key Laboratory of Malignant Tumor Epigenetics and Gene Regulation, Sun Yat-Sen Memorial Hospital, Sun Yat-Sen University(中山大学放射科、广东省恶性肿瘤表观遗传与基因调控重点实验室、中山纪念医院) Department of Computer Science and Engineering, HKUST(香港科技大学计算机科学与工程系)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02583 2025-05-06 cs.LG stat.ML 57%

Towards Cross-Modality Modeling for Time Series Analytics: A Survey in the LLM Era

Chenxi Liu, Shaowen Zhou, Qianxiong Xu, Hao Miao, Cheng Long, Ziyue Li, Rui Zhao

机构 * S-Lab, Nanyang Technological University, Singapore(南洋理工大学S实验室) Aalborg University, Denmark(奥尔堡大学) University of Cologne, Germany(科隆大学) SenseTime Research, China(SenseTime研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Accepted by IJCAI 2025 Survey Track

详情

展开后加载摘要…

URL PDF HTML 收藏