arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8057 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8057 篇

2411.15539 2025-05-06 cs.CV cs.AI 57%

Large Language Model with Region-guided Referring and Grounding for CT Report Generation

Zhixuan Chen, Yequan Bie, Haibo Jin, Hao Chen

机构 * Hong Kong University of Science and Technology(香港科技大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12404 2025-05-06 cs.LG 57%

Analyzing the Generalization and Reliability of Steering Vectors

Daniel Tan, David Chanin, Aengus Lynch, Dimitrios Kanoulas, Brooks Paige, Adria Garriga-Alonso, Robert Kirk

机构 * AI Centre, Department of Computer Science, University College London(人工智能中心、计算机科学系、伦敦大学学院)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01488 2025-05-06 cs.LG cs.CR 57%

Explainable Machine Learning for Cyberattack Identification from Traffic Flows

Yujing Zhou, Marc L. Jacquet, Robel Dawit, Skyler Fabre, Dev Sarawat, Faheem Khan, Madison Newell, Yongxin Liu, Dahai Liu, Hongyun Chen, Jian Wang, Huihui Wang

机构 * Embry-Riddle Aeronautical University(埃姆布里-瑞尔德航空航天大学) University of Tennessee at Martin(田纳西大学马丁分校) Northeastern University(东北大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00968 2025-05-06 cs.CV cs.LG 57%

CoDe: Blockwise Control for Denoising Diffusion Models

Anuj Singh, Sayak Mukherjee, Ahmad Beirami, Hadi Jamali-Rad

机构 * Delft University of Technology(代尔夫特理工大学) Shell Global Solutions International B.V.(壳牌全球解决方案国际有限公司) Massachusetts Institute of Technology(麻省理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Journal ref Transactions on Machine Learning Research, 2025. ISSN: 2835-8856

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00690 2025-05-02 cs.CV cs.AI cs.RO 57%

Towards Autonomous Micromobility through Scalable Urban Simulation

Wayne Wu, Honglin He, Chaoyuan Zhang, Jack He, Seth Z. Zhao, Ran Gong, Quanyi Li, Bolei Zhou

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments CVPR 2025 Highlight. Project page: https://metadriverse.github.io/urban-sim/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00603 2025-05-02 cs.AI cs.HC 57%

Can LLMs Help Improve Analogical Reasoning For Strategic Decisions? Experimental Evidence from Humans and GPT-4

Phanish Puranam, Prothit Sen, Maciej Workiewicz

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00389 2025-05-02 cs.CL 57%

CSE-SFP: Enabling Unsupervised Sentence Representation Learning via a Single Forward Pass

Bowen Zhang, Zixin Song, Chunping Li

机构 * Tsinghua University(清华大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted by SIGIR 2025 (Full)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00196 2025-05-02 cs.LG q-bio.NC 57%

Mapping minds not averages: a scalable subject-specific manifold learning framework for neuroimaging data

Eloy Geenjaar, Vince Calhoun

机构 * Electrical and Computer Engineering, Georgia Institute of Technology(电子与计算机工程系,佐治亚理工学院) (TReNDS) Translational Research in Neuroimaging & Data Science center, Georgia State University, Georgia Institute of Technology, & Emory University((TReNDS) 神经影像与数据科学转化研究中心,佐治亚州立大学,佐治亚理工学院,及埃默里大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments 20 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20340 2025-04-30 cs.AI cs.CV cs.HC 57%

A Picture is Worth a Thousand Prompts? Efficacy of Iterative Human-Driven Prompt Refinement in Image Regeneration Tasks

Khoi Trinh, Scott Seidenberger, Raveen Wijewickrama, Murtuza Jadliwala, Anindya Maiti

机构 * University of Oklahoma(俄克拉荷马大学) University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17365 2025-04-30 cs.CV cs.CL 57%

TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation

Ling You, Wenxuan Huang, Xinni Xie, Xiangyi Wei, Bangyan Li, Shaohui Lin, Yang Li, Changbo Wang

机构 * East China Normal University(东华大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16940 2025-04-29 q-bio.NC cs.AI cs.CV 57%

Better artificial intelligence does not mean better models of biology

Drew Linsley, Pinyuan Feng, Thomas Serre

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18849 2025-04-29 cs.LG eess.IV 57%

Theoretical Framework for Tempered Fractional Gradient Descent: Application to Breast Cancer Classification

Omar Naifar

机构 * Control and Energy Management Laboratory, National School of Engineering, University of Sfax(工程控制与能源管理实验室,国家工程学院,突尼斯苏菲夫大学) Higher Institute of Applied Sciences and Technology of Kairouan, University of Kairouan(凯鲁安应用科学和技术高等学院,凯鲁安大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18628 2025-04-29 cs.AR cs.LG 57%

Periodic Online Testing for Sparse Systolic Tensor Arrays

Christodoulos Peltekis, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos

机构 * Electrical and Computer Engineering Democritus University of Thrace, Greece(电子与计算机工程系德米特里乌斯大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments International Conference on Modern Circuits and Systems Technologies (MOCAST) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13926 2025-04-29 cs.HC cs.AI 57%

A Multi-Layered Research Framework for Human-Centered AI: Defining the Path to Explainability and Trust

Chameera De Silva, Thilina Halloluwa, Dhaval Vyas

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments I am requesting this withdrawal because I believe the current version requires significant revisions and restructuring to better reflect the intended research contributions. I plan to substantially improve the work and may resubmit a revised version in the future. Thank you for your understanding and support

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11419 2025-04-29 cs.AI cs.NE 57%

Embodied World Models Emerge from Navigational Task in Open-Ended Environments

Li Jin, Liu Jia

机构 * Tsinghua Laboratory of Brain and Intelligence(清华大学脑科学与智能实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Research on explainable meta-reinforcement learning AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13019 2025-04-29 cs.CL 57%

Oreo: A Plug-in Context Reconstructor to Enhance Retrieval-Augmented Generation

Sha Li, Naren Ramakrishnan

机构 * Virginia Tech(弗吉尼亚理工大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04107 2025-04-29 cs.IR cs.AI 57%

Pre-train, Align, and Disentangle: Empowering Sequential Recommendation with Large Language Models

Yuhao Wang, Junwei Pan, Pengyue Jia, Wanyu Wang, Maolin Wang, Zhixiang Feng, Xiaotian Li, Jie Jiang, Xiangyu Zhao

机构 * City University of Hong Kong(香港城市大学) Tencent Inc.(腾讯公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments accepted to SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07149 2025-04-29 cs.CV cs.LG 57%

Towards Interpreting Visual Information Processing in Vision-Language Models

Clement Neo, Luke Ong, Philip Torr, Mor Geva, David Krueger, Fazl Barez

机构 * Nanyang Technological University(南洋理工大学) University of Oxford(牛津大学) Tel Aviv University(特拉维夫大学) MILA(蒙特利尔人工智能研究院) ERA-Krueger AI Safety Lab(ERA-Krueger人工智能安全实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Published at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18412 2025-04-28 cs.CL 57%

Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers

Jared Moore, Declan Grabb, William Agnew, Kevin Klyman, Stevie Chancellor, Desmond C. Ong, Nick Haber

机构 * Stanford University(斯坦福大学) Carnegie Mellon University(卡内基梅隆大学) University of Minnesota(明尼苏达大学) University of Texas(德克萨斯大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17448 2025-04-25 cs.LG cs.DB cs.DC 57%

CHASe: Client Heterogeneity-Aware Data Selection for Effective Federated Active Learning

Jun Zhang, Jue Wang, Huan Li, Zhongle Xie, Ke Chen, Lidan Shou

机构 * State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Accepted by TKDE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17025 2025-04-25 cs.CL 57%

Optimizing LLMs for Italian: Reducing Token Fertility and Enhancing Efficiency Through Vocabulary Adaptation

Luca Moroni, Giovanni Puccetti, Pere-Lluis Huguet Cabot, Andrei Stefan Bejgu, Edoardo Barba, Alessio Miaschi, Felice Dell'Orletta, Andrea Esuli, Roberto Navigli

机构 * Babelscape(Babelscape公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19572 2025-04-24 cs.CL 57%

ChunkRAG: Novel LLM-Chunk Filtering Method for RAG Systems

Ishneet Sukhvinder Singh, Ritvik Aggarwal, Ibrahim Allahverdiyev, Muhammad Taha, Aslihan Akalin, Kevin Zhu, Sean O'Brien

机构 * Algoverse AI Research(Algoverse AI研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted at Conference of the North American Chapter of the Association for Computational Linguistics, Student Research Workshop 2025 (NAACL SRW 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14320 2025-04-23 cs.HC cs.AI 57%

Expanding the Generative AI Design Space through Structured Prompting and Multimodal Interfaces

Nimisha Karnatak, Adrien Baranes, Rob Marchant, Huinan Zeng, Tríona Butler, Kristen Olson

机构 * University of Oxford(牛津大学) Google DeepMind(谷歌DeepMind) King’s College London(伦敦大学学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accepted at CHI'25 Workshop on Designing and Developing User Interfaces with AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15848 2025-04-23 cs.CL 57%

Exploring Cognitive and Aesthetic Causality for Multimodal Aspect-Based Sentiment Analysis

Luwei Xiao, Rui Mao, Shuai Zhao, Qika Lin, Yanhao Jia, Liang He, Erik Cambria

机构 * School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) Saw Swee Hock School of Public Health, National University of Singapore(新加坡国立大学 Saw Swee Hock 公共卫生学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted by TAFFC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15552 2025-04-23 cs.AI 57%

A Multi-Agent Framework for Automated Qinqiang Opera Script Generation Using Large Language Models

Gengxian Cao, Fengyuan Li, Hong Duan, Ye Yang, Bofeng Wang, Donghe Li

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 17 pages,7 figures,1 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01801 2025-04-23 cs.CL 57%

Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training

Zhijun Wang, Jiahuan Li, Hao Zhou, Rongxiang Weng, Jingang Wang, Xin Huang, Xue Han, Junlan Feng, Chao Deng, Shujian Huang

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14955 2025-04-22 cs.LG 57%

Efficient Document Retrieval with G-Retriever

Manthankumar Solanki

机构 * University of Stuttgart(斯图加特大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Extended version of a paper presented at NeurIPS 2024 (arXiv:2402.07630)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14782 2025-04-22 cs.LG cond-mat.mtrl-sci 57%

Novel Concept-Oriented Synthetic Data approach for Training Generative AI-Driven Crystal Grain Analysis Using Diffusion Model

Ahmed Sobhi Saleh, Kristof Croes, Hajdin Ceric, Ingrid De Wolf, Houman Zahedmanesh

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments 19 Pages, 5 Figures

Journal ref Computational Materials Science, Vol 251 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01019 2025-04-22 cs.CV cs.AI 57%

MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations

Ziyang Zhang, Yang Yu, Yucheng Chen, Xulei Yang, Si Yong Yeo

机构 * MedVisAI Lab(MedVisAI实验室) ECE, Northwestern University(电子工程系,西北大学) Institute for Infocomm Research (I 2 R), A*STAR, Singapore(信息与通信研究所(I2R),A*STAR,新加坡) Lee Kong Chian School of Medicine, Nanyang Technological University(Lee Kong Chian医学院,南洋理工大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments To be pubilshed in CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17496 2025-04-18 cs.AI cs.SY eess.SY 57%

SemML: Enhancing Automata-Theoretic LTL Synthesis with Machine Learning

Jan Kretinsky, Tobias Meggendorfer, Maximilian Prokop, Ashkan Zarkhah

机构 * Masaryk University(马萨里克大学) Technical University of Munich(慕尼黑工业大学) Lancaster University Leipzig(莱斯特大学莱比锡分校)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏