arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8044 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8044 篇

2205.12961 2025-03-18 cs.LG cs.AI 62%

Position: Tensor Networks are a Valuable Asset for Green AI

Eva Memmel, Clara Menzen, Jetze Schuurmans, Frederiek Wesel, Kim Batselier

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments This paper has been accepted for presentation at the International Conference on Machine Learning (ICML) 2024 and will appear in the conference proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09638 2025-03-14 cs.RO cs.AI cs.LG 62%

Edge AI-Powered Real-Time Decision-Making for Autonomous Vehicles in Adverse Weather Conditions

Milad Rahmati

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06729 2025-03-11 cs.HC cs.AI cs.CY cs.ET 62%

ACAI for SBOs: AI Co-creation for Advertising and Inspiration for Small Business Owners

Nimisha Karnatak, Adrien Baranes, Rob Marchant, Triona Butler, Kristen Olson

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04684 2025-03-11 cs.LG cs.AI 62%

G2PDiffusion: Cross-Species Genotype-to-Phenotype Prediction via Evolutionary Diffusion

Mengdi Liu, Zhangyang Gao, Hong Chang, Stan Z. Li, Shiguang Shan, Xilin Chen

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05966 2025-03-11 cs.LG cs.AI 62%

FLOPS: Forward Learning with OPtimal Sampling

Tao Ren, Zishi Zhang, Jinyang Jiang, Guanghao Li, Zeliang Zhang, Mingqian Feng, Yijie Peng

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Published in the Thirteenth International Conference on Learning Representations(ICLR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05870 2025-03-11 cs.CR cs.CL cs.LG 62%

Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents

Avital Shafran, Roei Schuster, Vitaly Shmatikov

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.LG

Comments To appear in USENIX Security Symposium 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04780 2025-03-10 cs.CL cs.AI physics.atom-ph 62%

MV-CLAM: Multi-View Molecular Interpretation with Cross-Modal Projection via Language Model

Sumin Ha, Jun Hyeong Kim, Yinhua Piao, Sun Kim

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07335 2025-03-10 cs.LG cs.AI 62%

TempoGPT: Enhancing Time Series Reasoning via Quantizing Embedding

Haochuan Zhang, Chunhua Yang, Jie Han, Liyang Qin, Xiaoli Wang

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12360 2025-03-07 cs.CV cs.AI cs.LG 62%

Detecting Systematic Weaknesses in Vision Models along Predefined Human-Understandable Dimensions

Sujan Sai Gannamaneni, Rohil Prakash Rao, Michael Mock, Maram Akila, Stefan Wrobel

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10877 2025-03-07 cs.CL cs.AI 62%

Improving Data Efficiency via Curating LLM-Driven Rating Systems

Jinlong Pang, Jiaheng Wei, Ankit Parag Shah, Zhaowei Zhu, Yaxuan Wang, Chen Qian, Yang Liu, Yujia Bao, Wei Wei

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12106 2025-03-07 cs.CL cs.AI 62%

Measuring Human and AI Values Based on Generative Psychometrics with Large Language Models

Haoran Ye, Yuhang Xie, Yuanyi Ren, Hanjun Fang, Xin Zhang, Guojie Song

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17671 2025-03-06 cs.LG cs.AI 62%

Transfer of Reinforcement Learning-Based Controllers from Model- to Hardware-in-the-Loop

Mario Picerno, Lucas Koch, Kevin Badalian, Marius Wegener, Joschka Schaub, Charles Robert Koch, Jakob Andert

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Journal ref IEEE Transactions on Vehicular Technology (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05316 2025-03-06 cs.LG cs.AI cs.CE q-bio.BM 62%

Aligning Large Language Models and Geometric Deep Models for Protein Representation

Dong Shu, Bingbing Duan, Kai Guo, Kaixiong Zhou, Jiliang Tang, Mengnan Du

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 37 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17037 2025-03-04 cs.CY cs.AI cs.HC 62%

Standardised schema and taxonomy for AI incident databases in critical digital infrastructure

Avinash Agarwal, Manisha J. Nene

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY

Comments 6 pages, 3 tables. Accepted at the 2024 IEEE Pune Section International Conference (PuneCon)

Journal ref IEEE Pune Section International Conference (PuneCon), Pune, India, 2024, pp. 1-6

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13213 2025-03-04 cs.AI cs.LG 62%

LLMOPT: Learning to Define and Solve General Optimization Problems from Scratch

Caigao Jiang, Xiang Shu, Hong Qian, Xingyu Lu, Jun Zhou, Aimin Zhou, Yang Yu

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15998 2025-03-04 cs.CV cs.AI cs.LG cs.RO 62%

Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Min Shi, Fuxiao Liu, Shihao Wang, Shijia Liao, Subhashree Radhakrishnan, Yilin Zhao, De-An Huang, Hongxu Yin, Karan Sapra, Yaser Yacoob, Humphrey Shi, Bryan Catanzaro, Andrew Tao, Jan Kautz, Zhiding Yu, Guilin Liu

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Github: https://github.com/NVlabs/Eagle, HuggingFace: https://huggingface.co/NVEagle

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00177 2025-03-04 cs.LG cs.AI 62%

Steering Large Language Model Activations in Sparse Spaces

Reza Bayat, Ali Rahimi-Kalahroudi, Mohammad Pezeshki, Sarath Chandar, Pascal Vincent

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12373 2025-03-03 cs.LG cs.AI 62%

Cell-ontology guided transcriptome foundation model

Xinyu Yuan, Zhihao Zhan, Zuobai Zhang, Manqi Zhou, Jianan Zhao, Boyu Han, Yue Li, Jian Tang

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to NeurIPS 2024 as Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20186 2025-02-28 cs.CL cs.LG 62%

Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge

Yan-Lun Chen, Yi-Ru Wei, Chia-Yi Hsu, Chia-Mu Yu, Chun-Ying Huang, Ying-Dar Lin, Yu-Sung Wu, Wei-Bin Lee

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19347 2025-02-27 cs.CL cs.AI 62%

Controlled Diversity: Length-optimized Natural Language Generation

Diana Marie Schenke, Timo Baumann

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ISCA/ITG Workshop on Diversity in Large Speech and Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15766 2025-02-27 cs.LG cs.CL 62%

Learning Harmonized Representations for Speculative Sampling

Lefan Zhang, Xiaodan Wang, Yanhua Huang, Ruiwen Xu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Published as a conference paper at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17900 2025-02-26 cs.LG cs.AI 62%

Knowledge-enhanced Multimodal ECG Representation Learning with Arbitrary-Lead Inputs

Che Liu, Cheng Ouyang, Zhongwei Wan, Haozhe Wang, Wenjia Bai, Rossella Arcucci

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03777 2025-02-26 cs.CL cs.AI 62%

Determine-Then-Ensemble: Necessity of Top-k Union for Large Language Model Ensembling

Yuxuan Yao, Han Wu, Mingyang Liu, Sichun Luo, Xiongwei Han, Jie Liu, Zhijiang Guo, Linqi Song

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08038 2025-02-25 cs.LG cs.CL cs.SI 62%

Bootstrapping Heterogeneous Graph Representation Learning via Large Language Models: A Generalized Approach

Hang Gao, Chenhao Zhang, Fengge Wu, Junsuo Zhao, Changwen Zheng, Huaping Liu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16797 2025-02-25 cs.CL cs.AI 62%

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models

Alireza Amiri-Margavi, Iman Jebellat, Ehsan Jebellat, Seyed Pouyan Mousavi Davoudi

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 14 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16198 2025-02-25 cs.NI cs.AI cs.ET cs.LG 62%

An Autonomous Network Orchestration Framework Integrating Large Language Models with Continual Reinforcement Learning

Masoud Shokrnezhad, Tarik Taleb

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments IEEE Communications Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15700 2025-02-25 cs.IR cs.AI cs.CL 62%

Sustainable Digitalization of Business with Multi-Agent RAG and LLM

Muhammad Arslan, Saba Munawar, Christophe Cruz

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18328 2025-02-24 cs.CL cs.CY 62%

Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring

Xuansheng Wu, Padmaja Pravin Saraf, Gyeonggeon Lee, Ehsan Latif, Ninghao Liu, Xiaoming Zhai

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY

Comments Accepted by Technology, Knowledge, and Learning (TKNL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08877 2025-02-24 cs.SE cs.CL cs.LG 62%

Aligning the Objective of LLM-based Program Repair

Junjielong Xu, Ying Fu, Shin Hwei Tan, Pinjia He

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted by ICSE'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14893 2025-02-24 cs.CV cs.AI cs.LG cs.SD eess.AS 62%

NOTA: Multimodal Music Notation Understanding for Visual Large Language Model

Mingni Tang, Jiajia Li, Lu Yang, Zhiqiang Zhang, Jinghao Tian, Zuchao Li, Lefei Zhang, Ping Wang

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏