arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8057 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8057 篇

2407.18328 2025-02-24 cs.CL cs.CY 62%

Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring

Xuansheng Wu, Padmaja Pravin Saraf, Gyeonggeon Lee, Ehsan Latif, Ninghao Liu, Xiaoming Zhai

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY

Comments Accepted by Technology, Knowledge, and Learning (TKNL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08877 2025-02-24 cs.SE cs.CL cs.LG 62%

Aligning the Objective of LLM-based Program Repair

Junjielong Xu, Ying Fu, Shin Hwei Tan, Pinjia He

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted by ICSE'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14893 2025-02-24 cs.CV cs.AI cs.LG cs.SD eess.AS 62%

NOTA: Multimodal Music Notation Understanding for Visual Large Language Model

Mingni Tang, Jiajia Li, Lu Yang, Zhiqiang Zhang, Jinghao Tian, Zuchao Li, Lefei Zhang, Ping Wang

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13619 2025-02-20 cs.CL cs.AI 62%

Complex Ontology Matching with Large Language Model Embeddings

Guilherme Sousa, Rinaldo Lima, Cassia Trojahn

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14745 2025-02-20 cs.CL cs.AI 62%

Semi-supervised Fine-tuning for Large Language Models

Junyu Luo, Xiao Luo, Xiusi Chen, Zhiping Xiao, Wei Ju, Ming Zhang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Github Repo: https://github.com/luo-junyu/SemiEvol

Journal ref NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11447 2025-02-20 cs.LG cs.AI 62%

Does Editing Provide Evidence for Localization?

Zihao Wang, Victor Veitch

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11969 2025-02-18 cs.AI cs.CV cs.LG 62%

Learning Generalizable Prompt for CLIP with Class Similarity Knowledge

Sehun Jung, Hyang-won Lee

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11843 2025-02-18 cs.CL cs.AI cs.SI 62%

Can LLM Agents Maintain a Persona in Discourse?

Pranav Bhandari, Nicolas Fay, Michael Wise, Amitava Datta, Stephanie Meek, Usman Naseem, Mehwish Nasim

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11560 2025-02-18 cs.AI cs.LG 62%

A Survey of Automatic Prompt Engineering: An Optimization Perspective

Wenwu Li, Xiangfeng Wang, Wenhao Li, Bo Jin

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 19 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11059 2025-02-18 cs.LG cs.AI 62%

ClimateLLM: Efficient Weather Forecasting via Frequency-Aware Large Language Models

Shixuan Li, Wei Yang, Peiyu Zhang, Xiongye Xiao, Defu Cao, Yuehan Qin, Xiaole Zhang, Yue Zhao, Paul Bogdan

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12228 2025-02-18 cs.IR cs.AI cs.CL 62%

Triple Modality Fusion: Aligning Visual, Textual, and Graph Data with Large Language Models for Multi-Behavior Recommendations

Luyi Ma, Xiaohan Li, Zezhong Fan, Kai Zhao, Jianpeng Xu, Jason Cho, Praveen Kanumala, Kaushiki Nag, Sushant Kumar, Kannan Achan

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10863 2025-02-18 cs.CL cs.AI 62%

Exploring the Personality Traits of LLMs through Latent Features Steering

Shu Yang, Shenzhe Zhu, Liang Liu, Lijie Hu, Mengdi Li, Di Wang

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05933 2025-02-18 cs.CL cs.AI 62%

Learning to Substitute Words with Model-based Score Ranking

Hongye Liu, Ricardo Henao

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at NAACL 2025 (main, long)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10266 2025-02-17 cs.CL cs.AI 62%

Are Large Language Models the future crowd workers of Linguistics?

Iris Ferrazzo

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16205 2025-02-13 cs.LG cs.AI cs.RO eess.SP 62%

Machine Learning-Based Estimation Of Wave Direction For Unmanned Surface Vehicles

Manele Ait Habouche, Mickaël Kerboeuf, Goulven Guillou, Jean-Philippe Babau

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12517 2025-02-11 cs.LG cs.AI 62%

Scaling FP8 training to trillion-token LLMs

Maxim Fishman, Brian Chmiel, Ron Banner, Daniel Soudry

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.07033 2025-02-11 cs.LG cs.AI 62%

Interpreting What Typical Fault Signals Look Like via Prototype-matching

Qian Chen, Xingjian Dong, Zhike Peng

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 17 pages, 12 figures, 6 tables

Journal ref Advanced Engineering Informatics, vol. 62, p. 102849, Oct. 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02945 2025-02-06 cs.CL cs.AI 62%

LLM-KT: Aligning Large Language Models with Knowledge Tracing using a Plug-and-Play Instruction

Ziwei Wang, Jie Zhou, Qin Chen, Min Zhang, Bo Jiang, Aimin Zhou, Qinchun Bai, Liang He

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13554 2025-02-06 cs.CV cs.AI cs.LG 62%

One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt

Tao Liu, Kai Wang, Senmao Li, Joost van de Weijer, Fahad Shahbaz Khan, Shiqi Yang, Yaxing Wang, Jian Yang, Ming-Ming Cheng

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 28 pages, 22 figures, ICLR2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08854 2025-02-06 cs.LG cs.AI cs.NI cs.SY eess.SY 62%

Hybrid LLM-DDQN based Joint Optimization of V2I Communication and Autonomous Driving

Zijiang Yan, Hao Zhou, Hina Tabassum, Xue Liu

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by IEEE Wireless Communications Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00989 2025-02-04 cs.CL cs.AI 62%

ChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution

Kanika Goswami, Puneet Mathur, Ryan Rossi, Franck Dernoncourt

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02626 2025-02-04 cs.CL cs.AI 62%

Time-Reversal Provides Unsupervised Feedback to LLMs

Yerram Varun, Rahul Madhavan, Sravanti Addepalli, Arun Suggala, Karthikeyan Shanmugam, Prateek Jain

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted as a spotlight in NeurIPS 2024

Journal ref The Thirty-Eighth Annual Conference on Neural Information Processing Systems (NeurIPS), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11276 2025-02-03 cs.LG cs.AI eess.SP 62%

Wearable Accelerometer Foundation Models for Health via Knowledge Distillation

Salar Abbaspourazad, Anshuman Mishra, Joseph Futoma, Andrew C. Miller, Ian Shapiro

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments updated format

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12473 2025-01-30 cs.CL cs.LG 62%

Large Language Models for Biomedical Knowledge Graph Construction: Information extraction from EMR notes

Vahan Arsenyan, Spartak Bughdaryan, Fadi Shaya, Kent Small, Davit Shahnazaryan

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16471 2025-01-29 cs.LG cs.AI eess.AS eess.IV q-bio.NC 62%

SIM: Surface-based fMRI Analysis for Inter-Subject Multimodal Decoding from Movie-Watching Experiments

Simon Dahan, Gabriel Bénédict, Logan Z. J. Williams, Yourong Guo, Daniel Rueckert, Robert Leech, Emma C. Robinson

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 27 pages, accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16377 2025-01-29 cs.LG cs.AI 62%

Optimal Signal Decomposition-based Multi-Stage Learning for Battery Health Estimation

Vijay Babu Pamshetti, Wei Zhang, King Jet Tseng, Bor Kiat Ng, Qingyu Yan

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11770 2025-01-29 cs.CL cs.AI 62%

CNMBERT: A Model for Converting Hanyu Pinyin Abbreviations to Chinese Characters

Zishuo Feng, Feng Cao

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 8 pages, 5 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00715 2025-01-28 cs.CV cs.AI cs.LG 62%

B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable

Shreyash Arya, Sukrut Rao, Moritz Böhle, Bernt Schiele

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 31 pages, 9 figures, 12 tables, Neural Information Processing Systems (NeurIPS) 2024; added references, corrected typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12651 2025-01-23 cs.CL cs.AI 62%

The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories

Raj Sanjay Shah, Sashank Varma

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12408 2025-01-23 cs.AI cs.LG cs.RO cs.SY eess.SY stat.ML 62%

Control-ITRA: Controlling the Behavior of a Driving Model

Vasileios Lioutas, Adam Scibior, Matthew Niedoba, Berend Zwartsenberg, Frank Wood

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 16 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏