arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8044 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8044 篇

2506.23055 2025-07-01 cs.LG cs.AI 62%

Measuring How LLMs Internalize Human Psychological Concepts: A preliminary analysis

Hiro Taiyo Hamada, Ippei Fujisawa, Genji Kawakita, Yuki Yamada

机构 * R&D Department Araya inc.(阿莱亚公司研发部) Department of Computational Neuroscience(计算神经科学系) Kyusyu University(九州大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22520 2025-07-01 cs.HC cs.AI cs.CE cs.CY 62%

Exploring Artificial Intelligence Tutor Teammate Adaptability to Harness Discovery Curiosity and Promote Learning in the Context of Interactive Molecular Dynamics

Mustafa Demir, Jacob Miratsky, Jonathan Nguyen, Chun Kit Chan, Punya Mishra, Abhishek Singharoy

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21805 2025-06-30 cs.AI cs.CL 62%

CitySim: Modeling Urban Behaviors and City Dynamics with Large-Scale LLM-Driven Agent Simulation

Nicolas Bougie, Narimasa Watanabe

机构 * Woven by Toyota(丰田织造)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21803 2025-06-30 eess.SP cs.AI cs.LG 62%

From Token to Rhythm: A Multi-Scale Approach for ECG-Language Pretraining

Fuying Wang, Jiacheng Xu, Lequan Yu

机构 * School of Computing and Data Science, The University of Hong Kong, Hong Kong SAR, China(计算与数据科学学院,香港大学,香港特别行政区,中国)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21616 2025-06-30 cs.CL cs.CY 62%

TIM: A Large-Scale Dataset and large Timeline Intelligence Model for Open-domain Timeline Summarization

Chuanrui Hu, Wei Hu, Penghang Yu, Hua Zhang, Bing-Kun Bao

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20167 2025-06-26 cs.CL cs.AI 62%

SEED: A Structural Encoder for Embedding-Driven Decoding in Time Series Prediction with LLMs

Fengze Li, Yue Wang, Yangle Liu, Ming Huang, Dou Hong, Jieming Ma

机构 * Xi'an Jiaotong-Liverpool University(西安交通大学利物浦大学) University of Liverpool(利物浦大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19494 2025-06-26 cs.CL cs.LG 62%

Graph Linearization Methods for Reasoning on Graphs with Large Language Models

Christos Xypolopoulos, Guokan Shang, Xiao Fei, Giannis Nikolentzos, Hadi Abdine, Iakovos Evdaimon, Michail Chatzianastasis, Giorgos Stamou, Michalis Vazirgiannis

机构 * Ecole Polytechnique(巴黎高等理工学院) MBZUAI(马克斯·普朗克人工智能研究所) NTUA(希腊国家技术研究中心) University of Peloponnese(希腊皮洛斯大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19599 2025-06-25 cs.CL cs.AI 62%

ECCoT: A Framework for Enhancing Effective Cognition via Chain of Thought in Large Language Model

Zhenke Duan, Jiqun Pan, Jiani Tu, Xiaoyi Wang, Yanqing Wang

机构 * School of Statistics and Mathematics, Zhongnan University of Economics and Law(统计与数学学院,中国政法大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19525 2025-06-25 cs.CL cs.AI 62%

Automatic Posology Structuration : What role for LLMs?

Natalia Bobkova, Laura Zanella-Calzada, Anyes Tafoughalt, Raphaël Teboul, François Plesse, Félix Gaschi

机构 * SAS Posos Sorbonne Université(索邦大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18532 2025-06-25 cs.CL cs.LG cs.SD eess.AS 62%

End-to-End Spoken Grammatical Error Correction

Mengjie Qian, Rao Ma, Stefano Bannò, Mark J. F. Gales, Kate M. Knill

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16490 2025-06-25 cs.CR cs.AI cs.LG 62%

Towards Robust Stability Prediction in Smart Grids: GAN-based Approach under Data Constraints and Adversarial Challenges

Emad Efatinasab, Alessandro Brighente, Denis Donadel, Mauro Conti, Mirco Rampazzo

机构 * University of Padova, Department of Information Engineering(帕多瓦大学信息工程系) University of Padova, Department of Mathematics(帕多瓦大学数学系)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Journal ref Internet of Things, Volume 33, 2025, ISSN 2542-6605

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13642 2025-06-24 cs.AI cs.CL cs.CV cs.SD eess.AS 62%

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

Shaolei Zhang, Shoutao Guo, Qingkai Fang, Yan Zhou, Yang Feng

机构 * Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences (ICT/CAS)(中国科学院智能信息处理重点实验室) Key Laboratory of AI Safety, Chinese Academy of Sciences(中国科学院人工智能安全重点实验室) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Code: https://github.com/ictnlp/Stream-Omni , Model: https://huggingface.co/ICTNLP/stream-omni-8b

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01235 2025-06-24 stat.ML cs.AI cs.LG 62%

LoRA-One: One-Step Full Gradient Could Suffice for Fine-Tuning Large Language Models, Provably and Efficiently

Yuanhe Zhang, Fanghui Liu, Yudong Chen

机构 * Department of Statistics, University of Warwick, UK.(威斯敏斯特大学统计学系) Department of Computer Science, University of Warwick, UK.(威斯敏斯特大学计算机科学系) Centre for Discrete Mathematics and its Applications (DIMAP), University of Warwick, UK.(威斯敏斯特大学离散数学及其应用中心(DIMAP))

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted by ICML 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.04532 2025-06-24 cs.LG cs.AI cs.CV 62%

Recent Trends in Artificial Intelligence Technology: A Scoping Review

Teemu Niskanen, Tuomo Sipola, Olli Väänänen

机构 * Institute of Information Technology(信息技术学院)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17367 2025-06-24 cs.CL cs.AI cs.MA 62%

Cash or Comfort? How LLMs Value Your Inconvenience

Mateusz Cedro, Timour Ichmoukhamedov, Sofie Goethals, Yifan He, James Hinns, David Martens

机构 * University of Antwerp(安特卫普大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 12 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17267 2025-06-24 cs.LG cs.AI 62%

CF-VLM:CounterFactual Vision-Language Fine-tuning

Jusheng Zhang, Kaitong Cai, Yijia Fan, Jian Wang, Keze Wang

机构 * Sun Yat-sen University(中山大学) Snap Inc.(Snap公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17019 2025-06-23 cs.CL cs.AI 62%

Instituto de Telecomunicações at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning

Giuseppe Attanasio, Sonal Sannigrahi, Ben Peters, André F. T. Martins

机构 * Instituto de Telecomunicações(电信研究所) Instituto Superior Técnico(技术高等学院) Unbabel

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 7 pages, 1 figure, IWSLT 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16653 2025-06-23 cs.SE cs.AI cs.LG 62%

LLMs in Coding and their Impact on the Commercial Software Engineering Landscape

Vladislav Belozerov, Peter J Barclay, Askhan Sami

机构 * School of Engineering, Computing, \& the Built Environment Edinburgh Napier University, Scotland WWW home page: https://www.napier.ac.uk

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11618 2025-06-23 cs.LG cs.AI 62%

Convergent Linear Representations of Emergent Misalignment

Anna Soligo, Edward Turner, Senthooran Rajamanoharan, Neel Nanda

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15722 2025-06-23 cs.LG cs.AI 62%

UniMate: A Unified Model for Mechanical Metamaterial Generation, Property Prediction, and Condition Confirmation

Wangzhi Zhan, Jianpeng Chen, Dongqi Fu, Dawei Zhou

机构 * Department of Computer Science, Virginia Polytechnic Institute and State University, Virginia, US(计算机科学系,弗吉尼亚理工大学和州立大学,弗吉尼亚,美国) Meta AI

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15710 2025-06-23 cs.LG cs.AI 62%

RAST: Reasoning Activation in LLMs via Small-model Transfer

Siru Ouyang, Xinyu Zhu, Zilin Xiao, Minhao Jiang, Yu Meng, Jiawei Han

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Virginia(弗吉尼亚大学) Rice University(里德大学) GE HealthCare(通用电气医疗)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15468 2025-06-19 cs.HC cs.AI cs.LG 62%

Co-Creative Learning via Metropolis-Hastings Interaction between Humans and AI

Ryota Okumura, Tadahiro Taniguchi, Akira Taniguchi, Yoshinobu Hagiwara

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21794 2025-06-19 cs.CV cs.AI cs.LG 62%

Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey

Atsuyuki Miyai, Jingkang Yang, Jingyang Zhang, Yifei Ming, Yueqian Lin, Qing Yu, Go Irie, Shafiq Joty, Yixuan Li, Hai Li, Ziwei Liu, Toshihiko Yamasaki, Kiyoharu Aizawa

机构 * The University of Tokyo(东京大学) S-Lab, Nanyang Technological University(南洋理工大学S实验室) Duke University(杜克大学) Salesforce AI Research(Salesforce人工智能研究) Nanyang Technological University(南洋理工大学) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Tokyo University of Science(东京科学大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted at TMLR2025. Survey paper. We welcome questions, issues, and paper requests via https://github.com/AtsuMiyai/Awesome-OOD-VLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14577 2025-06-18 cs.LG cs.AI 62%

Object-Centric Neuro-Argumentative Learning

Abdul Rahman Jacob, Avinash Kori, Emanuele De Angelis, Ben Glocker, Maurizio Proietti, Francesca Toni

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Proceedings of Machine Learning Research, 2025 19th Conference on Neurosymbolic Learning and Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13811 2025-06-18 cs.MA cs.AI cs.CL 62%

Investigating the Potential of Large Language Model-Based Router Multi-Agent Architectures for Foundation Design Automation: A Task Classification and Expert Selection Study

Sompote Youwai, David Phim, Vianne Gayl Murcia, Rianne Clair Onas

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12093 2025-06-17 cs.CY cs.AI econ.GN q-fin.EC 62%

Intelligent Automation for FDI Facilitation: Optimizing Tariff Exemption Processes with OCR And Large Language Models

Muhammad Sukri Bin Ramli

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12042 2025-06-17 cs.LG cs.AI stat.ML 62%

CRITS: Convolutional Rectifier for Interpretable Time Series Classification

Alejandro Kuratomi, Zed Lee, Guilherme Dinis Chaliane Junior, Tony Lindgren, Diego García Pérez

机构 * Department of Computer and Systems Sciences, Stockholm University(斯德哥尔摩大学计算机与系统科学系) Department of Computer Science and Engineering, University of Oviedo(奥维多大学计算机科学与工程系)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments This paper was presented at the 2024 European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML-PKDD), as part of the XKDD workshop on interpretability. However it was not published in the LNCSI proceedings of the conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11104 2025-06-16 cs.CL cs.AI 62%

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration

Hanzhi Zhang, Heng Fan, Kewei Sha, Yan Huang, Yunhe Feng

机构 * LLaVi Lab, Department of Computer Science & Engineering(LLaVi实验室,计算机科学与工程系) Department of Data Science(数据科学系)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10953 2025-06-13 cs.LG cs.CL 62%

Build the web for agents, not agents for the web

Xing Han Lù, Gaurav Kamath, Marius Mosbach, Siva Reddy

机构 * McGill University(麦吉尔大学) Mila – Quebec AI Institute(蒙特利尔人工智能研究所)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10328 2025-06-13 cs.CV cs.AI cs.LG 62%

Towards Scalable SOAP Note Generation: A Weakly Supervised Multimodal Framework

Sadia Kamal, Tim Oates, Joy Wan

机构 * Department of Computer Science, University of Maryland, Baltimore County(计算机科学系,马里兰大学巴尔的摩县分校) Department of Dermatology, Johns Hopkins University School of Medicine(皮肤科系,约翰霍普金斯大学医学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted at IEEE/CVF Computer Society Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)

详情

展开后加载摘要…

URL PDF HTML 收藏