arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8057 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8057 篇

2504.11977 2025-04-17 cs.AI 57%

Leveraging Machine Learning Models to Predict the Outcome of Digital Medical Triage Interviews

Sofia Krylova, Fabian Schmidt, Vladimir Vlassov

机构 * Platform24 AB(普拉特福姆24公司) KTH Royal Institute of Technology(瑞典皇家理工学院)

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments 8 pages, 4 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02696 2025-04-17 cs.CL 57%

How Inclusively do LMs Perceive Social and Moral Norms?

Michael Galarnyk, Agam Shah, Dipanwita Guhathakurta, Poojitha Nandigam, Sudheer Chava

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted at NAACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11658 2025-04-17 cs.IR cs.AI 57%

Improving LLM Interpretability and Performance via Guided Embedding Refinement for Sequential Recommendation

Nanshan Jia, Chenfei Yuan, Yuhang Wu, Zeyu Zheng

机构 * University of California, Berkeley(加州大学伯克利分校) Berkeley AI Research Lab (BAIR)(伯克利人工智能研究实验室) Tsinghua University(清华大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11477 2025-04-17 cs.CV cs.AI 57%

SDIGLM: Leveraging Large Language Models and Multi-Modal Chain of Thought for Structural Damage Identification

Yunkai Zhang, Shiyin Wei, Yong Huang, Yawu Su, Shanshan Lu, Hui Li

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10647 2025-04-16 cs.CL 57%

Improving In-Context Learning with Reasoning Distillation

Nafis Sadeq, Xin Xu, Zhouhang Xie, Julian McAuley, Byungkyu Kang, Prarit Lamba, Xiang Gao

机构 * UC San Diego(加州大学圣迭戈分校) Intuit(英图伊特公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10538 2025-04-16 cs.IR cs.AI 57%

Distilling Transitional Pattern to Large Language Models for Multimodal Session-based Recommendation

Jiajie Su, Qiyong Zhong, Yunshan Ma, Weiming Liu, Chaochao Chen, Xiaolin Zheng, Jianwei Yin, Tat-Seng Chua

机构 * Zhejiang University(浙江大学) Singapore Management University(新加坡管理大学) National University of Singapore(新加坡国立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10169 2025-04-15 cs.LG stat.ML 57%

Challenges in interpretability of additive models

Xinyu Zhang, Julien Martinelli, ST John

机构 * Aalto University(阿尔托大学) Inserm(法国国家健康与医学研究院) Vaccine Research Institute(疫苗研究所) Université de Bordeaux(波尔多大学) Inria Bordeaux Sud-ouest(法国国家信息与自动化研究所波尔多西南分部)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Journal ref XAI-IJCAI24: Explainable AI workshop @ IJCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09488 2025-04-15 cs.CL 57%

Kongzi: A Historical Large Language Model with Fact Enhancement

Jiashu Yang, Ningning Wang, Yian Zhao, Chaoran Feng, Junjia Du, Hao Pang, Zhirui Fang, Xuxin Cheng

机构 * Dalian University of Technology(大连理工大学) Nanyang Technological University(南洋理工大学) Peking University(北京大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments 22 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09135 2025-04-15 cs.CL 57%

Efficient and Asymptotically Unbiased Constrained Decoding for Large Language Models

Haotian Ye, Himanshu Jain, Chong You, Ananda Theertha Suresh, Haowei Lin, James Zou, Felix Yu

机构 * Stanford University(斯坦福大学) Google(谷歌公司) Peking University(北京大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL

Journal ref AISTATS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05471 2025-04-09 cs.LG 57%

Graph Neural Networks for Enhancing Ensemble Forecasts of Extreme Rainfall

Christopher Bülte, Sohir Maskey, Philipp Scholl, Jonas von Berg, Gitta Kutyniok

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments Accepted paper at ICLR 2025 - Tackling Climate Change with Machine Learning Workshop (https://www.climatechange.ai/events/iclr2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04346 2025-04-09 cs.AI cs.SI 57%

Crowdsourcing-Based Knowledge Graph Construction for Drug Side Effects Using Large Language Models with an Application on Semaglutide

Zhijie Duan, Kai Wei, Zhaoqian Xue, Jiayan Zhou, Shu Yang, Siyuan Ma, Jin Jin, Lingyao li

机构 * University of Pennsylvania(宾夕法尼亚大学) University of Michigan(密歇根大学) Georgetown University(乔治城大学) Stanford University(斯坦福大学) Vanderbilt University(范德堡大学) University of South Florida(南佛罗里达大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02953 2025-04-07 cs.CL 57%

Cultural Learning-Based Culture Adaptation of Language Models

Chen Cecilia Liu, Anna Korhonen, Iryna Gurevych

机构 * Technical University of Darmstadt(达姆施塔特工业大学) Hessian Center for AI (hessian.AI)(黑森人工智能中心) University of Cambridge(剑桥大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12593 2025-04-07 cs.CV cs.AI 57%

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction

Yuanbin Man, Ying Huang, Chengming Zhang, Bingzhe Li, Wei Niu, Miao Yin

机构 * University of Texas at Arlington(德克萨斯大学阿灵顿分校) University of Houston(休斯顿大学) University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Georgia(佐治亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments CVPR 2025 Highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02688 2025-04-04 cs.NI cs.LG eess.SP 57%

Handover and SINR-Aware Path Optimization in 5G-UAV mmWave Communication using DRL

Achilles Kiwanuka Machumilane, Alberto Gotta, Pietro Cassarà

机构 * Institute of Information Science and Technologies (ISTI) CNR(国家研究委员会信息科学与技术研究所(ISTI))

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00223 2025-04-02 cs.LG cond-mat.mtrl-sci 57%

A machine learning platform for development of low flammability polymers

Duy Nhat Phan, Alexander B. Morgan, Lokendra Poudel, Rahul Bhowmik

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16872 2025-04-02 cs.LG cs.CV 57%

Lie Detector: Unified Backdoor Detection via Cross-Examination Framework

Xuan Wang, Siyuan Liang, Dongping Liao, Han Fang, Aishan Liu, Xiaochun Cao, Yu-liang Lu, Ee-Chien Chang, Xitong Gao

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15314 2025-04-02 cs.CL 57%

KTCR: Improving Implicit Hate Detection with Knowledge Transfer driven Concept Refinement

Samarth Garg, Vivek Hruday Kavuri, Gargi Shroff, Rahul Mishra

机构 * ABV-IIITM Gwalior(ABV-IIITM 格瓦利尔) IIIT Hyderabad(印度信息技术研究所海得拉巴分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments 9 pages, 4 figures, 2 algorithms, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07181 2025-04-02 cond-mat.mtrl-sci cs.LG physics.comp-ph 57%

A predictive machine learning force field framework for liquid electrolyte development

Sheng Gong, Yumin Zhang, Zhenliang Mu, Zhichen Pu, Hongyi Wang, Zhiao Yu, Mengyi Chen, Tianze Zheng, Zhi Wang, Lifei Chen, Zhenze Yang, Xiaojie Wu, Shaochen Shi, Weihao Gao, Wen Yan, Liang Xiang

机构 * ByteDance Research(字节跳动研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Figures provided as the tex source files

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24190 2025-04-01 cs.CL 57%

Implicit In-Context Learning: Evidence from Artificial Language Experiments

Xiaomeng Ma, Qihui Xu

机构 * Amazon(亚马逊) Ohio State University(俄亥俄州立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23601 2025-04-01 cs.RO cs.LG 57%

Exploring GPT-4 for Robotic Agent Strategy with Real-Time State Feedback and a Reactive Behaviour Framework

Thomas O'Brien, Ysobel Sims

机构 * University of Newcastle(纽卡斯尔大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Journal ref Australasian Conference on Robotics and Automation (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23407 2025-04-01 cs.CV cs.AI 57%

GMapLatent: Geometric Mapping in Latent Space

Wei Zeng, Xuebin Chang, Jianghao Su, Xiang Gu, Jian Sun, Zongben Xu

机构 * School of Mathematics and Statistics(数学与统计学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23306 2025-04-01 cs.CL 57%

Focus Directions Make Your Language Models Pay More Attention to Relevant Contexts

Youxiang Zhu, Ruochen Li, Danqing Wang, Daniel Haehn, Xiaohui Liang

机构 * University of Massachusetts Boston(马萨诸塞大学波士顿分校) Technische Universität München(慕尼黑工业大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21834 2025-03-31 cs.CV cs.AI 57%

A Multi-Modal Knowledge-Enhanced Framework for Vessel Trajectory Prediction

Haomin Yu, Tianyi Li, Kristian Torp, Christian S. Jensen

机构 * Aalborg University(奥尔堡大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21766 2025-03-28 cs.CV cs.AI 57%

Stable-SCore: A Stable Registration-based Framework for 3D Shape Correspondence

Haolin Liu, Xiaohang Zhan, Zizheng Yan, Zhongjin Luo, Yuxin Wen, Xiaoguang Han

机构 * FNii, CUHKSZ(香港中文大学(深圳)未来智人研究院) Tencent(腾讯) Tencent-Hunyuan3D(腾讯混元3D实验室) SSE, CUHKSZ(香港中文大学(深圳)理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accepted by CVPR 2025. Homepage: https://haolinliu97.github.io/Stable-Score/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03384 2025-03-28 cs.LG 57%

GNNMerge: Merging of GNN Models Without Accessing Training Data

Vipul Garg, Ishita Thakre, Sayan Ranu

机构 * Indian Institute of Technology Delhi(印度理工学院德里分校)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06608 2025-03-28 cs.CV cs.AI 57%

TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models

Yangguang Li, Zi-Xin Zou, Zexiang Liu, Dehu Wang, Yuan Liang, Zhipeng Yu, Xingchao Liu, Yuan-Chen Guo, Ding Liang, Wanli Ouyang, Yan-Pei Cao

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20662 2025-03-27 cs.CV cs.LG eess.IV 57%

AutoRad-Lung: A Radiomic-Guided Prompting Autoregressive Vision-Language Model for Lung Nodule Malignancy Prediction

Sadaf Khademi, Mehran Shabanpour, Reza Taleei, Anastasia Oikonomou, Arash Mohammadi

机构 * Concordia University(康考迪亚大学) Thomas Jefferson University Hospital(托马斯杰斐逊大学医院) Sunnybrook Health Sciences Centre(桑尼布鲁克健康科学中心)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20113 2025-03-27 cs.LG 57%

Domain Adaptation Framework for Turning Movement Count Estimation with Limited Data

Xiaobo Ma, Hyunsoo Noh, Ryan Hatch, James Tokishi, Zepu Wang

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments arXiv admin note: substantial text overlap with arXiv:2412.09861

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06302 2025-03-27 cs.CL 57%

Latent Convergence Modulation in Large Language Models: A Novel Approach to Iterative Contextual Realignment

Patricia Porretta, Sylvester Pakenham, Huxley Ainsworth, Gregory Chatten, Godfrey Allerton, Simon Hollingsworth, Vance Periwinkle

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18205 2025-03-27 cs.CL 57%

Contextually Structured Token Dependency Encoding for Large Language Models

James Blades, Frederick Somerfield, William Langley, Susan Everingham, Maurice Witherington

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏