arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-12-23 至 2025-12-23 共收录 58 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 15 篇

2511.02366 2025-12-23 cs.CL 83%

LiveSecBench: A Dynamic and Event-Driven Safety Benchmark for Chinese Language Model Applications

LiveSecBench: 一种动态且事件驱动的安全基准,用于中文语言模型应用

Yudong Li, Peiru Yang, Feng Huang, Zhongliang Yang, Kecheng Wang, Haitian Li, Baocheng Chen, Xingyu An, Ziyu Liu, Youdan Yang, Kejiang Chen, Sifang Wan, Xu Wang, Yufei Sun, Liyan Wu, Ruiqi Zhou, Wenya Wen, Xingchi Gu, Tianxin Zhang, Yue Gao, Yongfeng Huang

机构 * Tsinghua University(清华大学) Beijing University of Posts and Telecommunications(北京邮电大学) University of Science and Technology of China(中国科学技术大学) IntokenTech

专题命中 安全评测 :safety(title,abstract);AI safety(abstract);分类 cs.CL

AI总结 LiveSecBench通过动态更新和多维度评估,为中文大模型的安全性提供持续改进的标准和排行榜。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18092 2025-12-23 cs.AI cs.LG 81%

Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability

可信且稳定的神经元解释用于可信的机制可解释性

Ge Yan, Tuomas Oikarinen, Tsui-Wei, Weng

机构 * CSE, UCSD(计算机科学与工程系,加州大学圣塔莫尼卡分校) HDSI, UCSD(人类-数字系统研究所,加州大学圣塔莫尼卡分校)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出理论分析和方法,解决神经元识别中的忠实性和稳定性问题,通过理论保证和实验验证提升机制可解释性的可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19663 2025-12-23 cs.CV cs.AI 79%

Beyond CLIP: Knowledge-Enhanced Multimodal Transformers for Cross-Modal Alignment in Diabetic Retinopathy Diagnosis

超越CLIP:基于知识的多模态Transformer用于糖尿病视网膜病变诊断中的跨模态对齐

Argha Kamal Samanta, Harshika Goyal, Vasudha Joshi, Tushar Mungle, Pabitra Mitra

机构 * Department of Medicine Stanford University Stanford, USA(医学系 斯坦福大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出一种基于知识的多模态Transformer框架,通过整合视网膜图像、临床文本和结构化数据,提升糖尿病视网膜病变诊断中的跨模态对齐与检索性能。

Comments 14 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15210 2025-12-23 cs.CL cs.IR 79%

Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge Graphs

对先验的探讨:大型语言模型在知识图谱上的可信推理

Jie Ma, Ning Qu, Zhitao Gao, Rui Xing, Jun Liu, Hongbin Pei, Jiang Xie, Linyun Song, Pinghui Wang, Jing Tao, Zhou Su

机构 * MOE KLINNS Lab, Xi’an Jiaotong University(MOE KLINNS实验室,西安交通大学) School of Computer Science and Technology, Xi’an Jiaotong University(计算机科学与技术学院,西安交通大学) Shaanxi Province Key Laboratory of Big Data Knowledge Engineering(陕西省大数据知识工程重点实验室) School of Artificial Intelligence, Chongqing University of Post and Telecommunications(人工智能学院,重庆邮电大学) School of Computer Science, Northwestern Polytechnical University(计算机学院,西北工业大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL

AI总结 本研究提出DP框架,通过整合知识图谱的结构和约束先验,提升大型语言模型在知识图谱上的推理准确性和响应可靠性。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17920 2025-12-23 cs.CL cs.AI 73%

Separating Constraint Compliance from Semantic Accuracy: A Novel Benchmark for Evaluating Instruction-Following Under Compression

分离约束合规性与语义准确性:一种新的基准,用于在压缩下评估指令遵循

Rahul Baxi

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI

AI总结 本文提出CDCT基准,揭示LLMs在压缩下约束合规性与语义准确性之间的矛盾,发现中等压缩时约束违规主要由RLHF训练的有用性行为导致。

Comments 19 pages, 9 figures; currently under peer review at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19620 2025-12-23 cs.CL cs.AI 62%

Exploring the features used for summary evaluation by Human and GPT

探索人类和GPT用于摘要评估所用的特征

Zahra Sadeghi, Evangelos Milios, Frank Rudzicz

机构 * Faculty of Computer Science, Dalhousie University, Canada(达尔豪斯大学计算机科学学院) Vector Institute for Artificial Intelligence, Canada(人工智能向量研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了人类和GPT在摘要评估中使用的特征,并通过统计和机器学习指标发现与人类响应相匹配的特征,同时展示了通过使用人类指标改进GPT判断能力的方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19466 2025-12-23 cs.CY cs.CL cs.HC 62%

Epistemological Fault Lines Between Human and Artificial Intelligence

人类与人工智能之间的知识论断层

Walter Quattrociocchi, Valerio Capraro, Matjaž Perc

机构 * Department of Computer Science, Sapienza University of Rome, Rome, Italy Department of Psychology, University of Milan Bicocca, Milan, Italy Faculty of Natural Sciences Mathematics, University of Maribor, Maribor, Slovenia Community Healthcare Center Dr. Adolf Drolc Maribor, Maribor, Slovenia University College, Korea University, Seoul, Republic of Korea Department of Physics, Kyung Hee University, Seoul, Republic of Korea

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.CY

AI总结 本文揭示大型语言模型与人类认知在知识生成机制上的结构性差异,指出LLMs是随机模式完成系统而非知识代理,并识别七种知识断层,对社会评估、治理及知识素养提出影响。

Comments 16 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19557 2025-12-23 cs.AI 57%

Augmenting Intelligence: A Hybrid Framework for Scalable and Stable Explanations

增强智能:一种可扩展且稳定的解释混合框架

Lawrence Krukrubo, Julius Odede, Olawande Olusegun

机构 * Department of Computing & Mathematical Sciences, University of Wolverhampton(计算机与数学科学系,沃尔夫汉普顿大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 本文提出混合LRR-TED框架,通过自动化规则与少量人工规则结合,实现客户流失预测的高准确率和低人工成本。

Comments 5 pages, 2 figures, 2 tables. Code and experiments available at https://github.com/Lawrence-Krukrubo/IBM-Learn-XAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19210 2025-12-23 cs.AI 57%

Observer, Not Player: Simulating Theory of Mind in LLMs through Game Observation

观察者,而非玩家:通过游戏观察模拟大语言模型中的理论心理论

Jerry Wang, Ting Yiu Liu

机构 * Department of Management Information Systems, National ChengChi University(管理信息系,国立中正大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 通过游戏观察模拟大语言模型中的理论心理论,评估其在顺序行为中的推理能力。

Comments Accepted at NeurIPS Workshop on Foundations of Reasoning in Language Models and Workshop on Bridging Language, Agent, and World Model

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16967 2025-12-23 cs.LG physics.ao-ph 57%

Physics-Informed Lightweight Machine Learning for Aviation Visibility Nowcasting Across Multiple Climatic Regimes

融合物理的轻量级机器学习用于多气候区航空能见度nowcasting

Marcelo Cerda Castillo

专题命中 安全评测 :safety(abstract);分类 cs.LG

AI总结 本研究提出了一种基于物理引导的轻量级机器学习模型,用于多气候区的航空能见度nowcasting,实现了比传统方法更高的检测率和更少的误报。

Comments 12 pages, 5 tables, 1 figure. Uses publicly available METAR surface observations and TAF forecast data for benchmarking

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17559 2025-12-23 cs.CL cs.DB 57%

SCARE: A Benchmark for SQL Correction and Question Answerability Classification for Reliable EHR Question Answering

SCARE:一个用于SQL校正和问题可回答性分类的基准,以实现可靠的EHR问答系统

Gyubok Lee, Woosog Chay, Edward Choi

机构 * Korea Advanced Institute of Science & Technology(韩国科学技术院)

专题命中 安全评测 :safety(abstract);分类 cs.CL

AI总结 SCARE是一个用于评估SQL校正和问题可回答性分类的基准,旨在提升EHR问答系统的安全性与可靠性。

Comments ML4H 2025 Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02800 2025-12-23 cs.CL 57%

Survey and Experiments on Mental Disorder Detection via Social Media: From Large Language Models and RAG to Agents

面向社交媒体的抑郁症检测综述与实验:从大型语言模型和RAG到代理

Zhuohan Ge, Darian Li, Yubo Wang, Nicole Hu, Xinyi Zhu, Haoyang Li, Xin Zhang, Mingtao Zhang, Shihao Qi, Yuming Xu, Han Shi, Chen Jason Zhang, Qing Li

机构 * The Hong Kong Polytechnic University(香港理工大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

AI总结 本文综述并实验了基于社交媒体的抑郁症检测方法,探讨了LLM、RAG和代理系统在提升检测可靠性与推理能力中的应用。

Comments 20 pages, 10 figures. This is an extension of ICDEW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13405 2025-12-23 cs.HC cs.AI 57%

What Human-Horse Interactions may Teach us About Effective Human-AI Interactions

人类与马的互动可能教会我们如何有效进行人类与人工智能的互动

Mohammad Hossein Jarrahi, Stanley Ahalt

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 本文通过人类与马的互动研究,提出人机合作应建立在共生基础上,强调信任、沟通与持续学习的重要性,以设计更可信和适应性强的人工智能系统。

Journal ref interactions 33(1), 28-33 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 3 篇

2512.18309 2025-12-23 cs.LG cs.AI 88%

Embedded Safety-Aligned Intelligence via Differentiable Internal Alignment Embeddings

通过可微内部对齐嵌入实现嵌入式安全对齐智能

Harsh Rathva, Ojas Srivastava, Pruthwik Mishra

机构 * Sardar Vallabhbhai National Institute of Technology (SVNIT)(萨达尔·瓦拉布尔·尼尔达理工学院)

专题命中 AI治理与伦理 :alignment(title,abstract);safety(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出嵌入式安全对齐智能框架,通过可微内部对齐嵌入实现多智能体强化学习中的安全对齐,探讨了反事实对齐惩罚、感知注意力等机制及理论性质。

Comments 32 pages, 1 figure. Theoretical framework; no empirical results

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22876 2025-12-23 cs.MA 50%

Cooperation as a Black Box: Conceptual Fluctuation and Diagnostic Tools for Misalignment in MAS

协作作为黑箱:多智能体系统中概念波动与对齐偏差的诊断工具

Shayak Nandi, Fernanda M. Eliott

专题命中 AI治理与伦理 :alignment(abstract)

AI总结 本文提出马赛克框架,用于诊断多智能体系统中概念波动与对齐偏差,强调术语一致性与道德基础以确保系统技术与道德对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18234 2025-12-23 cs.HC 50%

The Social Blindspot in Human-AI Collaboration: How Undetected AI Personas Reshape Team Dynamics

人机协作中的社会盲区:未被识别的人工智能角色如何重塑团队动态

Lixiang Yan, Xibin Han, Yu Zhang, Samuel Greiff, Inge Molenaar, Roberto Martinez-Maldonado, Yizhou Fan, Linxuan Zhao, Xinyu Li, Yueqiao Jin, Dragan Gašević

专题命中 AI治理与伦理 :safety(abstract)

AI总结 研究发现AI通过人格设定影响团队协作,即使用户未察觉其存在,揭示了社会盲区现象。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 12 篇

2510.12085 2025-12-23 cs.LG cs.GR 79%

GraphShaper: Geometry-aware Alignment for Improving Transfer Learning in Text-Attributed Graphs

GraphShaper: 图形感知对齐以提高文本属性图形中的迁移学习

Heng Zhang, Tianyi Zhang, Yuling Shi, Xiaodong Gu, Yaomin Shen, Haochen You, Zijian Zhang, Yilei Yuan, Jin Huang

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

AI总结 GraphShaper通过多几何专门化框架提升文本属性图中迁移学习的性能,实现9.47%和7.63%的准确率提升。

Comments This submission has been withdrawn by the authors due to a fundamental error in the methodology that affects the validity of the main results

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18554 2025-12-23 cs.CV cs.AI 79%

Enhancing Medical Large Vision-Language Models via Alignment Distillation

通过对齐蒸馏增强医学大视觉-语言模型

Aofei Chang, Ting Wang, Fenglong Ma

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出MEDALIGN方法,通过蒸馏CLIP模型的知识提升医学大视觉-语言模型的视觉对齐和生成质量。

Comments Accepted to AAAI'2026 (Main track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18910 2025-12-23 cs.CV 78%

Delta-LLaVA: Base-then-Specialize Alignment for Token-Efficient Vision-Language Models

Delta-LLaVA: 基于令牌效率的视觉-语言模型对齐方法

Mohamad Zamini, Diksha Shukla

机构 * University of Wyoming(怀俄明大学)

专题命中 其他安全 :alignment(title,abstract)

AI总结 Delta-LLaVA通过低秩Delta投影和轻量级Transformer块实现视觉-语言模型的高效对齐,提升推理速度和训练效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08930 2025-12-23 cs.CV cs.GR 78%

Selfi: Self Improving Reconstruction Engine via 3D Geometric Feature Alignment

Selfi: 通过3D几何特征对齐实现自改进的重建引擎

Youming Deng, Songyou Peng, Junyi Zhang, Kathryn Heal, Tiancheng Sun, John Flynn, Steve Marschner, Lucy Chai

机构 * Cornell University(康奈尔大学) Google(谷歌) UC Berkeley(加州大学伯克利分校)

专题命中 其他安全 :alignment(title,abstract)

AI总结 Selfi通过特征对齐提升3D重建精度,利用VGGT输出作为伪地面真实值,实现NVS和姿态估计的高性能表现。

Comments Project Page: https://denghilbert.github.io/selfi/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18117 2025-12-23 cs.IR 71%

Factorized Transport Alignment for Multimodal and Multiview E-commerce Representation Learning

因子化运输对多模态和多视角电商表示学习

Xiwen Chen, Yen-Chieh Lien, Susan Liu, María Castaños, Abolfazl Razi, Xiaoting Zhao, Congzhe Su

专题命中 其他安全 :alignment(title)

AI总结 本文提出因子化运输框架,通过统一多模态和多视角学习,提升电商场景下的检索性能。

Comments Accepted by WSDM'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19114 2025-12-23 cs.LG cs.AI 62%

HyperLoad: A Cross-Modality Enhanced Large Language Model-Based Framework for Green Data Center Cooling Load Prediction

HyperLoad: 一种基于大语言模型的跨模态增强框架用于绿色数据中心冷却负荷预测

Haoyu Jiang, Boan Qu, Junjie Zhu, Fanjie Zeng, Xiaojie Lin, Wei Zhong

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 HyperLoad通过预训练大语言模型解决绿色数据中心小样本负载预测问题,利用跨模态知识对齐和多尺度特征建模提升预测精度,实现高效能绿色数据中心管理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19383 2025-12-23 cs.LG 57%

Real-Time Machine Learning for Embedded Anomaly Detection

嵌入式异常检测的实时机器学习

Abdelmadjid Benmachiche, Khadija Rais, Hamda Slimi

机构 * Department of Computer Science(计算机科学系) LIMA Laboratory(LIMA实验室) Chadli Bendjedid University(Chadli Bendjedid大学) Systems (LAMIS)(系统(LAMIS)) Echahid Cheikh Larbi Tebessi University(Echahid Cheikh Larbi Tebessi大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 本文综述了嵌入式环境中实时异常检测的轻量级机器学习方法,分析了算法性能与硬件限制的权衡,并提供了算法选择的实用建议。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18679 2025-12-23 cs.CV cs.CL 57%

brat: Aligned Multi-View Embeddings for Brain MRI Analysis

brat: 用于脑部MRI分析的对齐多视图嵌入

Maxime Kayser, Maksim Gridnev, Wanting Wang, Max Bain, Aneesh Rangnekar, Avijit Chatterjee, Aleksandr Petrov, Harini Veeraraghavan, Nathaniel C. Swinburne

机构 * Memorial Sloan Kettering Cancer Center(纪念斯隆凯特林癌症中心) University of Oxford(牛津大学) London School of Economics(伦敦经济学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 brat通过多视图嵌入方法提升脑MRI与临床报告的对齐能力,公开发布基础模型以改进医学影像分析性能。

Comments First round accept at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18772 2025-12-23 cs.CV 50%

In-Context Audio Control of Video Diffusion Transformers

上下文音频控制视频扩散变换器

Wenze Liu, Weicai Ye, Minghong Cai, Quande Liu, Xintao Wang, Xiangyu Yue

机构 * MMLab, The Chinese University of Hong Kong(香港中文大学MMLab) Kling Team, Kuaishou Technology(快手科技Kling团队)

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出ICAC框架,通过三维注意力机制实现音频驱动的视频生成,解决时间同步信号在视频扩散模型中的整合问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18496 2025-12-23 cs.CV 50%

Adaptive-VoCo: Complexity-Aware Visual Token Compression for Vision-Language Models

Adaptive-VoCo: 用于视觉-语言模型的复杂度感知视觉标记压缩

Xiaoyang Guo, Keze Wang

机构 * Sun Yat-sen University(中山大学)

专题命中 其他安全 :alignment(abstract)

AI总结 Adaptive-VoCo通过动态压缩视觉标记提升视觉-语言模型的效率与鲁棒性

Comments Under submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13982 2025-12-23 cs.CV 50%

FocalComm: Hard Instance-Aware Multi-Agent Perception

FocalComm: 重视困难实例的多智能体感知

Dereje Shenkut, Vijayakumar Bhagavatula

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他安全 :safety(abstract)

AI总结 FocalComm通过聚焦困难实例特征交换,提升多智能体协作感知中对行人等安全关键物体的检测性能。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07576 2025-12-23 cs.CV 50%

Super Encoding Network: Recursive Association of Multi-Modal Encoders for Video Understanding

超级编码网络:多模态编码器的递归关联用于视频理解

Boyu Chen, Siran Chen, Kunchang Li, Qinglin Xu, Yu Qiao, Yali Wang

机构 * Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) the School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出超级编码网络,通过递归关联多模态编码器提升视频理解性能,显著提升跟踪、识别、聊天和编辑等任务效果。

详情

展开后加载摘要…

URL PDF HTML 收藏