arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-12-23 至 2025-12-23 共收录 15 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 15 篇

2505.14972 2025-12-23 cs.CL 88%

Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies

多模态文化安全:评估框架与对齐策略

Haoyi Qiu, Kung-Hsiang Huang, Ruichen Zheng, Jiao Sun, Nanyun Peng

机构 * University of California, Los Angeles(加州大学洛杉矶分校) Salesforce AI Research(Salesforce人工智能研究) Google DeepMind(谷歌DeepMind)

专题命中 安全评测 :alignment(title,abstract);safety(title,abstract);分类 cs.CL

AI总结 本文提出CROSS基准和CROSS-Eval框架,评估多模态模型的文化安全能力,发现提升推理能力可改善文化对齐,但需结合监督微调和偏好微调策略以增强文化合规性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11361 2025-12-23 cs.CL 87%

VLDBench Evaluating Multimodal Disinformation with Regulatory Alignment

VLDBench:评估具有监管对齐的多模态虚假信息

Shaina Raza, Ashmal Vayani, Aditya Jain, Aravind Narayanan, Vahid Reza Khazaie, Syed Raza Bashir, Elham Dolatabadi, Gias Uddin, Christos Emmanouilidis, Rizwan Qureshi, Mubarak Shah

专题命中 安全评测 :alignment(title,abstract);safety(abstract);trustworthy(abstract);AI safety(abstract)

AI总结 VLDBench 是首个多模态虚假信息检测基准,通过大规模标注数据提升检测准确率,支持 AI 管治框架下的可信虚假信息分析。

Comments Accepted in Information Fusion Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02366 2025-12-23 cs.CL 83%

LiveSecBench: A Dynamic and Event-Driven Safety Benchmark for Chinese Language Model Applications

LiveSecBench: 一种动态且事件驱动的安全基准,用于中文语言模型应用

Yudong Li, Peiru Yang, Feng Huang, Zhongliang Yang, Kecheng Wang, Haitian Li, Baocheng Chen, Xingyu An, Ziyu Liu, Youdan Yang, Kejiang Chen, Sifang Wan, Xu Wang, Yufei Sun, Liyan Wu, Ruiqi Zhou, Wenya Wen, Xingchi Gu, Tianxin Zhang, Yue Gao, Yongfeng Huang

机构 * Tsinghua University(清华大学) Beijing University of Posts and Telecommunications(北京邮电大学) University of Science and Technology of China(中国科学技术大学) IntokenTech

专题命中 安全评测 :safety(title,abstract);AI safety(abstract);分类 cs.CL

AI总结 LiveSecBench通过动态更新和多维度评估,为中文大模型的安全性提供持续改进的标准和排行榜。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18092 2025-12-23 cs.AI cs.LG 81%

Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability

可信且稳定的神经元解释用于可信的机制可解释性

Ge Yan, Tuomas Oikarinen, Tsui-Wei, Weng

机构 * CSE, UCSD(计算机科学与工程系,加州大学圣塔莫尼卡分校) HDSI, UCSD(人类-数字系统研究所,加州大学圣塔莫尼卡分校)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出理论分析和方法,解决神经元识别中的忠实性和稳定性问题,通过理论保证和实验验证提升机制可解释性的可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19663 2025-12-23 cs.CV cs.AI 79%

Beyond CLIP: Knowledge-Enhanced Multimodal Transformers for Cross-Modal Alignment in Diabetic Retinopathy Diagnosis

超越CLIP:基于知识的多模态Transformer用于糖尿病视网膜病变诊断中的跨模态对齐

Argha Kamal Samanta, Harshika Goyal, Vasudha Joshi, Tushar Mungle, Pabitra Mitra

机构 * Department of Medicine Stanford University Stanford, USA(医学系 斯坦福大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出一种基于知识的多模态Transformer框架,通过整合视网膜图像、临床文本和结构化数据,提升糖尿病视网膜病变诊断中的跨模态对齐与检索性能。

Comments 14 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15210 2025-12-23 cs.CL cs.IR 79%

Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge Graphs

对先验的探讨:大型语言模型在知识图谱上的可信推理

Jie Ma, Ning Qu, Zhitao Gao, Rui Xing, Jun Liu, Hongbin Pei, Jiang Xie, Linyun Song, Pinghui Wang, Jing Tao, Zhou Su

机构 * MOE KLINNS Lab, Xi’an Jiaotong University(MOE KLINNS实验室,西安交通大学) School of Computer Science and Technology, Xi’an Jiaotong University(计算机科学与技术学院,西安交通大学) Shaanxi Province Key Laboratory of Big Data Knowledge Engineering(陕西省大数据知识工程重点实验室) School of Artificial Intelligence, Chongqing University of Post and Telecommunications(人工智能学院,重庆邮电大学) School of Computer Science, Northwestern Polytechnical University(计算机学院,西北工业大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL

AI总结 本研究提出DP框架,通过整合知识图谱的结构和约束先验,提升大型语言模型在知识图谱上的推理准确性和响应可靠性。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17920 2025-12-23 cs.CL cs.AI 73%

Separating Constraint Compliance from Semantic Accuracy: A Novel Benchmark for Evaluating Instruction-Following Under Compression

分离约束合规性与语义准确性:一种新的基准,用于在压缩下评估指令遵循

Rahul Baxi

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI

AI总结 本文提出CDCT基准,揭示LLMs在压缩下约束合规性与语义准确性之间的矛盾,发现中等压缩时约束违规主要由RLHF训练的有用性行为导致。

Comments 19 pages, 9 figures; currently under peer review at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19620 2025-12-23 cs.CL cs.AI 62%

Exploring the features used for summary evaluation by Human and GPT

探索人类和GPT用于摘要评估所用的特征

Zahra Sadeghi, Evangelos Milios, Frank Rudzicz

机构 * Faculty of Computer Science, Dalhousie University, Canada(达尔豪斯大学计算机科学学院) Vector Institute for Artificial Intelligence, Canada(人工智能向量研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了人类和GPT在摘要评估中使用的特征,并通过统计和机器学习指标发现与人类响应相匹配的特征,同时展示了通过使用人类指标改进GPT判断能力的方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19466 2025-12-23 cs.CY cs.CL cs.HC 62%

Epistemological Fault Lines Between Human and Artificial Intelligence

人类与人工智能之间的知识论断层

Walter Quattrociocchi, Valerio Capraro, Matjaž Perc

机构 * Department of Computer Science, Sapienza University of Rome, Rome, Italy Department of Psychology, University of Milan Bicocca, Milan, Italy Faculty of Natural Sciences Mathematics, University of Maribor, Maribor, Slovenia Community Healthcare Center Dr. Adolf Drolc Maribor, Maribor, Slovenia University College, Korea University, Seoul, Republic of Korea Department of Physics, Kyung Hee University, Seoul, Republic of Korea

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.CY

AI总结 本文揭示大型语言模型与人类认知在知识生成机制上的结构性差异,指出LLMs是随机模式完成系统而非知识代理,并识别七种知识断层,对社会评估、治理及知识素养提出影响。

Comments 16 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19557 2025-12-23 cs.AI 57%

Augmenting Intelligence: A Hybrid Framework for Scalable and Stable Explanations

增强智能:一种可扩展且稳定的解释混合框架

Lawrence Krukrubo, Julius Odede, Olawande Olusegun

机构 * Department of Computing & Mathematical Sciences, University of Wolverhampton(计算机与数学科学系,沃尔夫汉普顿大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 本文提出混合LRR-TED框架,通过自动化规则与少量人工规则结合,实现客户流失预测的高准确率和低人工成本。

Comments 5 pages, 2 figures, 2 tables. Code and experiments available at https://github.com/Lawrence-Krukrubo/IBM-Learn-XAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19210 2025-12-23 cs.AI 57%

Observer, Not Player: Simulating Theory of Mind in LLMs through Game Observation

观察者,而非玩家:通过游戏观察模拟大语言模型中的理论心理论

Jerry Wang, Ting Yiu Liu

机构 * Department of Management Information Systems, National ChengChi University(管理信息系,国立中正大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 通过游戏观察模拟大语言模型中的理论心理论,评估其在顺序行为中的推理能力。

Comments Accepted at NeurIPS Workshop on Foundations of Reasoning in Language Models and Workshop on Bridging Language, Agent, and World Model

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16967 2025-12-23 cs.LG physics.ao-ph 57%

Physics-Informed Lightweight Machine Learning for Aviation Visibility Nowcasting Across Multiple Climatic Regimes

融合物理的轻量级机器学习用于多气候区航空能见度nowcasting

Marcelo Cerda Castillo

专题命中 安全评测 :safety(abstract);分类 cs.LG

AI总结 本研究提出了一种基于物理引导的轻量级机器学习模型,用于多气候区的航空能见度nowcasting,实现了比传统方法更高的检测率和更少的误报。

Comments 12 pages, 5 tables, 1 figure. Uses publicly available METAR surface observations and TAF forecast data for benchmarking

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17559 2025-12-23 cs.CL cs.DB 57%

SCARE: A Benchmark for SQL Correction and Question Answerability Classification for Reliable EHR Question Answering

SCARE:一个用于SQL校正和问题可回答性分类的基准,以实现可靠的EHR问答系统

Gyubok Lee, Woosog Chay, Edward Choi

机构 * Korea Advanced Institute of Science & Technology(韩国科学技术院)

专题命中 安全评测 :safety(abstract);分类 cs.CL

AI总结 SCARE是一个用于评估SQL校正和问题可回答性分类的基准,旨在提升EHR问答系统的安全性与可靠性。

Comments ML4H 2025 Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02800 2025-12-23 cs.CL 57%

Survey and Experiments on Mental Disorder Detection via Social Media: From Large Language Models and RAG to Agents

面向社交媒体的抑郁症检测综述与实验:从大型语言模型和RAG到代理

Zhuohan Ge, Darian Li, Yubo Wang, Nicole Hu, Xinyi Zhu, Haoyang Li, Xin Zhang, Mingtao Zhang, Shihao Qi, Yuming Xu, Han Shi, Chen Jason Zhang, Qing Li

机构 * The Hong Kong Polytechnic University(香港理工大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

AI总结 本文综述并实验了基于社交媒体的抑郁症检测方法,探讨了LLM、RAG和代理系统在提升检测可靠性与推理能力中的应用。

Comments 20 pages, 10 figures. This is an extension of ICDEW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13405 2025-12-23 cs.HC cs.AI 57%

What Human-Horse Interactions may Teach us About Effective Human-AI Interactions

人类与马的互动可能教会我们如何有效进行人类与人工智能的互动

Mohammad Hossein Jarrahi, Stanley Ahalt

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 本文通过人类与马的互动研究,提出人机合作应建立在共生基础上,强调信任、沟通与持续学习的重要性,以设计更可信和适应性强的人工智能系统。

Journal ref interactions 33(1), 28-33 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏