arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-12 至 2026-01-12 共收录 9 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9 篇

2512.23243 2026-01-12 cs.CV 78%

Multimodal Interpretation of Remote Sensing Images: Dynamic Resolution Input Strategy and Multi-scale Vision-Language Alignment Mechanism

遥感图像的多模态解释:动态分辨率输入策略与多尺度视觉-语言对齐机制

Siyu Zhang, Lianlei Shan, Runhe Qiu

机构 * Tsinghua University(清华大学)

专题命中 安全评测 :alignment(title,abstract)

AI总结 本文提出动态分辨率输入策略与多尺度视觉-语言对齐机制,提升遥感图像多模态解释的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05570 2026-01-12 cs.AI cs.MA 70%

Crisis-Bench: Benchmarking Strategic Ambiguity and Reputation Management in Large Language Models

Crisis-Bench: 大型语言模型中战略模糊与声誉管理的基准测试

Cooper Lin, Maohao Ran, Yanting Zhang, Zhenglin Wan, Hongwei Fan, Yibo Xu, Yike Guo, Wei Xue, Jun Song

机构 * Hong Kong University of Science and Technology(香港科技大学) Hong Kong Baptist University(香港 Baptist 大学) National University of Singapore(新加坡国立大学) Imperial College London(伦敦帝国理工学院)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 Crisis-Bench通过多智能体POMDP评估LLM在高风险企业危机中的战略模糊与声誉管理能力,揭示模型在信息隐瞒与道德约束间的平衡问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05648 2026-01-12 q-bio.GN cs.AI cs.CL cs.LG 67%

Open World Knowledge Aided Single-Cell Foundation Model with Robust Cross-Modal Cell-Language Pre-training

开放世界知识辅助的单细胞基础模型与鲁棒跨模态细胞-语言预训练

Haoran Wang, Xuanyi Zhang, Shuangsang Fang, Longke Ran, Ziqing Deng, Yong Zhang, Yuxiang Li, Shaoshuai Li

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 OKR-CELL通过开放世界知识和鲁棒跨模态预训练,提升单细胞基础模型在细胞-语言跨模态任务中的表现。

Comments 41 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05904 2026-01-12 cs.CY cs.AI 62%

Can AI mediation improve democratic deliberation?

人工智能调解能否提升民主讨论?

Michael Henry Tessler, Georgina Evans, Michiel A. Bakker, Iason Gabriel, Sophie Bridgers, Rishub Jain, Raphael Koster, Verena Rieser, Anca Dragan, Matthew Botvinick, Christopher Summerfield

机构 * Google DeepMind(谷歌DeepMind) Massachusetts Institute of Technology(麻省理工学院) Yale Law School(耶鲁法学院) Department of Experimental Psychology, University of Oxford(牛津大学实验心理学系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

AI总结 本文探讨人工智能如何通过增强参与、公平调解和有意义讨论来提升民主讨论的质量。

Journal ref Knight Institute for the First Amendment at Columbia University Symposium on "AI and Democratic Freedoms", April 10-11, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05300 2026-01-12 cs.LG cs.CL 62%

TIME: Temporally Intelligent Meta-reasoning Engine for Context Triggered Explicit Reasoning

TIME:面向上下文触发的显式推理的时序智能元推理引擎

Susmit Das

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

AI总结 TIME通过引入时序智能元推理引擎,改进对话模型的显式推理能力,提升时间感知和推理效率。

Comments 14 pages, 3 figures with 27 page appendix. See https://github.com/The-Coherence-Initiative/TIME and https://github.com/The-Coherence-Initiative/TIMEBench for associated code

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05264 2026-01-12 cs.IR cs.AI 61%

Engineering the RAG Stack: A Comprehensive Review of the Architecture and Trust Frameworks for Retrieval-Augmented Generation Systems

构建RAG堆栈:对检索增强生成系统架构和信任框架的全面综述

Dean Wampler, Dave Nielson, Alireza Seddighi

机构 * The AI Alliance(AI联盟) IBM Research(IBM研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI;safety(comments)

AI总结 本文综述了RAG系统架构和信任框架,提出统一分类学和评估框架,为构建安全且领域适应性强的RAG系统提供指导。

Comments 86 pages, 2 figures, 37 tables. A comprehensive review of Retrieval-Augmented Generation (RAG) architectures and trust frameworks (2018-2025), encompassing a unified taxonomy, evaluation benchmarks, and trust-safety modeling

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19432 2026-01-12 cs.LG 57%

Advanced Long-term Earth System Forecasting

先进长期地球系统预测

Hao Wu, Yuan Gao, Ruijian Gou, Xian Wu, Chuhan Wu, Huahui Yi, Johannes Brandstetter, Fan Xu, Kun Wang, Penghao Zhao, Hao Jia, Qi Song, Xinliang Liu, Juncai He, Shuhao Cao, Huanshuo Dong, Yanfei Xiang, Fan Zhang, Haixin Wang, Xingjian Shi, Qiufeng Wang, Shuaipeng Li, Ruobing Xie, Feng Tao, Yuxu Lu, Yu Guo, Yuntian Chen, Yuxuan Liang, Qingsong Wen, Wanli Ouyang, Deliang Chen, Niklas Boers, Xiaomeng Huang

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

AI总结 TritonCast通过专用潜在动态核心和嵌套网格结构,实现长期稳定的地球系统预测,提升大气和海洋预测的准确性和稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07248 2026-01-12 cs.CL 57%

MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings

MedRiskEval: 医疗风险评估基准,语言模型在医疗领域中的重要性:用户视角的重要性

Jean-Philippe Corbeil, Minseon Kim, Maxime Griot, Sheela Agarwal, Alessandro Sordoni, Francois Beaulieu, Paul Vozila

机构 * Microsoft Healthcare & Life Sciences(微软医疗与生命科学) Microsoft Research Montréal, Canada(微软研究蒙特利尔分校) Université catholique de Louvain, Belgium(列日大学) Mila, Université de Montréal, Canada(蒙特利尔大学Mila)

专题命中 安全评测 :safety(abstract);分类 cs.CL

AI总结 MedRiskEval通过引入患者导向的数据集,评估了医疗领域中语言模型的风险,强调用户视角在医疗应用中的重要性。

Comments EACL2026 industry track

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05606 2026-01-12 cs.MA 50%

Conformity Dynamics in LLM Multi-Agent Systems: The Roles of Topology and Self-Social Weighting

LLM多智能体系统中的从众动力学:拓扑结构与自我-社会权重的作用

Chen Han, Jin Tan, Bohan Yu, Wenzhen Zheng, Xijin Tang

专题命中 安全评测 :alignment(abstract)

AI总结 本文研究了LLM多智能体系统中网络拓扑和自我-社会权重对集体决策效率、鲁棒性及失败模式的影响。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏