arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-03-03 至 2026-03-03 共收录 40 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 40 篇

2603.00088 2026-03-03 cs.CY 88%

AI Safety Evaluations Need To Consider Cascading Effects

AI安全评估需要考虑连锁效应

Anna Neumann, Jatinder Singh

专题命中 其他安全 :safety(title,abstract);AI safety(title,abstract);分类 cs.CY

AI总结 本文提出通过分析AI供应链中各组件的连锁效应,改进AI安全评估方法,以提升系统透明度和安全性。

Comments As accepted in the IASEAI 2026 Proceedings, Track 11: Epistemological challenges of AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01297 2026-03-03 cs.LG cs.CL 84%

I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift

我难以相信它不稳健:在嵌入漂移下安全分类器的灾难性崩溃

Subramanyam Sahoo, Vinija Jain, Divya Chaudhary, Aman Chadha

机构 * Independent(独立研究者) Meta AI AWS Generative AI Innovation Center, Amazon Web Services(AWS生成式AI创新中心,亚马逊网络服务) Northeastern University, Seattle, WA, USA(东北大学,西雅图,华盛顿州,美国) Stanford University(斯坦福大学)

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.CL、cs.LG

AI总结 研究发现嵌入漂移导致安全分类器性能大幅下降,揭示了生产AI安全架构的脆弱性并挑战了安全机制的转移假设。

Comments Accepted at the ICBINB: Where LLMs Need to Improve workshop at ICLR 2026. 12 pages and 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01096 2026-03-03 cs.CV cs.AI cs.CL cs.LG 82%

Unified Vision-Language Modeling via Concept Space Alignment

通过概念空间对齐实现统一的视觉-语言建模

Yifu Qiu, Paul-Ambroise Duquenne, Holger Schwenk

机构 * University of Edinburgh(爱丁堡大学) FAIR at Meta(Meta公司FAIR团队)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 通过概念空间对齐实现统一的视觉-语言建模,V-SONAR在多语言和多模态任务中超越现有模型。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02742 2026-03-03 cs.LG cs.AI 81%

Entropy-Guided Dynamic Tokens for Graph-LLM Alignment in Molecular Understanding

熵引导的动态令牌用于分子理解中的图-语言模型对齐

Zihao Jing, Qiuhao Zeng, Ruiyi Fang, Yan Sun, Boyu Wang, Pingzhao Hu

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 EDT-Former通过熵引导的动态令牌生成,在无需微调LLM主干的情况下实现图编码器与LLM的对齐,提升分子图理解的效率和泛化能力。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00097 2026-03-03 q-bio.BM cs.AI cs.CE cs.IR cs.LG 81%

Exploring Drug Safety Through Knowledge Graphs: Protein Kinase Inhibitors as a Case Study

通过知识图谱探索药物安全性:蛋白激酶抑制剂作为案例研究

David Jackson, Michael Gertz, Jürgen Hesser

机构 * David Jackson Institute of Informatics University of Amsterdam(大卫·杰克逊信息学研究所 阿姆斯特丹大学) Michael Gertz Institute of Computer Science Heidelberg University(迈克尔·格茨计算机科学研究所海德堡大学) Jürgen Hesser Data Analysis and Modeling in Medicine Mannheim Institute for Intelligent Systems in Medicine (MIISM) Medical Faculty Mannheim, Heidelberg University(朱尔根·赫瑟医学数据分析与建模 马尔姆研究所(MIISM)海德堡大学医学系)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于知识图谱的框架,整合多种数据源以预测蛋白激酶抑制剂的不良药物反应,通过分析疗效、靶点相似性和不良事件相关性,提升药物安全性研究。

Comments 14 pages, 5 figures. Code and data available at https://github.com/davidjackson99/PKI_KG

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00483 2026-03-03 cs.CV cs.AI 79%

RAISE: Requirement-Adaptive Evolutionary Refinement for Training-Free Text-to-Image Alignment

RAISE:基于需求的进化精炼用于无训练文本到图像对齐

Liyao Jiang, Ruichen Chen, Chao Gao, Di Niu

机构 * Department of ECE, University of Alberta, Canada(阿尔伯塔大学电子工程系) Huawei Technologies, Canada(华为技术有限公司)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

AI总结 RAISE是一种无训练、需求驱动的进化框架,通过动态调整生成过程实现高效且通用的文本到图像对齐。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00443 2026-03-03 cs.CV 78%

SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural Alignment

SesaHand: 通过语义和结构对齐可控生成增强3D手重建

Zhuoran Zhao, Xianghao Kong, Linlin Yang, Zheng Wei, Pan Hui, Anyi Rao

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学) Communication University of China(中国通信大学)

专题命中 其他安全 :alignment(title,abstract)

AI总结 SesaHand通过语义和结构对齐提升可控手部生成,优化3D手重建性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00412 2026-03-03 cs.CV 78%

PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language Models

PointAlign:用于3D视觉-语言模型的特征级对齐正则化

Yuanhao Su, Shaofeng Zhang, Xiaosong Jia, Qi Fan

机构 * University of Science and Technology of China(中国科学技术大学) Fuzhou University(福州大学) Fudan University(复旦大学) Nanjing University(南京大学)

专题命中 其他安全 :alignment(title,abstract)

AI总结 PointAlign通过特征级对齐正则化提升3D视觉-语言模型的几何信息保留与任务性能

Comments CVPR 2026 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00193 2026-03-03 q-bio.QM 78%

Multimodal Alignment Improves Generalizability of Genomic Biomarker Prediction in Computational Pathology

多模态对齐提升了计算病理学中基因组生物标志物预测的泛化能力

Ekaterina Redekop, Eric Zimmermann, Ava P Amini, Alex X Lu, Neil Tenenholtz, James Brian Hall, Lorin Crawford, Kristen A Severson

专题命中 其他安全 :alignment(title,abstract)

AI总结 MARBLE通过多模态对比预训练策略,将组织病理学图像与基因组生物标志物的表示对齐,提升计算病理学中基因组生物标志物预测的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01477 2026-03-03 cs.RO 71%

SFCo-Nav: Efficient Zero-Shot Visual Language Navigation via Collaboration of Slow LLM and Fast Attributed Graph Alignment

SFCo-Nav: 通过慢LLM与快速属性图对齐的协作实现高效的零样本视觉语言导航

Chaoran Xiong, Litao Wei, Xinhao Hu, Kehui Ma, Ziyi Xia, Zixin Jiang, Zhen Sun, Ling Pei

机构 * Shanghai Key Laboratory of Navigation and Location Based Services, Shanghai Jiao Tong University(上海导航与位置基于服务重点实验室,上海交通大学)

专题命中 其他安全 :alignment(title)

AI总结 SFCo-Nav通过慢LLM与快速属性图对齐的协作,实现了高效的零样本视觉语言导航,显著提升效率并降低计算成本。

Comments Accepted by 2026 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00498 2026-03-03 cs.LG 70%

Antibody: Strengthening Defense Against Harmful Fine-Tuning for Large Language Models via Attenuating Harmful Gradient Influence

Antibody: 通过削弱有害梯度影响来加强对抗有害微调的防御

Quoc Minh Nguyen, Trung Le, Jing Wu, Anh Tuan Bui, Mehrtash Harandi

机构 * Department of Electrical and Computer Systems Engineering, Monash University, Australia(莫纳什大学电气与计算机系统工程系) Department of Data Science and AI, Monash University, Australia(莫纳什大学数据科学与人工智能系)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.LG

AI总结 Antibody通过削弱有害梯度影响,提升大型语言模型对有害微调攻击的防御能力,并增强微调性能。

Comments Published at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00062 2026-03-03 cs.CY 70%

How much technical talent is there? A systematic estimate of the ML research pool among 3 million consultants

有多少技术人才?对300万顾问中机器学习研究池的系统估计

Maximilian Schons, Red Bermejo, Florian Aldehoff-Zeidler, Niccolò Zanichelli, Oliver Evans, Gavin Leech, Samuel Härgestam

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CY

AI总结 研究通过系统方法估计300万顾问中机器学习研究人才数量,发现技术人才数量超过MATS培训计划校友,且通过工作测试的AI模型尚未出现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21722 2026-03-03 cs.CL cs.AI cs.CY 67%

German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies

德国通用社会调查人设:一种基于德国通用社会调查的人员提示集合,用于与人口一致的LLM研究

Jens Rupprecht, Leon Fröhling, Claudia Wagner, Markus Strohmaier

机构 * University of Mannheim(曼海姆大学) GESIS -- Leibniz Institute for the Social Sciences(莱布尼茨社会科学研究所) RWTH Aachen University(亚琛工业大学) Complexity Science Hub(复杂性科学中心)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出德国通用社会调查人设集合,用于构建与人口一致的LLM研究,通过实验证明其在模拟调查响应方面的有效性。

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04067 2026-03-03 cs.LG cs.AI cs.CL 67%

What Scales in Cross-Entropy Scaling Law?

交叉熵缩放定律中什么因素具有可扩展性?

Junxi Yan, Zixi Wei, Qingyao Ai, Yiqun Liu, Jingtao Zhan

机构 * Tsinghua University(清华大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文通过分解交叉熵为误差熵、自我对齐和置信度三个部分,揭示了误差熵是唯一遵循幂律缩放的因素,从而解释了交叉熵缩放定律在小规模有效但在大尺度失效的原因。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01438 2026-03-03 cs.CL cs.AI 62%

Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agents

通过动态重要性估计增强解码时的个性跟随以用于角色扮演代理

Yuxin Liu, Mingye Zhu, Siyuan Liu, Bo Hu, Lei Zhang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种基于理论的动态重要性估计方法,通过加权奖励引导解码提升角色扮演代理在动态场景中的个性跟随能力。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01274 2026-03-03 cs.LG cs.AI 62%

GlassMol: Interpretable Molecular Property Prediction with Concept Bottleneck Models

GlassMol:通过概念瓶颈模型实现可解释的分子属性预测

Oscar Rivera, Ziqing Wang, Matthieu Dagommer, Abhishek Pandey, Kaize Ding

机构 * Northwestern University(西北大学) AbbVie(阿比维)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

AI总结 GlassMol通过自动化概念整理和LLM引导的概念选择,解决分子属性预测中的可解释性与性能权衡问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14760 2026-03-03 cs.CL cs.AI 62%

Residual Connections and the Causal Shift: Uncovering a Structural Misalignment in Transformers

残差连接与因果偏移:揭示变换器中的结构性不一致

Jonathan Lys, Vincent Gripon, Bastien Pasdeloup, Axel Marmoret, Lukas Mauch, Fabien Cardinaux, Ghouthi Boukli Hacene

机构 * IMT Atlantique, Lab-STICC, UMR CNRS 6285, Brest, France(IMT Atlantique,Lab-STICC,CNRS 6285研究所,布列塔尼,法国) Sony Europe Ltd. Stuttgart Technology Center, EUREC, Germany(索尼欧洲有限公司,斯图加特技术中心,EUREC,德国)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本研究揭示了自回归变换器中残差连接与因果偏移的结构性不一致,并提出基于残差衰减的轻量级缓解方法,通过实验验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05296 2026-03-03 cs.AI cs.LG 62%

Control Tax: The Price of Keeping AI in Check

控制税:保持AI受控的成本

Mikhail Terekhov, Zhen Ning David Liu, Caglar Gulcehre, Samuel Albanie

机构 * EPFL(瑞士联邦理工学院) MATS

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出控制税概念,通过理论框架量化控制成本,评估语言模型在对抗性环境下的安全性,并开发优化的监控策略以平衡安全与成本效益。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17692 2026-03-03 cs.LG cs.AI 62%

Agentic Unlearning: When LLM Agent Meets Machine Unlearning

代理反学习:当大语言模型代理遇见机器反学习

Bin Wang, Fan Wang, Pingping Wang, Jinyu Cong, Yang Yu, Yilong Yin, Zhongyi Han, Benzheng Wei

机构 * Center for Medical Artificial Intelligence, Shandong University of Traditional Chinese Medicine(山东中医药大学中医人工智能中心) School of Software, Shandong University(山东大学软件学院) School of Medical Information Engineering, Shandong University of Traditional Chinese Medicine(山东中医药大学医学信息工程学院) Shandong Huazhi Talent Technology Co., Ltd.(山东华智人才科技有限公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出SBU框架,通过同步双更新协议实现参数与记忆路径的联合反学习,有效减少敏感信息痕迹并保持数据保留质量。

Comments 9 pages, 6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22611 2026-03-03 cs.LG cs.AI 62%

Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning

分位数优势估计:稳定LLM推理的RLVR

Junkang Wu, Kexin Huang, Jiancan Wu, An Zhang, Xiang Wang, Xiangnan He

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

AI总结 通过分位数优势估计方法,稳定RLVR训练过程,提升LLM推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19791 2026-03-03 cs.CL cs.IR 57%

ToolDreamer: Instilling LLM Reasoning Into Tool Retrievers

ToolDreamer: 将大语言模型推理能力注入工具检索器

Saptarshi Sengupta, Zhengyu Zhou, Jun Araki, Xingbo Wang, Bingqing Wang, Suhang Wang, Zhe Feng

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Bosch Research North America(博世北美研究)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 ToolDreamer通过合成工具描述来改进工具检索,提升稀疏和密集检索器性能,减少LLM上下文窗口压力。

Comments Accepted to EACL 2026 (main/oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01822 2026-03-03 cs.AI 57%

Emerging Human-like Strategies for Semantic Memory Foraging in Large Language Models

新兴的人类样策略在大语言模型中的语义记忆采集

Eric Lacosse, Mariana Duarte, Peter M. Todd, Daniel C. McNamee

机构 * Champalimaud Research(恰帕拉德研究) Centre for Restorative Neurotechnology(修复性神经技术中心) Indiana University(印第安纳大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本研究通过分析大语言模型中的语义流畅任务,探索其语义记忆采集策略,揭示人类与AI在认知机制上的异同。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20089 2026-03-03 cs.CV cs.AI 57%

StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues

StructXLIP: 通过多模态结构线索增强视觉-语言模型

Zanxi Ruan, Songqun Gao, Qiuyu Kong, Yiming Wang, Marco Cristani

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 StructXLIP通过引入多模态结构线索提升视觉-语言模型的跨模态检索性能。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09463 2026-03-03 cs.AI 57%

SpotAgent: Grounding Visual Geo-localization in Large Vision-Language Models through Agentic Reasoning

SpotAgent:通过代理推理在大视觉语言模型中实现视觉地理定位

Furong Jia, Ling Dai, Wenjin Deng, Fan Zhang, Chen Hu, Daxin Jiang, Yu Liu

机构 * Peking University(北京大学) StepFun

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 SpotAgent通过代理推理提升大视觉语言模型在地理定位中的表现,结合外部工具验证和强化学习优化,实现更精确和可靠的定位结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04369 2026-03-03 cs.LG 57%

Multi-scale hypergraph meets LLMs: Aligning large language models for time series analysis

多尺度超图与大语言模型:面向时间序列分析的对齐方法

Zongjiang Shang, Dongliang Cui, Binqing Wu, Ling Chen

机构 * State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学) College of Computer Science and Technology, Zhejiang University(计算机科学与技术学院,浙江大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本文提出MSH-LLM方法,通过多尺度超图机制和跨模态对齐模块,提升大语言模型在时间序列分析中的表现。

Comments Accepted by ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00350 2026-03-03 cs.AI 57%

Monotropic Artificial Intelligence: Toward a Cognitive Taxonomy of Domain-Specialized Language Models

单调人工智能:面向领域专用语言模型的认知分类

Antonio de Sousa Leitão Filho, Allan Kardec Duailibe Barros Filho, Fabrício Saul Lima, Selby Mykael Lima dos Santos, Rejani Bandeira Vieira Sousa

机构 * Aia Context Federal University of Maranhão(马那瓜联邦大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 单调人工智能通过牺牲通用性实现领域内高精度,挑战通用智能主导的AI研究范式,提出专门化与通用系统互补共存的认知生态。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05064 2026-03-03 cs.LG 57%

Boomerang Distillation Enables Zero-Shot Model Size Interpolation

回环蒸馏使零样本模型大小插值

Sara Kangaslahti, Nihal V. Nayak, Jonathan Geuter, Marco Fumero, Francesco Locatello, David Alvarez-Melis

机构 * Harvard University(哈佛大学) Kempner Institute(凯普纳研究所) IST Austria(IST奥地利研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 回环蒸馏通过蒸馏和重构实现零样本模型大小插值,生成细粒度模型家族,降低训练成本并提升适应性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12490 2026-03-03 physics.ao-ph cs.LG 57%

SamudrACE: Fast and Accurate Coupled Climate Modeling with 3D Ocean and Atmosphere Emulators

SamudrACE:基于3D海洋和大气模拟器的快速准确耦合气候建模

James P. C. Duncan, Elynn Wu, Surya Dheeshjith, Adam Subel, Troy Arcomano, Spencer K. Clark, Brian Henn, Anna Kwa, Jeremy McGibbon, W. Andre Perkins, William Gregory, Carlos Fernandez-Granda, Julius Busecke, Oliver Watt-Meyer, William J. Hurlin, Alistair Adcroft, Laure Zanna, Christopher Bretherton

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 SamudrACE通过3D海洋和大气模拟器实现快速准确的耦合气候建模,能够模拟数百年长的高分辨率气候现象。

Comments 29 pages, 26 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18564 2026-03-03 cs.LG cs.CE 57%

Efficient Aircraft Design Optimization Using Multi-Fidelity Models and Multi-fidelity Physics Informed Neural Networks

利用多保真模型和多保真物理指导神经网络实现高效飞机设计优化

Apurba Sarker

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本研究利用多保真物理指导神经网络和生成对抗网络,实现高效飞机设计优化,提升设计迭代速度与经济性。

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00144 2026-03-03 cs.CV cs.AI 57%

Disentangled Hierarchical VAE for 3D Human-Human Interaction Generation

解耦层次变分自编码器用于3D人-人交互生成

Zichen Geng, Zeeshan Hayder, Bo Miao, Jian Liu, Wei Liu, Ajmal Mian

机构 * Department of CSSE, The University of Western Australia(西澳大学计算机科学与工程系) Data61, CSIRO(澳大利亚联邦科学工业研究组织Data61部门) Australian Institute for Machine Learning, The University of Adelaide(澳大利亚阿德莱德大学人工智能研究所) NERC-RVC, Hunan University(湖南大学NERC-RVC部门)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出DHVAE,通过解耦层次变分自编码器生成结构化且可控的3D人-人交互,提升运动保真度和物理合理性。

详情

展开后加载摘要…

URL PDF HTML 收藏