arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-02-27 至 2026-02-27 共收录 53 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 17 篇

2602.23363 2026-02-27 cs.CV 50%

MediX-R1: Open Ended Medical Reinforcement Learning

MediX-R1:开放端医疗强化学习

Sahal Shaji Mullappilly, Mohammed Irfan Kurpath, Omair Mohamed, Mohamed Zidan, Fahad Khan, Salman Khan, Rao Anwer, Hisham Cholakkal

机构 * Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI)(迈赫迈德·本·扎耶德人工智能大学) Jubilee Mission Medical College(jubilee mission 医学院) Research Institute(研究院) JJM Medical College(JJM 医学院)

专题命中 安全评测 :alignment(abstract)

AI总结 MediX-R1通过综合奖励信号和LLM评估,实现医疗多模态模型的开放端强化学习,提升临床推理可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22923 2026-02-27 cs.CV cs.RO 50%

WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents

WaterVideoQA: 以ASV为中心的感知与符合规则的推理 via 多模态智能体

Runwei Guan, Shaofeng Liang, Ningwei Ouyang, Weichen Fei, Shanliang Yao, Wei Dai, Chenhao Ge, Penglei Sun, Xiaohui Zhu, Tao Huang, Ryan Wen Liu, Hui Xiong

机构 * Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州)人工智能研究所) Hubei Key Laboratory of Inland Shipping Technology (Wuhan University of Technology)(湖北内河航运技术重点实验室(武汉理工大学)) School of Navigation, Wuhan University of Technology(武汉理工大学航海学院) School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学先进科技学院) School of Artificial Intelligence, Nanjing University(南京大学人工智能学院) School of Information Engineering, Yancheng Institute of Technology(盐城职业技术学院信息工程学院) School of Engineering, Stanford University(斯坦福大学工程学院) Centre for AI and Data Science Innovation and the School of Science and Engineering, James Cook University(詹姆斯库克大学人工智能与数据科学创新中心及科学与工程学院)

专题命中 安全评测 :trustworthy(abstract)

AI总结 WaterVideoQA通过多模态智能体系统,实现ASV在复杂水域环境中的感知与规则合规推理,提升自主航行的安全性和精确性。

Comments 11 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20903 2026-02-27 cs.CV 50%

TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering

TextPecker: 通过奖励结构异常量化提升视觉文本渲染

Hanshen Zhu, Yuliang Liu, Xuecheng Wu, An-Lan Wang, Hao Feng, Dingkang Yang, Chao Feng, Can Huang, Jingqun Tang, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) ByteDance(字节跳动)

专题命中 安全评测 :alignment(abstract)

AI总结 TextPecker通过结构异常感知强化学习策略提升视觉文本渲染的结构忠实度和语义对齐度。

Comments Accepted by CVPR 2026; Code: https://github.com/CIawevy/TextPecker

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22253 2026-02-27 cs.SD 50%

AR&D: A Framework for Retrieving and Describing Concepts for Interpreting AudioLLMs

AR&D: 一种用于音频大语言模型解释的检索与描述框架

Townim Faisal Chowdhury, Ta Duc Huy, Siqi Pan, Jeremy Stoddard, Zhibin Liao

机构 * Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) Dolby Laboratories(杜比实验室) School of Computer and Mathematical Sciences, University of Adelaide, Australia(计算机与数学科学学院,阿德莱德大学,澳大利亚)

专题命中 安全评测 :trustworthy(abstract)

AI总结 AR&D框架通过稀疏自编码器解构音频大语言模型的多义激活,实现对模型内部特征的可解释性增强,为高风险领域应用提供可靠部署基础。

Comments Accepted at International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 3 篇

2602.22557 2026-02-27 cs.AI cs.LG 81%

CourtGuard: A Model-Agnostic Framework for Zero-Shot Policy Adaptation in LLM Safety

CourtGuard:一种用于大语言模型安全的模型无关框架,用于零样本策略适应

Umid Suleymanov, Rufiz Bayramov, Suad Gafarli, Seljan Musayeva, Taghi Mammadov, Aynur Akhundlu, Murat Kantarcioglu

机构 * Department of Computer Science, Virginia Tech(弗吉尼亚理工大学计算机科学系) School of IT(信息技术学院) Engineering, ADA University(工程学院,ADA大学) School of Law, ADA University(法学院,ADA大学)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.LG

AI总结 CourtGuard通过证据辩论框架实现大语言模型的安全零样本适应,超越传统方法在多个安全基准测试中的表现,并具备跨领域泛化和自动数据审计能力。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22790 2026-02-27 cs.CL cs.AI 62%

Natural Language Declarative Prompting (NLD-P): A Modular Governance Method for Prompt Design Under Model Drift

自然语言声明式提示(NLD-P):一种模块化治理方法,用于在模型漂移下的提示设计

Hyunwoo Kim, Hanau Yi, Jaehee Bae, Yumin Kim

机构 * ddai Inc.(ddai公司)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出NLD-P作为模块化治理方法,通过自然语言声明式提示解决模型漂移下的提示设计问题,提供可解释的控制框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05435 2026-02-27 eess.AS cs.AI cs.LG 62%

Unbiased Sliced Wasserstein Kernels for High-Quality Audio Captioning

无偏切片Wasserstein核用于高质量音频描述生成

Manh Luong, Khai Nguyen, Dinh Phung, Gholamreza Haffari, Lizhen Qu

机构 * Monash University(墨尔本大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出无偏切片Wasserstein核,通过保留模态间时间信息,提升音频描述质量及推理能力。

Journal ref Manh Luong. (2025). Unbiased Sliced Wasserstein Kernels for High-Quality Audio Captioning. In Advances in Neural Information Processing Systems 38 (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 16 篇

2602.22775 2026-02-27 cs.HC cs.AI cs.CL 81%

TherapyProbe: Generating Design Knowledge for Relational Safety in Mental Health Chatbots Through Adversarial Simulation

TherapyProbe: 通过对抗性模拟生成关系安全设计知识以改进心理健康聊天机器人

Joydeep Chandra, Satyam Kumar Navneet, Yong Zhang

机构 * BNRIST, Dept. of CST, Tsinghua University(北京理工大学(Tsinghua大学)) Independent Researcher(独立研究者)

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI

AI总结 TherapyProbe通过对抗性模拟生成关系安全设计知识,帮助改进心理健康聊天机器人的交互模式质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22660 2026-02-27 cs.LG 79%

LEDA: Latent Semantic Distribution Alignment for Multi-domain Graph Pre-training

LEDA:用于多领域图预训练的潜在语义分布对齐

Lianze Shan, Jitao Zhao, Dongxiao He, Siqi Liu, Jiaxu Cui, Weixiong Zhang

机构 * Tianjin University(天津大学) Jilin University(吉林大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

AI总结 LEDA通过潜在语义分布对齐方法,提升多领域图预训练的通用性和跨领域性能。

Comments Accepted by WWW-26, 12 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22475 2026-02-27 cs.CL 79%

Mind the Gap in Cultural Alignment: Task-Aware Culture Management for Large Language Models

注意文化契合差距:面向任务的文化管理用于大语言模型

Binchi Zhang, Xujiang Zhao, Jundong Li, Haifeng Chen, Zhengzhang Chen

机构 * University of Virginia(弗吉尼亚大学) NEC Laboratories America(NEC美洲实验室)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

AI总结 本文提出CultureManager,一种面向任务的文化管理方法,通过任务感知的文化数据对齐和文化路由器管理,提升大语言模型在文化敏感任务中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22879 2026-02-27 cs.AI 74%

Towards LLM-Empowered Knowledge Tracing via LLM-Student Hierarchical Behavior Alignment in Hyperbolic Space

基于大语言模型的知识追踪:通过超几何空间中的LLM-学生层次行为对齐

Xingcheng Fu, Shengpeng Wang, Yisen Gao, Xianxian Li, Chunpei Li, Qingyun Sun, Dongran Yu

专题命中 其他安全 :alignment(title);分类 cs.AI

AI总结 本文提出L-HAKT框架,通过超几何空间中的层次行为对齐,提升知识追踪对认知状态层次演化和个性化问题难度感知的建模能力。

Comments 9 pages, 6 figures, Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23068 2026-02-27 cs.SD 71%

TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual Alignment

TADA: 一种通过文本-语音双模对齐生成语音模型的框架

Trung Dang, Sharath Rao, Ananya Gupta, Christopher Gagne, Panagiotis Tzirakis, Alice Baird, Jakub Piotr Cłapa, Peter Chin, Alan Cowen

机构 * Hume AI Dartmouth College(达特茅斯学院)

专题命中 其他安全 :alignment(title)

AI总结 TADA通过文本-语音双模对齐生成框架,实现语音与文本的一对一同步建模,提升语音生成的保真度和效率,减少幻觉并降低推理成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20031 2026-02-27 cs.AI cs.LG 62%

Latent Introspection: Models Can Detect Prior Concept Injections

潜在自我反思:模型可以检测先前概念注入

Theia Pearson-Vogel, Martin Vanek, Raymond Douglas, Jan Kulveit

机构 * ACS Research, CTS, Charles University(ACS研究机构、CTs、查尔斯大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

AI总结 Qwen 32B模型能检测并识别先前注入的概念,通过提供准确的AI自我反思机制信息可显著提升检测效果,同时提高注入概念间的互信息。

Comments 28 pages, 17 figures. Submitted to ICML 2026. Workshop version submitted to ICLR 2026 Workshop on Latent and Implicit Thinking

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22576 2026-02-27 cs.CL cs.IR cs.LG 62%

Search-P1: Path-Centric Reward Shaping for Stable and Efficient Agentic RAG Training

Search-P1: 基于路径的奖励塑造用于稳定且高效的代理RAG训练

Tianle Xia, Ming Xu, Lingxiang Hu, Yiding Sun, Wenwei Li, Linfang Shang, Liqun Liu, Peng Shu, Huan Yu, Jie Jiang

机构 * Tencent(腾讯)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

AI总结 Search-P1通过路径中心奖励塑造提升代理RAG训练的稳定性和效率,实现7.7个百分点的准确率提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23336 2026-02-27 cs.LG stat.ML 57%

Differentiable Zero-One Loss via Hypersimplex Projections

通过超简单面投影实现可微零一损失

Camilo Gomez, Pengyang Wang, Liansheng Tang

机构 * School of Data, Mathematical, and Statistical Sciences, University of Central Florida, Orlando, USA(数据、数学与统计科学学院,中央佛罗里达大学,奥兰多,美国) Department of CIS, University of Macau, Macao, China(信息与系统系,澳门大学,澳门,中国)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本文提出了一种可微的零一损失近似方法,通过超简单面投影和Soft-Binary-Argmax操作符,提升大批次训练下的模型泛化能力。

Comments To appear in PAKDD 2026 (Pacific-Asia Conference on Knowledge Discovery and Data Mining), 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20570 2026-02-27 cs.CV cs.AI 57%

Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP

Dyslexify: 一种针对CLIP中印刷攻击的机制性防御

Lorenz Hufe, Constantin Venhoff, Erblina Purelku, Maximilian Dreyer, Sebastian Lapuschkin, Wojciech Samek

机构 * Fraunhofer Heinrich Hertz Institute(弗劳恩霍夫海因里希·赫兹研究所) University of Oxford(牛津大学) Technological University Dublin(都柏林技术大学) Technische Universität Berlin(柏林技术大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 Dyslexify通过消融CLIP中的印刷电路,有效防御印刷攻击,提升性能并保持应用安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22404 2026-02-27 cs.CL 57%

SAFARI: A Community-Engaged Approach and Dataset of Stereotype Resources in the Sub-Saharan African Context

SAFARI:一种社区参与的方法和撒哈拉以南非洲语境下的刻板印象资源数据集

Aishwarya Verma, Laud Ammah, Olivia Nercy Ndlovu Lucas, Andrew Zaldivar, Vinodkumar Prabhakaran, Sunipa Dev

机构 * Google Research(谷歌研究) RAIN Africa(RAIN非洲) Mantaray Africa(Mantaray非洲)

专题命中 其他安全 :safety(abstract);分类 cs.CL

AI总结 SAFARI通过社区参与方法构建了覆盖四个撒哈拉以南非洲国家的多语言刻板印象资源数据集,旨在填补NLP资源的空白。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22267 2026-02-27 cs.LG 57%

Data-Driven Supervision of a Thermal-Hydraulic Process Towards a Physics-Based Digital Twin

数据驱动的热力过程监督以实现基于物理的数字孪生

Osimone Imhogiemhe, Yoann Jus, Hubert Lejeune, Saïd Moussaoui

机构 * Nantes Univ., Centrale Nantes LS2N, CNRS UMR 6004 F-44000 Nantes, France Fluid \& Sealing Technologies CETIM Nantes, France

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 本文提出了一种基于物理的数字孪生方法,通过数值模拟和机器学习实现热力过程的故障检测与诊断,验证了其在参数变化检测中的有效性。

Journal ref International Conference on Control, Automation and Diagnosis, Jul 2025, Barcelona (ES), Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22613 2026-02-27 cs.CV 50%

Spectrally Distilled Representations Aligned with Instruction-Augmented LLMs for Satellite Imagery

基于指令增强大语言模型的光谱蒸馏表示用于卫星图像

Minh Kha Do, Wei Xiang, Kang Han, Di Wu, Khoa Phan, Yi-Ping Phoebe Chen, Gaowen Liu, Ramana Rao Kompella

机构 * La Trobe University(拉特罗布大学) Cisco Research(思科研究)

专题命中 其他安全 :alignment(abstract)

AI总结 SATtxt通过光谱蒸馏和指令增强大语言模型实现卫星图像的光谱感知视觉-语言学习,提升了零样本分类、检索和线性探测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22562 2026-02-27 cs.CR 50%

Layer-Targeted Multilingual Knowledge Erasure in Large Language Models

多语言知识擦除中的层目标知识擦除

Taoran Li, Varun Chandrasekaran, Zhiyuan Yu

专题命中 其他安全 :alignment(abstract)

AI总结 MUTE通过识别语言无关的中间层实现多语言知识擦除,解决了LLMs中跨语言反向学习的泛化问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22416 2026-02-27 cs.HC 50%

Seeing Graphs Like Humans: Benchmarking Computational Measures and MLLMs for Similarity Assessment

像人类一样看图:通过计算度量和MLLMs进行相似性评估的基准测试

Seokweon Jung, Jeongmin Rhee, Seoyoung Doh, Hyeon Jeon, Ghulam Jilani Quadri, Jinwook Seo

专题命中 其他安全 :alignment(abstract)

AI总结 本文通过实验比较了计算度量和MLLMs在图相似性评估中的表现,发现MLLMs,特别是GPT-5,在匹配人类感知方面表现优异,能提供可解释的推理理由。

Comments 21 pages including 1 page of appendix, 9 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19929 2026-02-27 cs.NI cs.IT math.IT 50%

BeamVLM for Low-altitude Economy: Generative Beam Prediction via Vision-language Models

BeamVLM 用于低空经济:通过视觉-语言模型进行生成式波束预测

Chenran Kou, Changsheng You, Mingjiang Wu, Dingzhu Wen, Zezhong Zhang, Chengwen Xing

专题命中 其他安全 :alignment(abstract)

AI总结 BeamVLM 通过视觉-语言模型实现生成式波束预测,提升无人机与地面基站间通信的准确性和泛化能力。

Comments We propose a novel end-to-end generative framework for beam prediction by using vision-language models

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10585 2026-02-27 cs.GR 50%

D3MAS: Decompose, Deduce, and Distribute for Enhanced Knowledge Sharing in Multi-Agent Systems

D3MAS:分解、推断与分配以增强多智能体系统中的知识共享

Heng Zhang, Yuling Shi, Xiaodong Gu, Haochen You, Zijian Zhang, Lubin Gan, Yilei Yuan, Jin Huang

专题命中 其他安全 :alignment(abstract)

AI总结 D3MAS通过分层协调框架减少多智能体系统中的知识冗余,提升推理准确性。

Comments This submission has been withdrawn by the authors due to a fundamental error in the methodology that affects the validity of the main results

详情

展开后加载摘要…

URL PDF HTML 收藏