arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-14 至 2026-01-14 共收录 19 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 19 篇

2601.06663 2026-01-14 cs.AI 83%

SafePro: Evaluating the Safety of Professional-Level AI Agents

SafePro:评估专业级AI代理的安全性

Kaiwen Zhou, Shreedhar Jangam, Ashwin Nagarajan, Tejas Polu, Suhas Oruganti, Chengzhi Liu, Ching-Chen Kuo, Yuting Zheng, Sravana Narayanaraju, Xin Eric Wang

机构 * UCSC(加州大学圣塔 Cruz 分校) UCSB(加州大学圣塔 Barbara 分校) eBay(eBay 公司)

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.AI

AI总结 SafePro通过高复杂度专业任务数据集评估专业级AI代理的安全性,揭示了现有模型在安全判断和对齐方面的不足,并提出了改进安全性的策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18882 2026-01-14 cs.CY 83%

Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach

大语言模型中的个性化安全:一个基准和基于规划的代理方法

Yuchen Wu, Edward Sun, Kaijie Zhu, Jianxun Lian, Jose Hernandez-Orallo, Aylin Caliskan, Jindong Wang

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CY

AI总结 本文提出个性化安全框架PENGUIN和RAISE,通过获取用户背景信息提升LLM的安全性,实验显示安全评分提升达31.6%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08673 2026-01-14 cs.AI cs.CY 81%

Why AI Alignment Failure Is Structural: Learned Human Interaction Structures and AGI as an Endogenous Evolutionary Shock

为何AI对齐失败是结构性的:学习的人类互动结构与AGI作为内生性演化冲击

Didier Sornette, Sandro Claudio Lera, Ke Wu

机构 * Institute of Risk Analysis, Prediction and Management (Risks-X)(风险分析、预测与管理研究所) Southern University of Science and Technology (SUSTech)(南方科技大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.CY

AI总结 本文探讨AI对齐失败的结构性问题,指出AGI作为内生性演化冲击,其风险源于放大人类智能、权力与矛盾,而非对抗意图。

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08331 2026-01-14 cs.CL 74%

CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark

CLaS-Bench: 一种跨语言对齐与引导基准

Daniil Gurgurov, Yusser Al Ghussin, Tanja Baeumel, Cheng-Ting Chou, Patrick Schramowski, Marius Mosbach, Josef van Genabith, Simon Ostermann

机构 * Saarland University(萨尔兰大学) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI)) Centre for European Research in Trusted AI (CERTAIN)(可信AI欧洲研究中心(CERTAIN)) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) TU Darmstadt(德累斯顿技术大学) Mila - Quebec Artificial Intelligence Institute(魁北克人工智能研究所(Mila)) McGill University(麦吉尔大学)

专题命中 安全评测 :alignment(title);分类 cs.CL

AI总结 CLaS-Bench是首个多语言引导基准,通过评估32种语言中的语言强制行为,验证了残差基于DiffMean方法在跨语言引导中的有效性。

Comments pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08420 2026-01-14 cs.CV 71%

MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP

MMLGNet: 利用CLIP实现遥感数据的跨模态对齐

Aditya Chaudhary, Sneha Barman, Mainak Singha, Ankit Jha, Girish Mishra, Biplab Banerjee

专题命中 安全评测 :alignment(title)

AI总结 MMLGNet通过CLIP实现遥感数据的跨模态对齐,利用多模态语言引导网络有效融合光谱、空间和几何信息,提升语义理解能力。

Comments Accepted at InGARSS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08342 2026-01-14 cs.CL 70%

Detecting Mental Manipulation in Speech via Synthetic Multi-Speaker Dialogue

通过合成多说话对话检测心理操控

Run Chen, Wen Liang, Ziwei Gong, Lin Ai, Julia Hirschberg

机构 * Columbia University(哥伦比亚大学) Red Hat(红帽公司)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL

AI总结 本文提出首个通过合成多说话对话检测心理操控的研究,揭示语音与文本在检测准确性上的差异,强调多模态对话系统中模态意识的重要性。

Comments Accepted to IWSDS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21503 2026-01-14 cs.AI cs.CV 70%

MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions

MoHoBench: 通过无法回答的视觉问题评估多模态大语言模型的诚实性

Yanxu Zhu, Shitong Duan, Xiangxu Zhang, Jitao Sang, Peng Zhang, Tun Lu, Xiao Zhou, Jing Yao, Xiaoyuan Yi, Xing Xie

机构 * Beijing Jiaotong University(北京交通大学) Fudan University(复旦大学) Microsoft Research Asia(微软亚洲研究院) Renmin University of China(中国人民大学)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.AI

AI总结 MoHoBench通过评估多模态大语言模型在面对无法回答的视觉问题时的诚实行为,揭示其诚实性受视觉信息影响,提出改进方法以提升模型可信度。

Comments AAAI2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09583 2026-01-14 cs.LG stat.ML 70%

YRC-Bench: A Benchmark for Learning to Coordinate with Experts

YRC-Bench:一种学习与专家协调的基准

Mohamad H. Danesh, Nguyen X. Khanh, Tu Trinh, Benjamin Plaut

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.LG

AI总结 YRC-Bench通过提供开放源代码基准,支持在无监督环境下学习与专家协调的方法研究。

Comments Accepted at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08155 2026-01-14 cs.CV 67%

Instance-Aligned Captions for Explainable Video Anomaly Detection

实例对齐的描述用于可解释的视频异常检测

Inpyo Song, Minjun Joo, Joonhyung Kwon, Eunji Jeon, Jangwon Lee

机构 * SungKyunKwan University(顺 Kyun 峰大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

AI总结 本文提出实例对齐的描述方法,用于提升视频异常检测的可解释性和可信度,通过空间定位和实例关联增强解释的可验证性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08490 2026-01-14 cs.CL cs.AI 62%

BenchOverflow: Measuring Overflow in Large Language Models via Plain-Text Prompts

BenchOverflow: 通过纯文本提示测量大语言模型中的溢出现象

Erin Feiglin, Nir Hutnik, Raz Lapid

机构 * Deepkeep(深保持)

专题命中 安全评测 :prompt injection(abstract);分类 cs.CL、cs.AI

AI总结 BenchOverflow通过纯文本提示策略评估大语言模型的溢出现象,揭示长度控制对可靠性、成本和可持续性的影响,提供标准化比较框架以优化部署和防御措施。

Comments Accepted at TMLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08196 2026-01-14 cs.CL cs.AI cs.CR cs.LO cs.SE 62%

Evaluating Implicit Regulatory Compliance in LLM Tool Invocation via Logic-Guided Synthesis

通过逻辑引导合成评估LLM工具调用中的隐式监管合规性

Da Song, Yuheng Huang, Boqi Chen, Tianshuo Cong, Randy Goebel, Lei Ma, Foutse Khomh

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出LogiSafetyGen框架,通过逻辑引导合成评估LLM在工具调用中的隐式监管合规性,揭示大模型在安全约束上的不足。

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07871 2026-01-14 q-bio.QM cs.AI cs.CV cs.LG 62%

Imaging-anchored Multiomics in Cardiovascular Disease: Integrating Cardiac Imaging, Bulk, Single-cell, and Spatial Transcriptomics

心血管疾病中的成像锚定多组学:整合心脏成像、批量、单细胞和空间转录组学

Minh H. N. Le, Tuan Vinh, Thanh-Huy Nguyen, Tao Li, Bao Quang Gia Le, Han H. Huynh, Monika Raj, Carl Yang, Min Xu, Nguyen Quoc Khanh Le

机构 * International Ph.D. Program in Medicine, College of Medicine, Taipei Medical University, Taipei, Taiwan AIBioMed Research Group, Taipei Medical University, Taipei, Taiwan Medical Sciences Division, University of Oxford, Oxford, United Kingdom Computational Biology Department, School of Computer Science, Carnegie Mellon University, Pittsburgh, PA, USA Department of Computer Science, Emory University, Atlanta, GA, USA Department of Chemistry, Emory University, Atlanta, GA, USA International Master Program for Translational Science, College of Medical Science Technology, Taipei Medical University, Taipei 110, Taiwan In-Service Master Program in Artificial Intelligence in Medicine, College of Medicine, Taipei Medical University, Taipei, Taiwan Translational Imaging Research Center, Taipei Medical University Hospital, Taipei, Taiwan

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过整合心脏成像与多组学数据,推动心血管疾病研究的多模态融合方法,提升疾病诊断和治疗的精准性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00549 2026-01-14 cs.LG cs.AI 62%

Robust Single-Agent Reinforcement Learning for Regional Traffic Signal Control Under Demand Fluctuations

鲁棒的单智能体强化学习用于应对需求波动的区域交通信号控制

Qiang Li, Jin Niu, Lina Yu

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种鲁棒的单智能体强化学习框架,用于应对交通需求波动的区域交通信号控制,通过集中决策和高效学习模型有效减少交通队列长度。

Comments A critical error in the methodology. The reported congestion control effects were not caused by the proposed signal timing optimization, but by an incorrect traffic volume scaling factor during evaluation. The traffic demand was not properly amplified, resulting in misleading performance gains. Due to the substantial nature of the error, completion of revisions is not feasible in the short term

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20315 2026-01-14 cs.CL cs.AI 62%

Arctic-Text2SQL-R1: Simple Rewards, Strong Reasoning in Text-to-SQL

Arctic-Text2SQL-R1: 简单奖励,强推理的文本到SQL

Zhewei Yao, Guoheng Sun, Lukasz Borchmann, Gaurav Nuti, Zheyu Shen, Minghang Deng, Bohan Zhai, Hao Zhang, Ang Li, Yuxiong He

机构 * Snowflake AI Research(Snowflake AI研究院) University of Maryland(马里兰大学) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 Arctic-Text2SQL-R1通过简单奖励机制和强化学习框架,在文本到SQL任务中实现高准确率和高效性,优于现有大型模型。

Comments 22 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08733 2026-01-14 cs.LG quant-ph 57%

A Novel Approach to Explainable AI with Quantized Active Ingredients in Decision Making

可解释AI的新方法:决策中的量化活性成分

A. M. A. S. D. Alagiyawanna, Asoka Karunananda, Thushari Silva, A. Mahasinghe

机构 * Department of Computational Mathematics University of Moratuwa Sri Lanka(计算数学系 卢特瓦大学 斯里兰卡) Department of Mathematics University of Colombo Sri Lanka(数学系 科布姆大学 斯里兰卡)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

AI总结 本文提出基于量子玻尔兹曼机和经典玻尔兹曼机的可解释AI框架,通过量化活性成分提升模型的可解释性和预测准确性。

Comments Accepted and published in IEEE 2025. This is the authors manuscript version; final version available at IEEE Xplore: https://ieeexplore.ieee.org/document/11318441

Journal ref Proceedings of the 2025 9th SLAAI International Conference on Artificial Intelligence (SLAAI-ICAI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06575 2026-01-14 cs.CL 57%

Are Emotions Arranged in a Circle? Geometric Analysis of Emotion Representations via Hyperspherical Contrastive Learning

情感是否呈圆形?通过超球体对比学习进行情感表示的几何分析

Yusuke Yamauchi, Akiko Aizawa

机构 * The University of Tokyo(东京大学) National Institute of Informatics(信息处理研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 本文通过超球体对比学习在语言模型中诱导圆形情感表示,揭示了环形模型在可解释性和鲁棒性与高维设置下的性能权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10352 2026-01-14 cs.CL 57%

Cross-Prompt Encoder for Low-Performing Languages

跨提示编码器用于表现不佳的语言

Beso Mikaberidze, Teimuraz Saghinadze, Simon Ostermann, Philipp Muller

机构 * Muskhelishvili Institute of Computational Mathematics, GTU (MICM)(穆斯赫利什维利计算数学研究所(MICM)) Deutsches Forschungszentrum für Künstliche Intelligenz (DFKI)(德国人工智能研究中心(DFKI)) Center for European Research in Trusted AI (CERTAIN)(可信人工智能欧洲研究中心(CERTAIN)) Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 本文提出跨提示编码器(XPE)用于提升表现不佳语言的性能,并结合双软提示机制增强多语言适应能力。

Comments Accepted at Findings of IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04295 2026-01-14 cs.SE cs.AI 57%

EvoC2Rust: A Skeleton-guided Framework for Project-Level C-to-Rust Translation

EvoC2Rust:一个指导骨架的项目级C到Rust翻译框架

Chaofan Wang, Tingrui Yu, Beijun Shen, Jie Wang, Dong Chen, Wenrui Zhang, Yuling Shi, Chen Xie, Xiaodong Gu

机构 * Shanghai Jiao Tong University(上海交通大学) Huawei Technologies Co., Ltd(华为技术有限公司)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 EvoC2Rust通过结合规则和LLM方法,提升项目级C到Rust翻译的准确性和安全性

Comments Accepted by ICSE 2026 SEIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08241 2026-01-14 cs.CV cs.DC 50%

Improving Zero-shot ADL Recognition with Large Language Models through Event-based Context and Confidence

通过基于事件的上下文和置信度提升零样本ADL识别

Michele Fiori, Gabriele Civitarese, Marco Colussi, Claudio Bettini

机构 * Dept. of Computer Science University of Milan, Milan, Italy(计算机科学系米兰大学)

专题命中 安全评测 :safety(abstract)

AI总结 本文提出通过基于事件的上下文分割和新的置信度估计方法,改进零样本ADL识别,实验表明其在复杂数据集上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏