arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-14 至 2026-01-14 共收录 55 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 19 篇

2601.07871 2026-01-14 q-bio.QM cs.AI cs.CV cs.LG 62%

Imaging-anchored Multiomics in Cardiovascular Disease: Integrating Cardiac Imaging, Bulk, Single-cell, and Spatial Transcriptomics

心血管疾病中的成像锚定多组学:整合心脏成像、批量、单细胞和空间转录组学

Minh H. N. Le, Tuan Vinh, Thanh-Huy Nguyen, Tao Li, Bao Quang Gia Le, Han H. Huynh, Monika Raj, Carl Yang, Min Xu, Nguyen Quoc Khanh Le

机构 * International Ph.D. Program in Medicine, College of Medicine, Taipei Medical University, Taipei, Taiwan AIBioMed Research Group, Taipei Medical University, Taipei, Taiwan Medical Sciences Division, University of Oxford, Oxford, United Kingdom Computational Biology Department, School of Computer Science, Carnegie Mellon University, Pittsburgh, PA, USA Department of Computer Science, Emory University, Atlanta, GA, USA Department of Chemistry, Emory University, Atlanta, GA, USA International Master Program for Translational Science, College of Medical Science Technology, Taipei Medical University, Taipei 110, Taiwan In-Service Master Program in Artificial Intelligence in Medicine, College of Medicine, Taipei Medical University, Taipei, Taiwan Translational Imaging Research Center, Taipei Medical University Hospital, Taipei, Taiwan

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过整合心脏成像与多组学数据,推动心血管疾病研究的多模态融合方法,提升疾病诊断和治疗的精准性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00549 2026-01-14 cs.LG cs.AI 62%

Robust Single-Agent Reinforcement Learning for Regional Traffic Signal Control Under Demand Fluctuations

鲁棒的单智能体强化学习用于应对需求波动的区域交通信号控制

Qiang Li, Jin Niu, Lina Yu

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种鲁棒的单智能体强化学习框架,用于应对交通需求波动的区域交通信号控制,通过集中决策和高效学习模型有效减少交通队列长度。

Comments A critical error in the methodology. The reported congestion control effects were not caused by the proposed signal timing optimization, but by an incorrect traffic volume scaling factor during evaluation. The traffic demand was not properly amplified, resulting in misleading performance gains. Due to the substantial nature of the error, completion of revisions is not feasible in the short term

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20315 2026-01-14 cs.CL cs.AI 62%

Arctic-Text2SQL-R1: Simple Rewards, Strong Reasoning in Text-to-SQL

Arctic-Text2SQL-R1: 简单奖励,强推理的文本到SQL

Zhewei Yao, Guoheng Sun, Lukasz Borchmann, Gaurav Nuti, Zheyu Shen, Minghang Deng, Bohan Zhai, Hao Zhang, Ang Li, Yuxiong He

机构 * Snowflake AI Research(Snowflake AI研究院) University of Maryland(马里兰大学) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 Arctic-Text2SQL-R1通过简单奖励机制和强化学习框架,在文本到SQL任务中实现高准确率和高效性,优于现有大型模型。

Comments 22 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08733 2026-01-14 cs.LG quant-ph 57%

A Novel Approach to Explainable AI with Quantized Active Ingredients in Decision Making

可解释AI的新方法:决策中的量化活性成分

A. M. A. S. D. Alagiyawanna, Asoka Karunananda, Thushari Silva, A. Mahasinghe

机构 * Department of Computational Mathematics University of Moratuwa Sri Lanka(计算数学系 卢特瓦大学 斯里兰卡) Department of Mathematics University of Colombo Sri Lanka(数学系 科布姆大学 斯里兰卡)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

AI总结 本文提出基于量子玻尔兹曼机和经典玻尔兹曼机的可解释AI框架,通过量化活性成分提升模型的可解释性和预测准确性。

Comments Accepted and published in IEEE 2025. This is the authors manuscript version; final version available at IEEE Xplore: https://ieeexplore.ieee.org/document/11318441

Journal ref Proceedings of the 2025 9th SLAAI International Conference on Artificial Intelligence (SLAAI-ICAI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06575 2026-01-14 cs.CL 57%

Are Emotions Arranged in a Circle? Geometric Analysis of Emotion Representations via Hyperspherical Contrastive Learning

情感是否呈圆形?通过超球体对比学习进行情感表示的几何分析

Yusuke Yamauchi, Akiko Aizawa

机构 * The University of Tokyo(东京大学) National Institute of Informatics(信息处理研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 本文通过超球体对比学习在语言模型中诱导圆形情感表示,揭示了环形模型在可解释性和鲁棒性与高维设置下的性能权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10352 2026-01-14 cs.CL 57%

Cross-Prompt Encoder for Low-Performing Languages

跨提示编码器用于表现不佳的语言

Beso Mikaberidze, Teimuraz Saghinadze, Simon Ostermann, Philipp Muller

机构 * Muskhelishvili Institute of Computational Mathematics, GTU (MICM)(穆斯赫利什维利计算数学研究所(MICM)) Deutsches Forschungszentrum für Künstliche Intelligenz (DFKI)(德国人工智能研究中心(DFKI)) Center for European Research in Trusted AI (CERTAIN)(可信人工智能欧洲研究中心(CERTAIN)) Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 本文提出跨提示编码器(XPE)用于提升表现不佳语言的性能,并结合双软提示机制增强多语言适应能力。

Comments Accepted at Findings of IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04295 2026-01-14 cs.SE cs.AI 57%

EvoC2Rust: A Skeleton-guided Framework for Project-Level C-to-Rust Translation

EvoC2Rust:一个指导骨架的项目级C到Rust翻译框架

Chaofan Wang, Tingrui Yu, Beijun Shen, Jie Wang, Dong Chen, Wenrui Zhang, Yuling Shi, Chen Xie, Xiaodong Gu

机构 * Shanghai Jiao Tong University(上海交通大学) Huawei Technologies Co., Ltd(华为技术有限公司)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 EvoC2Rust通过结合规则和LLM方法,提升项目级C到Rust翻译的准确性和安全性

Comments Accepted by ICSE 2026 SEIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08241 2026-01-14 cs.CV cs.DC 50%

Improving Zero-shot ADL Recognition with Large Language Models through Event-based Context and Confidence

通过基于事件的上下文和置信度提升零样本ADL识别

Michele Fiori, Gabriele Civitarese, Marco Colussi, Claudio Bettini

机构 * Dept. of Computer Science University of Milan, Milan, Italy(计算机科学系米兰大学)

专题命中 安全评测 :safety(abstract)

AI总结 本文提出通过基于事件的上下文分割和新的置信度估计方法,改进零样本ADL识别,实验表明其在复杂数据集上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 3 篇

2505.13565 2026-01-14 cs.CY cs.AI cs.HC 81%

Aligning Trustworthy AI with Democracy: A Dual Taxonomy of Opportunities and Risks

将可信AI与民主对齐:机会与风险的双重分类

Oier Mentxaka, Natalia Díaz-Rodríguez, Mark Coeckelbergh, Marcos López de Prado, Emilia Gómez, David Fernández Llorca, Enrique Herrera-Viedma, Francisco Herrera

机构 * Dept. of Computer Science and Artificial Intelligence, DaSCI, University of Granada, Spain(计算机科学与人工智能系,DaSCI,格拉纳达大学) Dept. of Philosophy, University of Vienna, Vienna, Austria(哲学系,维也纳大学) School of Engineering, Cornell University, Ithaca, NY, United States(工程学院,康奈尔大学) Dept. of Mathematics, Khalifa University of Science and Technology, Abu Dhabi, UAE(数学系,科学与技术大学,阿布扎赫德,阿联酋) ADIA Lab, Al Maryah Island, Abu Dhabi, UAE(ADIA实验室,阿布扎赫德,阿联酋) Joint Research Centre, European Commission, Seville, Spain(联合研究中心,欧洲委员会,塞维利亚,西班牙) Computer Engineering Dept., University of Alcalá, Alcalá de Henares, Spain(计算机工程系,阿尔卡拉大学)

专题命中 AI治理与伦理 :trustworthy(title,abstract);分类 cs.AI、cs.CY

AI总结 本文提出双重分类框架,评估AI对民主的风险与机遇,为可信民主AI的发展提供规范性指导。

Comments 26 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09114 2026-01-14 cs.AI cs.CY 62%

AI TIPS 2.0: A Comprehensive Framework for Operationalizing AI Governance

AI TIPS 2.0: 一个全面的AI治理实施框架

Pamela Gupta

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

AI总结 AI TIPS 2.0是一个全面的AI治理实施框架,旨在解决AI部署中的三个关键治理挑战,通过提供定制化的治理方法和可操作的控制措施,提升AI系统的可信度和合规性。

Comments We have adjustments to make for higher effectiveness

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07954 2026-01-14 cs.CL 57%

A Human-Centric Pipeline for Aligning Large Language Models with Chinese Medical Ethics

面向中文医疗伦理的以人为本的大型语言模型对齐流水线

Haoan Jin, Han Ying, Jiacheng Ji, Hanhui Xu, Mengyue Wu

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 本文提出MedES基准和 guardian-in-the-loop 框架,通过监督微调和偏好优化,实现中文医疗伦理场景下的LLM对齐,提升伦理任务表现。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 14 篇

2505.16237 2026-01-14 cs.CL 79%

Align-GRAG: Anchor and Rationale Guided Dual Alignment for Graph Retrieval-Augmented Generation

Align-GRAG: 基于锚点和推理引导的双 Alignment 图检索增强生成

Derong Xu, Pengyue Jia, Xiaopeng Li, Yingyi Zhang, Maolin Wang, Qidong Liu, Xiangyu Zhao, Yichao Wang, Huifeng Guo, Ruiming Tang, Enhong Chen, Tong Xu

机构 * University of Science and Technology of China(中国科学技术大学) City University of Hong Kong(香港城市大学) Noah’s Ark Lab, Huawei(华为诺亚实验室)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

AI总结 Align-GRAG通过锚点和推理引导的双 Alignment 框架,提升图检索增强生成的准确性与语义对齐能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08139 2026-01-14 cs.CV cs.AI 74%

Subspace Alignment for Vision-Language Model Test-time Adaptation

子空间对齐用于视觉-语言模型测试时适应

Zhichen Zeng, Wenxuan Bao, Xiao Lin, Ruizhong Qiu, Tianxin Wei, Xuying Ning, Yuchen Yan, Chen Luo, Monica Xiao Cheng, Jingrui He, Hanghang Tong

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊)

专题命中 其他安全 :alignment(title);分类 cs.AI

AI总结 SubTTA通过子空间对齐提升视觉-语言模型测试时适应性能,有效解决模态差距和视觉噪声问题。

Comments 17 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04563 2026-01-14 cs.LG cs.AI cs.CL cs.CV 67%

A Vision for Multisensory Intelligence: Sensing, Science, and Synergy

多感官智能的愿景:感知、科学与协同

Paul Pu Liang

机构 * MIT Media Lab and MIT EECS(MIT媒体实验室和MIT电子工程与计算机科学系)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出未来十年多感官人工智能的研究愿景,强调通过感知、科学和协同三大主题推动多感官技术发展,提升人与AI的交互体验。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08634 2026-01-14 cs.CL cs.AI 62%

Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs

道德透镜,政治坐标:迈向道德条件化大语言模型的意识形态定位

Chenchen Yuan, Bolei Ma, Zheyu Zhang, Bardh Prenkaj, Frauke Kreuter, Gjergji Kasneci

机构 * Technical University of Munich(慕尼黑技术大学) LMU Munich(慕尼黑大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本研究通过道德条件化大语言模型,探索道德价值观与政治倾向之间的因果关系,揭示道德取向对政治坐标的影响,并提出更社会化的对齐方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08169 2026-01-14 cs.CL cs.LG 62%

Relational Knowledge Distillation Using Fine-tuned Function Vectors

基于微调函数向量的关系知识蒸馏

Andrea Kang, Yingnian Wu, Hongjing Lu

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

AI总结 通过微调函数向量提升关系知识表示,增强语言模型的类比推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08079 2026-01-14 cs.AI cs.CL cs.IR 62%

MemoBrain: Executive Memory as an Agentic Brain for Reasoning

MemoBrain: 作为推理代理的执行记忆脑

Hongjin Qian, Zhao Cao, Zheng Liu

机构 * Beijing Academy of Artificial Intelligence(北京人工智能研究院) Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 MemoBrain通过构建依赖感知的记忆模型,实现对长周期推理轨迹的显式认知控制,提升工具增强代理的推理连贯性和任务对齐性。

Comments Our codes are in https://github.com/qhjqhj00/MemoBrain

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07892 2026-01-14 cs.LG cs.AI 62%

Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization via Fine-grained Sparsification

Sherry:通过细粒度稀疏化实现高效的1.25位三元量化

Hong Huang, Decheng Wu, Qiangqiang Hu, Guanghua Yu, Jinhai Yang, Jianchen Zhu, Xue Liu, Dapeng Wu

机构 * City University of Hong Kong(香港城市大学) Tencent(腾讯) McGill University(麦吉尔大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 Sherry通过细粒度稀疏化实现高效的1.25位三元量化,显著减少模型大小并提升推理速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04954 2026-01-14 cs.LG cs.AI 62%

Precision over Diversity: High-Precision Reward Generalizes to Robust Instruction Following

精确性胜过多样性:高精度奖励能泛化到稳健的指令跟随

Yirong Zeng, Yufei Liu, Xiao Ding, Yutai Hou, Yuxian Wang, Haonan Song, Wu Ning, Dandan Tu, Qixun Zhang, Bibo Cai, Yuxiang He, Ting Liu

机构 * Harbin Institute of Technology, SCIR(哈尔滨工业大学,SCIR) Peking University(北京大学) Huawei Technologies Co., Ltd(华为技术有限公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本研究发现高精度奖励优于多样化的约束混合,提出数据导向的优化策略,提升指令跟随性能并减少训练时间。

Comments Under review, 13 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04948 2026-01-14 cs.CL cs.AI 62%

KaLM: Knowledge-aligned Autoregressive Language Modeling via Dual-view Knowledge Graph Contrastive Learning

KaLM: 通过双视角知识图谱对比学习实现知识对齐的自回归语言模型

Peng Yu, Cheng Deng, Beiya Dai, Xinbing Wang, Ying Wen

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 KaLM通过双视角知识图谱对比学习和三元组完成语言模型,实现LLMs与知识图谱的对齐,提升知识驱动任务的性能。

Comments The article has been accepted by Frontiers of Computer Science (FCS), with the DOI: {10.1007/s11704-026-50906-6}

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08401 2026-01-14 cs.CV cs.AI 57%

An Explainable Two Stage Deep Learning Framework for Pericoronitis Assessment in Panoramic Radiographs Using YOLOv8 and ResNet-50

可解释的两阶段深度学习框架用于全景X光片中牙周炎评估:使用YOLOv8和ResNet-50

Ajo Babu George, Pranav S, Kunal Agarwal

机构 * DiceMed College of Engineering Trivandrum(工程学院) S.C.B Dental College and Hospital(S.C.B牙科学院和医院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出了一种结合YOLOv8和ResNet-50的两阶段深度学习框架,用于可解释性地评估全景X光片中的牙周炎。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08017 2026-01-14 cs.CV cs.AI 57%

Representations of Text and Images Align From Layer One

文本和图像的表示从第一层对齐

Evžen Wybitul, Javier Rando, Florian Tramèr, Stanislav Fort

机构 * D-INFK, ETH Zurich, Switzerland(苏黎世联邦理工学院信息与知识系统研究所) Aisle Research(Aisle研究)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 该研究通过合成方法证明,视觉-语言模型中图像和文本表示在第一层即可实现对齐,为模型可解释性提供了新路径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07939 2026-01-14 cs.SE cs.AI 57%

SECite: Analyzing and Summarizing Citations in Software Engineering Literature

SECite: 分析和总结软件工程文献中的引用

Shireesh Reddy Pyreddy, Khaja Valli Pathan, Hasan Masum, Tarannum Shaila Zaman

机构 * Dept. of Computer Science SUNY Polytechnic Institute(计算机科学系圣尼古拉学院) Dept. of Information Systems University of Maryland Baltimore County(信息系统系马里兰大学巴尔的摩县)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 SECite通过分析文献引用中的情感倾向,结合生成式AI生成摘要,提供了一种评估学术贡献的综合框架。

Comments Accepted at IEEE CCWC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07976 2026-01-14 eess.IV cs.CV eess.SP physics.med-ph 50%

Application of Ideal Observer for Thresholded Data in Search Task

理想观察者在阈值化数据中的应用:搜索任务中的图像质量评估

Hongwei Lin, Howard C. Gifford

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出了一种基于阈值化数据的类人视觉搜索模型观察者,通过选择性处理高显著性特征提升图像质量评估性能,适用于临床真实任务和资源受限的模型训练。

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03486 2026-01-14 eess.SY cs.SY 50%

Adaptive Model-Based Reinforcement Learning for Orbit Feedback Control in NSLS-II Storage Ring

自适应模型驱动强化学习用于NSLS-II存储环轨道反馈控制

Zeyu Dong, Yuke Tian, Yu Sun

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出基于模型驱动强化学习的自适应框架,用于NSLS-II存储环轨道反馈控制,通过轨迹优化和在线模型优化实现束流稳定与对准误差最小化。

Comments Accepted by the 20th International Conference on Accelerator and Large Experimental Physics Control Systems (ICALEPCS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏