arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-02-10 至 2026-02-10 共收录 21 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 21 篇

2602.07540 2026-02-10 cs.CV cs.LG 79%

LLM-Guided Diagnostic Evidence Alignment for Medical Vision-Language Pretraining under Limited Pairing

基于LLM的诊断证据对齐用于有限配对情况下的医学视觉-语言预训练

Huimin Yan, Liang Bai, Xian Yang, Long Chen

机构 * Institute of Intelligent Information Processing, Shanxi University(山西大学智能信息处理研究所) Alliance Manchester Business School, The University of Manchester(曼彻斯特大学阿利安斯商学院) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

AI总结 本文提出LLM引导的诊断证据对齐方法,旨在解决有限配对数据下医学视觉-语言预训练的诊断表示学习问题,通过提取关键诊断证据提升跨模态对齐效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12215 2026-02-10 cs.CL 79%

GMSA: Enhancing Context Compression via Group Merging and Layer Semantic Alignment

GMSA:通过组合并和层语义对齐增强上下文压缩

Jiwei Tang, Zhicheng Zhang, Shunlong Wu, Jingheng Ye, Lichen Bai, Zitai Wang, Tingwei Lu, Lin Hai, Yiming Zhao, Hai-Tao Zheng, Hong-Gee Kim

机构 * Tsinghua University(清华大学) Pengcheng Laboratory(鹏城实验室) Seoul National University(首尔国立大学) Sun Yat-sen University(中山大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

AI总结 GMSA通过组合并和层语义对齐技术,提升长上下文场景下的上下文压缩效率,实现更高效的模型压缩与下游任务性能提升。

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08559 2026-02-10 cs.IR 71%

QARM V2: Quantitative Alignment Multi-Modal Recommendation for Reasoning User Sequence Modeling

QARM V2:量化对齐多模态推荐用于推理用户序列建模

Tian Xia, Jiaqi Zhang, Yueyang Liu, Hongjian Dou, Tingya Yin, Jiangxia Cao, Xulei Liang, Tianlu Xie, Lihao Liu, Xiang Chen, Shen Wang, Changxin Lao, Haixiang Gan, Jinkai Yu, Keting Cen, Lu Hao, Xu Zhang, Qiqiang Zhong, Zhongbo Sun, Yiyu Wang, Shuang Yang, Mingxin Wen, Xiangyu Wu, Shaoguo Liu, Tingting Gao, Zhaojie Liu, Han Li, Kun Gai

专题命中 其他安全 :alignment(title)

AI总结 QARM V2通过量化对齐多模态推荐方法,解决推荐系统中用户序列建模的语义理解与业务需求不匹配问题。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08630 2026-02-10 cs.AI cs.CC 70%

Debate is efficient with your time

与时间高效地进行辩论

Jonah Brown-Cohen, Geoffrey Irving, Simon C. Marshall, Ilan Newman, Georgios Piliouras, Mario Szegedy

机构 * Google DeepMind(谷歌DeepMind) UK AI Security Institute(英国人工智能安全研究所) University of Haifa(海法大学) Rutgers University(罗格斯大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 通过辩论实现人工智能安全,证明辩论查询复杂度与电路复杂性密切相关。

Comments 11 Pages, 0 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07547 2026-02-10 q-bio.NC cs.AI cs.CL cs.LG 67%

Linguistic properties and model scale in brain encoding: from small to compressed language models

语言属性与模型规模在脑编码中的作用:从小型到压缩语言模型

Subba Reddy Oota, Vijay Rowtula, Satya Sai Srinath Namburi, Khushbu Pahwa, Anant Khandelwal, Manish Gupta, Tanmoy Chakraborty, Bapi S. Raju

机构 * TU Berlin(柏林技术大学) IIIT-Hyderabad(海得拉巴理工学院) GE HealthCare(通用电气医疗) AWS AI Labs, Amazon(亚马逊AI实验室) Microsoft Research, Bangalore, India(微软研究院,班加罗尔,印度) Microsoft, Hyderabad, India(微软,海得拉巴,印度) IIT Delhi(德里理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究发现小型语言模型在脑一致性上可与大模型媲美,压缩不影响脑可预测性,挑战了神经扩展的常见假设。

Comments 40 pages, 33 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07276 2026-02-10 cs.AI cs.CL cs.LG 67%

Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs

Steer2Adapt:动态组合转向向量引发LLM高效适应

Pengrui Han, Xueqiang Xu, Keyang Xuan, Peiyang Song, Siru Ouyang, Runchu Tian, Yuqing Jiang, Cheng Qian, Pengcheng Jiang, Jiashuo Sun, Junxia Cui, Ming Zhong, Ge Liu, Jiawei Han, Jiaxuan You

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 STEER2ADAPT通过动态组合转向向量实现LLM高效适应,适用于需要多协调能力的复杂任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09003 2026-02-10 cs.AI cs.CL 62%

Data Science and Technology Towards AGI Part I: Tiered Data Management

数据科学与技术迈向通用人工智能 Part I:分层数据管理

Yudong Wang, Zixuan Fu, Hengyu Zhao, Chen Zhao, Chuyue Zhou, Xinle Lin, Hongya Lyu, Shuaikang Xue, Yi Yi, Yingjiao Wang, Zhi Zheng, Yuzhou Zhang, Jie Zhou, Chaojun Xiao, Xu Han, Zhiyuan Liu, Maosong Sun

机构 * Tsinghua University(清华大学) ModelBest Inc.(ModelBest公司) Beijing Institute of Technology(北京理工大学) South China Agricultural University(华南农业大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出分层数据管理框架,通过数据与模型的协同进化提升大语言模型训练效率和性能。

Comments 16 pages, 3 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03041 2026-02-10 cs.LG cs.AI 62%

Optimas: Optimizing Compound AI Systems with Globally Aligned Local Rewards

Optimas: 通过全局对齐的局部奖励优化复合AI系统

Shirley Wu, Parth Sarthi, Shiyu Zhao, Aaron Lee, Herumb Shandilya, Adrian Mladenic Grobelnik, Nurendra Choudhary, Eddie Huang, Karthik Subbian, Linjun Zhang, Diyi Yang, James Zou, Jure Leskovec

机构 * Stanford University(斯坦福大学) Amazon(亚马逊) Jožef Stefan Institute(乔泽夫·斯蒂芬研究所) Rutgers University(罗格斯大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 Optimas通过全局对齐的局部奖励优化复合AI系统,有效提升复合系统的性能。

Comments Accepted to ICLR 2026. 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07338 2026-02-10 cs.CL cs.AI 62%

Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation

意图不匹配导致大语言模型在多轮对话中迷失

Geng Liu, Fei Zhu, Rong Feng, Changyi Ma, Shiqi Wang, Gaofeng Meng

机构 * College of Computing, City University of Hong Kong(香港城市大学计算机学院) Centre for Artificial Intelligence and Robotics, HKISI, CAS(中国科学院香港中文大学人工智能与机器人中心) School of Artificial Intelligence, Jilin University(吉林大学人工智能学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出调解员-助手架构,通过解耦意图理解和任务执行,解决大语言模型在多轮对话中因意图不匹配导致的性能下降问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08295 2026-02-10 cs.AI 57%

The Vibe-Automation of Automation: A Proactive Education Framework for Computer Science in the Age of Generative AI

自动化之自动化:生成式人工智能时代计算机科学的前瞻性教育框架

Ilya Levin

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出Vibe-Automation概念,探讨生成式AI时代计算机科学教育的前瞻性框架,强调隐性规律的操作化与教育体系的变革。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08208 2026-02-10 cs.CL cs.HC 57%

LLMs and people both learn to form conventions -- just not with each other

大语言模型和人类都学会形成惯例——但并不是彼此之间

Cameron R. Jones, Agnese Lombardi, Kyle Mahowald, Benjamin K. Bergen

机构 * Department of Psychology, Stony Brook University(心理学系,石溪大学) Department of Cognitive Science, University of California San Diego(认知科学系,加州圣地亚哥大学) Department of Philology, Literature, and Linguistics, University of Pisa(philology、文学与语言学系,比萨大学) Department of Linguistics, University of Texas at Austin(语言学系,德克萨斯大学奥斯汀分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 研究发现人类和AI在同类型对话中能形成惯例,但人机对话效果较差,表明对话协调需要共同的解释偏见。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04413 2026-02-10 cs.LG quant-ph 57%

Distribution-Guided and Constrained Quantum Machine Unlearning

基于分布和约束的量子机器学习去学习

Nausherwan Malik, Zubair Khalid, Muhammad Faryad

机构 * Department of Electrical Engineering, Lahore University of Management Sciences, Lahore 54792, Pakistan.(电气工程系,拉合尔管理科学大学,巴基斯坦拉合尔) Department of Physics, Lahore University of Management Sciences, Lahore 54792, Pakistan.(物理系,拉合尔管理科学大学,巴基斯坦拉合尔)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本文提出了一种基于分布和约束的量子机器学习去学习框架,通过可调节的目标分布和保留约束,有效控制遗忘与保留行为之间的权衡,提升去学习的可靠性与可解释性。

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08099 2026-02-10 cs.CV cs.AI 57%

VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval

VidVec:解锁视频MLLM嵌入用于视频-文本检索

Issar Tzachor, Dvir Samuel, Rami Ben-Ari

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 VidVec通过利用预训练MLLM的中间层嵌入和校准头部,实现无需训练的视频-文本检索,超越现有方法,达到最佳性能。

Comments Project page: https://iyttor.github.io/VidVec/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07790 2026-02-10 cs.LG 57%

MaD-Mix: Multi-Modal Data Mixtures via Latent Space Coupling for Vision-Language Model Training

MaD-Mix: 通过潜在空间耦合实现多模态数据混合用于视觉-语言模型训练

Wanyun Xie, Francesco Tonin, Volkan Cevher

机构 * LIONS, EPFL(EPFL 雷奥尼实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 MaD-Mix通过潜在空间耦合实现多模态数据混合,提升视觉-语言模型训练效率,减少训练步骤并增强复杂场景下的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07378 2026-02-10 cs.LG physics.data-an stat.ML 57%

Dichotomy of Feature Learning and Unlearning: Fast-Slow Analysis on Neural Networks with Stochastic Gradient Descent

特征学习与遗忘的二元性:基于随机梯度下降的神经网络快慢分析

Shota Imai, Sota Nishiyama, Masaaki Imaizumi

机构 * The University of Tokyo(东京大学) RIKEN Center for Advanced Intelligence Project(RIKEN先进人工智能项目中心)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本文研究了神经网络中特征学习与遗忘的二元性,通过快慢动态分析揭示特征遗忘的机制与条件,提出特征遗忘由数据主非线性项强度和第二层权重初始尺度决定。

Comments 40 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18727 2026-02-10 cs.LG 57%

LogSyn: A Few-Shot LLM Framework for Structured Insight Extraction from Unstructured General Aviation Maintenance Logs

LogSyn: 一种基于少样本学习的LLM框架,用于从非结构化通用航空维护日志中提取结构化洞察

Devansh Agarwal, Maitreyi Chatterjee, Biplab Chatterjee

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 LogSyn通过少样本学习利用LLM将非结构化航空维护日志转化为结构化数据,实现故障模式识别与事件分类,提升航空维护流程和预测分析能力。

Comments Accepted in Proceedings of the 3rd INCOM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08571 2026-02-10 cs.RO 50%

Head-to-Head autonomous racing at the limits of handling in the A2RL challenge

极限操控下的头对头自动驾驶赛车挑战

Simon Hoffmann, Simon Sagmeister, Tobias Betz, Joscha Bongard, Sascha Büttner, Dominic Ebner, Daniel Esser, Georg Jank, Sven Goblirsch, Alexander Langmann, Maximilian Leitenstern, Levent Ögretmen, Phillip Pitschi, Ann-Kathrin Schwehn, Cornelius Schröder, Marcel Weinmann, Frederik Werner, Boris Lohmann, Johannes Betz, Markus Lienkamp

机构 * Institute of Automotive Technology, TUM(汽车技术研究所,慕尼黑技术大学) Chair of Automatic Control, TUM(自动控制教授座,慕尼黑技术大学) Professorship Autonomous Vehicle Systems, TUM(自主车辆系统教授职位,慕尼黑技术大学)

专题命中 其他安全 :safety(abstract)

AI总结 TUM团队通过模拟人类驾驶行为,在极限操控下开发算法赢得A2RL赛事,提升自动驾驶技术与道路安全。

Comments Submitted to Science Robotics for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08505 2026-02-10 cs.CV 50%

Are Vision Foundation Models Foundational for Electron Microscopy Image Segmentation?

视觉基础模型是否对电子显微镜图像分割具有基础性作用?

Caterina Fuster-Barceló, Virginie Uhlmann

机构 * Department of Molecular Life Sciences, Universität Zürich(分子生命科学系,苏黎世大学)

专题命中 其他安全 :alignment(abstract)

AI总结 本文研究了视觉基础模型在电子显微镜图像分割中的有效性,发现LoRA能提升单域性能,但多域训练导致性能下降,需额外领域对齐机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05825 2026-02-10 cs.HC 50%

ToMigo: Interpretable Design Concept Graphs for Aligning Generative AI with Creative Intent

ToMigo:可解释的设计概念图用于对齐生成式AI与创意意图

Lena Hegemann, Xinyi Wen, Michael A. Hedderich, Tarmo Nurmi, Hariharan Subramonyam

专题命中 其他安全 :alignment(abstract)

AI总结 ToMigo通过设计概念图实现生成式AI与创意意图的对齐,提供可解释的交互方式提升用户控制力。

Comments 18 pages, 10 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08004 2026-02-10 cs.SE cs.SI 50%

Agent Skills: A Data-Driven Analysis of Claude Skills for Extending Large Language Model Functionality

Agent Skills: 一种数据驱动的Claude技能分析,用于扩展大型语言模型功能

George Ling, Shanshan Zhong, Richard Huang

专题命中 其他安全 :safety(abstract)

AI总结 本文通过分析Claude技能的数据,揭示了技能在软件工程中的集中使用及安全风险,为技能重用和标准化提供参考。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07108 2026-02-10 physics.ao-ph astro-ph.EP astro-ph.SR 50%

Machine Learning-Ready Data Sets for the Analysis and Nowcasting of Atmospheric Radiation at Aviation Altitudes

为航空高度大气辐射分析和现在预报准备机器学习数据集

Viacheslav M Sadykov, Zachary M Watkins, Dustin Kempton, William Jones, Sanjib K C, Griffin T Goodwin, Xiaochun He, W Kent Tobiska, Irina Kitiashvili, Christopher Mertens, Shubha Ranjan, D Glenn Deardorff, Ryan Spaulding

专题命中 其他安全 :safety(abstract)

AI总结 本文提出为航空高度大气辐射分析和现在预报准备机器学习数据集,通过整合飞行数据和电离层环境参数,提升辐射环境预测的准确性。

Comments 24 pages, 8 figures, 1 table, accepted to Space Weather Journal

详情

展开后加载摘要…

URL PDF HTML 收藏