arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8057 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8057 篇

2603.02787 2026-03-04 cs.AI 57%

Rethinking Code Similarity for Automated Algorithm Design with LLMs

重新思考基于LLM的自动算法设计中的代码相似性

Rui Zhang, Zhichao Lu

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 BehaveSim通过分析问题解决行为轨迹,提升LLM-AAD性能并促进算法分析。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02553 2026-03-04 cs.RO cs.CV cs.HC cs.LG 57%

Give me scissors: Collision-Free Dual-Arm Surgical Assistive Robot for Instrument Delivery

给我剪刀:用于器械传递的碰撞自由双臂手术辅助机器人

Xuejin Luo, Shiquan Sun, Runshi Zhang, Ruizhi Zhang, Junchen Wang

机构 * School of Mechanical Engineering and Automation, Beihang University(机械工程与自动化学院,北航)

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 本文提出了一种碰撞自由的双臂手术辅助机器人,通过视觉-语言模型生成轨迹并实现动态环境中的安全器械传递。

Comments 8 pages, 10 figures. Accepted by IEEE International Conference on Robotics and Automation (ICRA), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23652 2026-03-04 cs.CV cs.AI 57%

3D Modality-Aware Pre-training for Vision-Language Model in MRI Multi-organ Abnormality Detection

面向MRI多器官异常检测的3D模态感知预训练

Haowen Zhu, Ning Yin, Xiaogen Zhou

机构 * School of Electronic, Electrical Engineering and Physics, Fujian University of Technology(福建工程学院电子电气工程学院) School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院) Department of Medical Imaging, Suzhou Traditional Chinese Medicine Hospital, China(苏州中医医院影像科)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出MedMAP框架,通过3D MRI的模态感知预训练提升多器官异常检测性能,实验表明其优于现有视觉-语言模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19791 2026-03-03 cs.CL cs.IR 57%

ToolDreamer: Instilling LLM Reasoning Into Tool Retrievers

ToolDreamer: 将大语言模型推理能力注入工具检索器

Saptarshi Sengupta, Zhengyu Zhou, Jun Araki, Xingbo Wang, Bingqing Wang, Suhang Wang, Zhe Feng

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Bosch Research North America(博世北美研究)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 ToolDreamer通过合成工具描述来改进工具检索,提升稀疏和密集检索器性能,减少LLM上下文窗口压力。

Comments Accepted to EACL 2026 (main/oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01822 2026-03-03 cs.AI 57%

Emerging Human-like Strategies for Semantic Memory Foraging in Large Language Models

新兴的人类样策略在大语言模型中的语义记忆采集

Eric Lacosse, Mariana Duarte, Peter M. Todd, Daniel C. McNamee

机构 * Champalimaud Research(恰帕拉德研究) Centre for Restorative Neurotechnology(修复性神经技术中心) Indiana University(印第安纳大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本研究通过分析大语言模型中的语义流畅任务,探索其语义记忆采集策略,揭示人类与AI在认知机制上的异同。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20089 2026-03-03 cs.CV cs.AI 57%

StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues

StructXLIP: 通过多模态结构线索增强视觉-语言模型

Zanxi Ruan, Songqun Gao, Qiuyu Kong, Yiming Wang, Marco Cristani

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 StructXLIP通过引入多模态结构线索提升视觉-语言模型的跨模态检索性能。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09463 2026-03-03 cs.AI 57%

SpotAgent: Grounding Visual Geo-localization in Large Vision-Language Models through Agentic Reasoning

SpotAgent:通过代理推理在大视觉语言模型中实现视觉地理定位

Furong Jia, Ling Dai, Wenjin Deng, Fan Zhang, Chen Hu, Daxin Jiang, Yu Liu

机构 * Peking University(北京大学) StepFun

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 SpotAgent通过代理推理提升大视觉语言模型在地理定位中的表现,结合外部工具验证和强化学习优化,实现更精确和可靠的定位结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04369 2026-03-03 cs.LG 57%

Multi-scale hypergraph meets LLMs: Aligning large language models for time series analysis

多尺度超图与大语言模型:面向时间序列分析的对齐方法

Zongjiang Shang, Dongliang Cui, Binqing Wu, Ling Chen

机构 * State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学) College of Computer Science and Technology, Zhejiang University(计算机科学与技术学院,浙江大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本文提出MSH-LLM方法,通过多尺度超图机制和跨模态对齐模块,提升大语言模型在时间序列分析中的表现。

Comments Accepted by ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00350 2026-03-03 cs.AI 57%

Monotropic Artificial Intelligence: Toward a Cognitive Taxonomy of Domain-Specialized Language Models

单调人工智能:面向领域专用语言模型的认知分类

Antonio de Sousa Leitão Filho, Allan Kardec Duailibe Barros Filho, Fabrício Saul Lima, Selby Mykael Lima dos Santos, Rejani Bandeira Vieira Sousa

机构 * Aia Context Federal University of Maranhão(马那瓜联邦大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 单调人工智能通过牺牲通用性实现领域内高精度,挑战通用智能主导的AI研究范式,提出专门化与通用系统互补共存的认知生态。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05064 2026-03-03 cs.LG 57%

Boomerang Distillation Enables Zero-Shot Model Size Interpolation

回环蒸馏使零样本模型大小插值

Sara Kangaslahti, Nihal V. Nayak, Jonathan Geuter, Marco Fumero, Francesco Locatello, David Alvarez-Melis

机构 * Harvard University(哈佛大学) Kempner Institute(凯普纳研究所) IST Austria(IST奥地利研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 回环蒸馏通过蒸馏和重构实现零样本模型大小插值,生成细粒度模型家族,降低训练成本并提升适应性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12490 2026-03-03 physics.ao-ph cs.LG 57%

SamudrACE: Fast and Accurate Coupled Climate Modeling with 3D Ocean and Atmosphere Emulators

SamudrACE:基于3D海洋和大气模拟器的快速准确耦合气候建模

James P. C. Duncan, Elynn Wu, Surya Dheeshjith, Adam Subel, Troy Arcomano, Spencer K. Clark, Brian Henn, Anna Kwa, Jeremy McGibbon, W. Andre Perkins, William Gregory, Carlos Fernandez-Granda, Julius Busecke, Oliver Watt-Meyer, William J. Hurlin, Alistair Adcroft, Laure Zanna, Christopher Bretherton

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 SamudrACE通过3D海洋和大气模拟器实现快速准确的耦合气候建模,能够模拟数百年长的高分辨率气候现象。

Comments 29 pages, 26 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18564 2026-03-03 cs.LG cs.CE 57%

Efficient Aircraft Design Optimization Using Multi-Fidelity Models and Multi-fidelity Physics Informed Neural Networks

利用多保真模型和多保真物理指导神经网络实现高效飞机设计优化

Apurba Sarker

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本研究利用多保真物理指导神经网络和生成对抗网络,实现高效飞机设计优化,提升设计迭代速度与经济性。

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00144 2026-03-03 cs.CV cs.AI 57%

Disentangled Hierarchical VAE for 3D Human-Human Interaction Generation

解耦层次变分自编码器用于3D人-人交互生成

Zichen Geng, Zeeshan Hayder, Bo Miao, Jian Liu, Wei Liu, Ajmal Mian

机构 * Department of CSSE, The University of Western Australia(西澳大学计算机科学与工程系) Data61, CSIRO(澳大利亚联邦科学工业研究组织Data61部门) Australian Institute for Machine Learning, The University of Adelaide(澳大利亚阿德莱德大学人工智能研究所) NERC-RVC, Hunan University(湖南大学NERC-RVC部门)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出DHVAE,通过解耦层次变分自编码器生成结构化且可控的3D人-人交互,提升运动保真度和物理合理性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24183 2026-03-02 cs.CV cs.LG 57%

A multimodal slice discovery framework for systematic failure detection and explanation in medical image classification

一种用于医学图像分类中系统性故障检测和解释的多模态切片发现框架

Yixuan Liu, Kanwal K. Bhatia, Ahmed E. Fetit

机构 * Department of Computing, Imperial College London, UK(帝国理工学院 computing 部,英国) Aival, London, UK(Aival,伦敦,英国)

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 该研究提出了一种多模态切片发现框架,用于医学图像分类中的系统性故障检测与解释,展示了其在故障发现和解释生成方面的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24097 2026-03-02 cs.AI 57%

Bi-level RL-Heuristic Optimization for Real-world Winter Road Maintenance

双层强化学习启发式优化用于现实世界冬季道路维护

Yue Xie, Zizhen Xu, William Beazley, Fumiya Iida

机构 * School of Science, Loughborough University, UK(洛辛厄姆大学科学学院) Department of Engineering, University of Cambridge, UK(剑桥大学工程系) National Highways, 199 Wharfside St, Birmingham B1 1RN, UK(国家公路局)

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 本研究提出双层强化学习优化方法,用于高效解决现实世界冬季道路维护中的大规模路由问题,实现资源优化与排放降低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06940 2026-03-02 cs.DB cs.AI 57%

VISTA: Knowledge-Driven Vessel Trajectory Imputation with Repair Provenance

VISTA: 基于知识的船舶轨迹补全与修复溯源

Hengyu Liu, Tianyi Li, Haoyu Wang, Kristian Torp, Tiancheng Zhang, Yushuai Li, Christian S. Jensen

机构 * Department of Computer Science, Aalborg University, Denmark(计算机科学系,奥胡斯大学,丹麦) School of Computer Science and Engineering, Northeastern University, Shenyang, China(计算机科学与工程学院,东北大学,沈阳,中国)

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 VISTA通过基于知识的可解释方法实现船舶轨迹补全,生成修复溯源以提升下游决策的可信度和效率。

Comments 24 pages, 14 figures, 4 algorithms, 8 tables. Code available at https://github.com/hyLiu1994/VISTA

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23336 2026-02-27 cs.LG stat.ML 57%

Differentiable Zero-One Loss via Hypersimplex Projections

通过超简单面投影实现可微零一损失

Camilo Gomez, Pengyang Wang, Liansheng Tang

机构 * School of Data, Mathematical, and Statistical Sciences, University of Central Florida, Orlando, USA(数据、数学与统计科学学院,中央佛罗里达大学,奥兰多,美国) Department of CIS, University of Macau, Macao, China(信息与系统系,澳门大学,澳门,中国)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本文提出了一种可微的零一损失近似方法,通过超简单面投影和Soft-Binary-Argmax操作符,提升大批次训练下的模型泛化能力。

Comments To appear in PAKDD 2026 (Pacific-Asia Conference on Knowledge Discovery and Data Mining), 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20570 2026-02-27 cs.CV cs.AI 57%

Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP

Dyslexify: 一种针对CLIP中印刷攻击的机制性防御

Lorenz Hufe, Constantin Venhoff, Erblina Purelku, Maximilian Dreyer, Sebastian Lapuschkin, Wojciech Samek

机构 * Fraunhofer Heinrich Hertz Institute(弗劳恩霍夫海因里希·赫兹研究所) University of Oxford(牛津大学) Technological University Dublin(都柏林技术大学) Technische Universität Berlin(柏林技术大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 Dyslexify通过消融CLIP中的印刷电路,有效防御印刷攻击,提升性能并保持应用安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22404 2026-02-27 cs.CL 57%

SAFARI: A Community-Engaged Approach and Dataset of Stereotype Resources in the Sub-Saharan African Context

SAFARI:一种社区参与的方法和撒哈拉以南非洲语境下的刻板印象资源数据集

Aishwarya Verma, Laud Ammah, Olivia Nercy Ndlovu Lucas, Andrew Zaldivar, Vinodkumar Prabhakaran, Sunipa Dev

机构 * Google Research(谷歌研究) RAIN Africa(RAIN非洲) Mantaray Africa(Mantaray非洲)

专题命中 其他安全 :safety(abstract);分类 cs.CL

AI总结 SAFARI通过社区参与方法构建了覆盖四个撒哈拉以南非洲国家的多语言刻板印象资源数据集,旨在填补NLP资源的空白。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22267 2026-02-27 cs.LG 57%

Data-Driven Supervision of a Thermal-Hydraulic Process Towards a Physics-Based Digital Twin

数据驱动的热力过程监督以实现基于物理的数字孪生

Osimone Imhogiemhe, Yoann Jus, Hubert Lejeune, Saïd Moussaoui

机构 * Nantes Univ., Centrale Nantes LS2N, CNRS UMR 6004 F-44000 Nantes, France Fluid \& Sealing Technologies CETIM Nantes, France

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 本文提出了一种基于物理的数字孪生方法,通过数值模拟和机器学习实现热力过程的故障检测与诊断,验证了其在参数变化检测中的有效性。

Journal ref International Conference on Control, Automation and Diagnosis, Jul 2025, Barcelona (ES), Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24072 2026-02-26 cs.CV cs.AI 57%

Uncovering Grounding IDs: How External Cues Shape Multimodal Binding

揭示地面ID:外部线索如何塑造多模态绑定

Hosein Hasani, Amirmohammad Izadi, Fatemeh Askari, Mobin Bagherian, Sadegh Mohammadian, Mohammad Izadi, Mahdieh Soleymani Baghshah

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出地面ID概念,揭示外部线索通过增强多模态绑定的注意力机制,提升跨模态定位精度并减少幻觉。

Comments Under review as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.13126 2026-02-26 cs.CY 57%

Generative agents in the streets: Exploring the use of Large Language Models (LLMs) in collecting urban perceptions

街道中的生成代理:探索大型语言模型(LLMs)在收集城市感知中的应用

Deepank Verma, Olaf Mumm, Vanessa Miriam Carlow

专题命中 其他安全 :safety(abstract);分类 cs.CY

AI总结 本研究利用生成代理探索大型语言模型在模拟城市环境中人类行为中的应用,通过街景图像交互和感知评估提升AI在城市感知中的能力。

Comments 30 Pages, 15 Figures, Submitted in a Journal for Peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20934 2026-02-25 cs.AI 57%

Architecting AgentOS: From Token-Level Context to Emergent System-Level Intelligence

构建AgentOS:从令牌级上下文到涌现的系统级智能

ChengYou Li, XiaoDong Liu, XiangBao Meng, XinYu Zhao

机构 * Yishu Research(亿舒研究) Fukuoka Institute of Technology(福冈技术学院) National University of Singapore(新加坡国立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出AgentOS框架,通过重构LLM为推理内核,引入深度上下文管理和语义切片机制,旨在构建具备系统级智能的自主认知环境。

Comments 16 pages,9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02276 2026-02-25 cs.AI 57%

BioX-Bridge: Model Bridging for Unsupervised Cross-Modal Knowledge Transfer across Biosignals

BioX-Bridge:生物信号跨模态知识迁移的模型桥接

Chenqi Li, Yu Liu, Timothy Denison, Tingting Zhu

机构 * Department of Engineering Science University of Oxford(工程科学系 奥克大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 BioX-Bridge通过轻量级桥接网络实现生物信号跨模态无监督知识迁移,显著减少可训练参数并提升迁移性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20486 2026-02-25 cs.HC cs.AI 57%

Hybrid LLM-Embedded Dialogue Agents for Learner Reflection: Designing Responsive and Theory-Driven Interactions

混合LLM嵌入对话代理用于学习者反思:设计响应性且理论驱动的互动

Paras Sharma, YuePing Sha, Janet Shufor Bih Epse Fofang, Brayden Yan, Jess A. Turner, Nicole Balay, Hubert O. Asare, Angela E. B. Stewart, Erin Walker

机构 * University of Pittsburgh(匹兹堡大学) Carnegie Mellon University(卡内基梅隆大学) Bowie State University(鲍威州立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出一种混合对话系统,结合LLM的响应能力与理论驱动的规则框架,以支持学习者反思,但发现其在提高反思深度的同时也带来了一些挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14370 2026-02-24 cs.AI physics.app-ph physics.soc-ph 57%

Competition for attention predicts good-to-bad tipping in AI

注意力竞争预测AI中的好到坏的 tipping

Neil F. Johnson, Frank Y. Huo

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 本文提出通过注意力竞争预测AI中的 tipping 点,揭示新的控制方法,适用于多领域和法律环境。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19620 2026-02-24 cs.AI 57%

Rules or Weights? Comparing User Understanding of Explainable AI Techniques with the Cognitive XAI-Adaptive Model

规则还是权重?比较用户对可解释AI技术的理解与认知XAI自适应模型

Louth Bin Rawshan, Zhuoyu Wang, Brian Y Lim

机构 * National University of Singapore(新加坡国立大学) Department of Computer Science(计算机科学系)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出CoXAM模型,通过认知框架比较规则与权重两种XAI技术的可解释性,并验证其在决策任务中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19441 2026-02-24 cs.SE cs.AI 57%

When AI Teammates Meet Code Review: Collaboration Signals Shaping the Integration of Agent-Authored Pull Requests

当AI队友遇见代码审查:协作信号塑造代理编写的拉取请求整合

Costain Nachuma, Minhaz Zibran

机构 * Idaho State University USA(爱达荷州立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 研究探讨了AI代理提交的拉取请求在代码审查中的整合效果,发现审查员参与和协作信号对整合成功至关重要。

Comments 5 pages, 2 figures, 1 table. Accepted at the 23rd International Conference on Mining Software Repositories (MSR 2026), Rio de Janeiro, Brazil

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07078 2026-02-24 cs.LG cs.SD eess.AS 57%

E-BATS: Efficient Backpropagation-Free Test-Time Adaptation for Speech Foundation Models

E-BATS:高效的无反向传播测试时间适应方法用于语音基础模型

Jiaheng Dong, Hong Jia, Soumyajit Chatterjee, Abhirup Ghosh, James Bailey, Ting Dang

机构 * The University of Melbourne(墨尔本大学) University of Auckland(奥克兰大学) Nokia Bell Labs, UK(诺基亚贝尔实验室(英国)) University of Birmingham(伯明翰大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 E-BATS是一种专为语音基础模型设计的高效无反向传播测试时间适应框架,通过轻量级提示适应、多尺度损失和测试时间指数移动平均机制,在适应效果和内存效率之间取得平衡,实现了4.1%-13.5%的准确率提升和2.0-6.4倍的GPU内存节省。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10996 2026-02-24 cs.RO cs.LG cs.MA 57%

Mixed-Reality Digital Twins: Leveraging the Physical and Virtual Worlds for Hybrid Sim2Real Transition of Multi-Agent Reinforcement Learning Policies

混合现实数字孪生:利用物理与虚拟世界实现多智能体强化学习策略的混合仿真到现实过渡

Chinmay Vilas Samak, Tanmay Vilas Samak, Venkat Narayan Krovi

机构 * Department of Automotive Engineering, Clemson University International Center for Automotive Research (CU-ICAR)(汽车工程系,克莱姆森大学国际汽车研究中心(CU-ICAR))

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 本文提出混合现实数字孪生框架,通过并行化和域随机化技术,显著提升多智能体强化学习策略的训练效率和仿真到现实迁移性能。

Comments Accepted in IEEE Robotics and Automation Letters (RA-L) and additionally accepted to be presented at IEEE International Conference on Robotics and Automation (ICRA) 2026

Journal ref IEEE Robotics and Automation Letters, vol. 10, no. 9, pp. 9040-9047, Sept. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏