arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8034 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8034 篇

2205.08184 2022-05-18 cs.CL cs.AI cs.LG 67%

SKILL: Structured Knowledge Infusion for Large Language Models

Fedor Moiseev, Zhe Dong, Enrique Alfonseca, Martin Jaggi

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.10761 2021-12-28 cs.CL cs.AI cs.LG 67%

Bilingual Lexicon Induction through Unsupervised Machine Translation

Mikel Artetxe, Gorka Labaka, Eneko Agirre

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ACL 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.07401 2021-09-16 cs.CL cs.AI cs.IR cs.LG 67%

Matching with Transformers in MELT

Sven Hertling, Jan Portisch, Heiko Paulheim

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments accepted at the Ontology Matching Workshop at the International Semantic Web Conference (ISWC 2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.06324 2021-09-15 cs.CL cs.AI cs.LG 67%

A Massively Multilingual Analysis of Cross-linguality in Shared Embedding Space

Alex Jones, William Yang Wang, Kyle Mahowald

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 15 pages, 8 figures, EMNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.04319 2021-09-10 cs.CL cs.AI cs.LG 67%

Translate & Fill: Improving Zero-Shot Multilingual Semantic Parsing with Synthetic Data

Massimo Nicosia, Zhongdi Qu, Yasemin Altun

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to EMNLP 2021 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.02164 2020-03-04 cs.CL cs.AI cs.LG 67%

Plug and Play Language Models: A Simple Approach to Controlled Text Generation

Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, Rosanne Liu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICLR 2020 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.04914 2019-05-14 cs.AI cs.CL cs.LG 67%

Learning to Exploit Long-term Relational Dependencies in Knowledge Graphs

Lingbing Guo, Zequn Sun, Wei Hu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by the 36th International Conference on Machine Learning (ICML 2019)

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.07178 2019-02-20 eess.AS cs.AI cs.CL cs.LG cs.SD 67%

A spelling correction model for end-to-end speech recognition

Jinxi Guo, Tara N. Sainath, Ron J. Weiss

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to ICASSP 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.02306 2018-09-10 cs.CL cs.AI cs.LG 67%

Unsupervised Cross-lingual Word Embedding by Multilingual Neural Language Models

Takashi Wada, Tomoharu Iwata

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.03938 2017-11-15 cs.CL cs.AI cs.LG 67%

Representation Learning for Grounded Spatial Reasoning

Michael Janner, Karthik Narasimhan, Regina Barzilay

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to TACL 2017, code: https://github.com/jannerm/spatial-reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
1508.04395 2016-03-16 cs.CL cs.AI cs.LG cs.NE 67%

End-to-End Attention-based Large Vocabulary Speech Recognition

Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, Yoshua Bengio

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09984 2025-04-22 cs.LG cs.AI 66%

Symmetry-Breaking Augmentations for Ad Hoc Teamwork

Ravi Hammond, Dustin Craggs, Mingyu Guo, Jakob Foerster, Ian Reid

机构 * Foerster Lab for AI Research, University of Oxford(牛津大学人工智能研究实验室) Australian Institute for Machine Learning, University of Adelaide(阿德莱德大学人工智能研究所) Meta AI Research, UK(英国Meta人工智能研究) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他安全 :alignment(abstract,comments);分类 cs.AI、cs.LG

Comments 21 pages, 12 figures, Bidirectional Human-AI Alignment workshop, ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.08093 2021-12-16 cs.LG cs.AI 66%

Towards Controllable Agent in MOBA Games with Generative Modeling

Shubao Zhang

专题命中 其他安全 :alignment(abstract,comments);分类 cs.AI、cs.LG

Comments Human-Compatible AI; Human-AI Cooperation; AI control; AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.10318 2021-10-22 cs.CL cs.LG 66%

Improved Multilingual Language Model Pretraining for Social Media Text via Translation Pair Prediction

Shubhanshu Mishra, Aria Haghighi

专题命中 其他安全 :alignment(abstract,comments);分类 cs.CL、cs.LG

Comments Camera ready version. Accepted to WNUT 2021. Code for reproducing the experiments can be found at: https://github.com/twitter-research/multilingual-alignment-tpp

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07237 2026-08-26 cs.LG cs.AI 62%

GWT: Scalable Optimizer State Compression for Large Language Model Training

GWT: 大规模语言模型训练中的可扩展优化器状态压缩

Ziqing Wen, Ping Luo, Jiahuan Wang, Junlin Zeng, Kun Yuan, Dongsheng Li, Tao Sun

机构 * National Key Laboratory of Parallel and Distributed Computing(并行与分布式计算国家重点实验室) National University of Defense Technology(国防科技大学) Center for Machine Learning Research(机器学习研究中心)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出GWT,一种基于小波变换的优化器状态压缩方法,有效减少内存开销,同时保持模型精度,适用于大规模预训练和任务微调。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22908 2026-08-25 cs.CL cs.AI 新提交 62%

Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text

口语语言模型在阅读文本时是否“听到”语音?弥合语音与文本之间的结构差距

Hyeonyu Kim, Hwayeon Kim, Youngwon Choi, Myeongkyun Cho, Huu-Kim Nguyen

机构 * Maum AI Inc.(Maum AI公司) KAIST(韩国科学技术院) Atmanity Inc.(Atmanity公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 该研究针对现有口语语言模型(SLMs)未充分解决语音与文本结构差异的问题,提出解耦长度不匹配与语义对齐的框架,经多基准实验验证其性能可与强基线媲美,凸显了明确解决语音文本结构差异的重要性。

Comments Accepted to EMNLP 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22417 2026-08-25 cs.AI cs.CL cs.HC 新提交 62%

LLMs for Survey Text Analysis - A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis

用于调查文本分析的大语言模型(LLMs):人类与GPT-5在归纳内容分析上的性能比较

Leonardo Bergmann, Renata Gheorghiu, Ana Gvritishvili, Alex Mican, Chris Stewart, Topias Tolonen-Weckström

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本研究对比人类与GPT-5.4在903份欧洲博士生调查开放式回答的归纳内容分析中的性能,发现GPT-5.4编码的ARI为0.61,可近似人类编码表现,或可作为归纳质性分析的可扩展支持工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22076 2026-08-25 cs.LG cs.AI 新提交 62%

Improving Energy Efficiency of Oil Platforms Through Optimal Loading of Diesel Generators Using Machine Learning and Search Algorithms

基于机器学习与搜索算法优化柴油发电机负载,提升石油平台能效

Khivishta Boodhoo, Josh Plumbly, Nicholas Watson

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本研究针对海上石油平台能源低效问题,采用机器学习构建柴油消耗预测模型,结合搜索算法优化发电机负载,实现日均27%的柴油消耗降低,为平台能效提升提供可行方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12196 2026-08-25 cs.LG cs.AI 62%

Dynamic Relational Priming Improves Transformer in Multivariate Time Series

动态关系先验提升Transformer在多变量时间序列中的表现

Hunjae Lee, Corey Clark

机构 * Department of Computer Science, Southern Methodist University, Dallas TX USA(计算机科学系,南方 Methodist 大学,德克萨斯州达拉斯)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 提出动态关系先验注意力机制(prime attention),通过为每个token对动态调整表示,有效捕捉多变量时间序列中异构的通道间依赖关系,在保持相同计算复杂度下提升预测精度达6.5%。

Journal ref Proceedings of the 43rd International Conference on Machine Learning (ICML). 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24255 2026-08-25 cs.CL cs.AI cs.HC 版本更新 62%

Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games

心理理论与亲社会信念对最后通牒博弈中大型语言模型的人类对齐行为的影响

Neemesh Yadav, Yihuai Lan, Shan Dong, Mai Hieu Hien, Palakorn Achananuparp, Jing Jiang, Ee-Peng Lim

机构 * Singapore Management University(新加坡管理大学) Australian National University(澳大利亚国立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本研究以最后通牒博弈为任务,赋予LLM不同亲社会信念与推理方法,开展2700次模拟,发现ToM推理可增强LLM行为对齐度等,且不同游戏角色适配不同ToM顺序,Llama 3.3 70B的推理与行动信念最一致。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20281 2026-08-21 cs.CL cs.AI 新提交 62%

Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

注入、对齐、恢复:用于无需检索的文档知识内化的分阶段后训练

Qian Kou, Xiaofeng Shi, Xiaosong Qiu, Hua Zhou

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 针对大型语言模型无需检索文档时无法回答相关问题的痛点,提出IAR三阶段后训练框架,在多数据集、多模型上显著提升了文档知识内化的领域与通用性能。

Comments 21 pages, 4 figures. Includes Supplementary Material Sections A--G. Qian Kou and Xiaofeng Shi contributed equally and are co-corresponding authors. Hua Zhou is the project leader

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19875 2026-08-21 cs.CL cs.AI 新提交 62%

A knowledge-guided agentic framework for mitigating patient-context ambiguity in health queries

一种用于缓解医疗查询中患者上下文歧义的知识引导智能体框架

Mahyar Abbasian, Saba A. Farahani, Arshia Ilaty, Hung Cao, Ramesh Jain, Amir M. Rahmani

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

AI总结 该研究提出知识引导智能体框架,通过解析医疗查询、构建假设、识别缺失上下文并提问,缓解患者上下文歧义,在诊断检索和饮食安全分类任务中显著提升了多种语言模型的性能。

Comments 48 pages, 3 figures, 6 tables, journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24856 2026-08-21 cs.LG cs.AI 版本更新 62%

The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth

概念分配区:追踪概念如何跨越Transformer深度形成

James Henry

机构 * Independent Researcher(独立研究者)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 提出概念分配区(CAZ)框架,通过层间度量(分离度、概念一致性、概念速度)检测概念在残差流中逐渐形成的深度区间,并在34个模型上验证了分离曲线的多模态性及温和CAZ的因果活性。

Comments v2: substantial revision. Cross-architecture ordering statistic and MHA/GQA cohort-split claim retracted per recomputation; corpus and companion-paper citations refreshed. See paper's "Changes from Version 1" section for the full list of superseded values

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03519 2026-08-21 cs.CL cs.AI 版本更新 62%

TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning

TS-Reasoner:将时间序列基础模型与大语言模型推理能力对齐

Fangxu Yu, Hongyu Zhao, Tianyi Zhou

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 TS-Reasoner通过两阶段训练将时间序列基础模型与大语言模型对齐,在多个基准上性能优于同类模型且数据效率更高,解决了时间序列模型推理能力不足的问题。

Comments Accepted to Transactions on Machine Learning Research, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12627 2026-08-20 cs.CV cs.AI cs.CL cs.HC 版本更新 62%

EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory

EgoCITE:面向长时程自我中心记忆的上下文增强索引与时序感知检索

Le Zhang, Hao Chen, Vlad Roznyatovskiy, Jianzhong Zhang, Ke Sun

机构 * University of Michigan(密歇根大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本研究针对长时程自我中心记忆系统的索引不可靠、忽略时序意图的问题,提出EgoCITE框架,经多数据集评估,其准确率优于基线且成本显著低于长上下文LLM智能体。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17117 2026-08-20 cs.CL cs.AI cs.IT math.IT 62%

From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning

从标记到思考:LLMs和人类如何在压缩与意义之间进行权衡

Chen Shani, Liron Soffer, Dan Jurafsky, Yann LeCun, Ravid Shwartz-Ziv

机构 * Stanford University(斯坦福大学) Tel Aviv University(特拉维夫大学) New York University(纽约大学) Meta - FAIR

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文通过信息瓶颈框架比较人类与LLMs的概念结构,发现LLMs在压缩效率上优于人类,但牺牲了语义丰富性,揭示了人工与自然智能的本质差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14763 2026-08-18 eess.IV cs.AI cs.CV cs.LG 新提交 62%

Cross-Modal Ultrasound-MRI Learning for Fetal Brain Ventricular Volumetry and Abnormality Screening

用于胎儿脑室体积测量与异常筛查的跨模态超声-MRI学习

Yuhao Huang, Yuanji Zhang, Yuhuan Lu, Dong Ni, P. Ellen Grant, Davood Karimi

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本研究提出跨模态学习框架VIFBA,基于超声视频实现胎儿脑室体积预测、VM严重程度分类及非VM异常筛查,性能优于基线与现有模型,为产前脑筛查提供实用方案。

Comments 17 pages, 11 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16514 2026-08-18 cs.CV cs.AI cs.CL cs.HC cs.MM 新提交 62%

Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans

匹配结果,不同注视:中央凹多模态大语言模型(MLLM)的搜索方式与人类的对比

Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Ulas Bagci, Alessandro Bruno

机构 * F-initiatives(F计划) Université Sorbonne Paris Nord(巴黎北索邦大学) Northwestern University(西北大学) IULM university(IULM大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 研究对比三款通用MLLM与人类在目标导向视觉搜索中的表现,发现模型在决策和目标获取上优于人类,但注视过程与人类不同,现有指标无法验证类人视觉,零样本模型不适用于过程层面问题。

Comments Paper accepted at 3rd HCV workshop at ECCV 2026. 12 pages main text, 16 pages supp

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16347 2026-08-18 cs.CL cs.LG 新提交 62%

Architecture-Dependent Causal Transfer of Activation States Across Large Language Models

大语言模型间激活状态的架构依赖型因果迁移

Fernando Cardenas Piepereit

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

AI总结 本文探究大语言模型间激活状态的因果迁移,发现该迁移依赖模型架构,仅部分仅解码器模型对可实现具统计显著性的因果效应,且迁移的是表征载体而非意义。

Comments 13 pages, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15693 2026-08-18 cs.AI cs.LG 新提交 62%

Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment

面向小型设备的大模型:边缘AI部署的最新进展与实证分析

Subhransu Das, Jiaming Cheng, Arnav Kumar, Sadia Afrose, Mingzhe Han, Michael Silagy, Shreya Palande, Brijesh Soni, Rajiv Ramnath

机构 * The Ohio State University(俄亥俄州立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文调研近期边缘AI部署研究,提炼指南并在多平台测试,发现不同任务适用不同压缩技术,剪枝可能提升分割性能但会增加延迟,相关成果已开源。

Comments Parts of this work were presented at the IEEE Consumer Communications & Networking Conference (CCNC), Las Vegas, NV, USA, January 2026

详情

展开后加载摘要…

URL PDF HTML 收藏