arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-02-03 至 2026-02-03 共收录 148 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 39 篇

2602.00092 2026-02-03 cs.LG cs.AI cs.CL cs.CV 67%

Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits

通过宪法进行原子概念编辑来解释和控制模型行为

Neha Kalibhat, Zi Wang, Prasoon Bajpai, Drew Proud, Wenjun Zeng, Been Kim, Mani Malek

机构 * Google DeepMind(谷歌DeepMind)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 该研究提出通过原子概念编辑学习模型宪法,以解释和控制模型行为,实验证明其在提升模型成功率方面效果显著。

Journal ref Twenty-Ninth Annual Conference on Artificial Intelligence and Statistics (AISTATS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01610 2026-02-03 cs.AI cs.LG 62%

ToPT: Task-Oriented Prompt Tuning for Urban Region Representation Learning

为城市区域表示学习设计的任务导向提示微调:ToPT

Zitao Guo, Changyang Jiang, Tianhong Zhao, Jinzhou Cao, Genan Dai, Bowen Zhang

机构 * College of Applied Science, Shenzhen University, Shenzhen, China(深圳大学应用科学学院) School of Artificial Intelligence, Shenzhen Technology University, Shenzhen, China(深圳科技大学人工智能学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 ToPT通过空间一致融合和任务对齐提升城市区域表示学习,实现任务导向的提示微调,提升多个城市任务的性能。

Comments The paper has been accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.13268 2026-02-03 q-bio.GN cs.AI cs.LG q-bio.QM 62%

Revolutionizing Genomics with Reinforcement Learning Techniques

用强化学习技术革新基因组学

Mohsen Karami, Khadijeh, Jahanian, Roohallah Alizadehsani, Iman Dehzangi, Juan M Gorriz, Yudong Zhang, Jia Wang, Farshid Hajati, Min Yang, Thantrira Porntaveetus, Hamid Alinejad-Rokny

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文探讨了强化学习在基因组学中的应用,包括基因调控网络、基因组组装和序列比对,并讨论了未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20245 2026-02-03 cs.CY cs.AI cs.HC 62%

How AI Impacts Skill Formation

人工智能如何影响技能形成

Judy Hanwen Shen, Alex Tamkin

机构 * Anthropic

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY

AI总结 研究发现人工智能辅助虽提升生产力,但损害技能学习,需谨慎使用以保护技能形成。

Comments V2: Fixed typos in Table 1 (Balance Table), corrected osf link

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01734 2026-02-03 cs.LG 57%

MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration

MSign: 通过稳定秩恢复防止大语言模型训练不稳定

Lianhai Ren, Yucheng Ding, Xiao Liu, Qianxiao Li, Peng Cheng, Yeyun Gong

机构 * National University of Singapore(新加坡国立大学) Microsoft Research(微软研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 MSign通过稳定秩恢复机制防止大语言模型训练不稳定,有效减少训练失败并降低计算开销

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01598 2026-02-03 cs.CL 57%

The Art of Socratic Inquiry: A Framework for Proactive Template-Guided Therapeutic Conversation Generation

苏格拉底式探究的艺术:一种主动模板引导治疗对话生成的框架

Mingwen Zhang, Minqiang Yang, Changsheng Ma, Yang Yu, Hui Bai, Chen Xu, Xiangzhen Kong, Bin Hu

机构 * School of Information Science and Engineering, Lanzhou University(信息科学与工程学院,兰州大学) School of Medical Technology, Beijing Institute of Technology(医学技术学院,北京理工大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 本文提出Socratic Inquiry Framework,通过策略锚定和模板检索分离提问时机与内容,提升LLMs的主动提问能力,推动心理引导型大语言模型的发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01525 2026-02-03 cs.HC cs.CY 57%

How Notations Evolve: A Historical Analysis with Implications for Supporting User-Defined Abstractions

符号如何演变:一种历史分析及其对支持用户定义抽象的启示

Jingyue Zhang, J. D. Zamfirescu-Pereira, Elena L. Glassman, Damien Masson, Ian Arawjo

专题命中 其他安全 :alignment(abstract);分类 cs.CY

AI总结 本文通过历史分析探讨符号演变的三个社会阶段和三个功能阶段,提出如何设计支持用户定义抽象的新系统。

Comments 23 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05825 2026-02-03 cs.HC cs.AI 57%

Decoding Workload and Agreement From EEG During Spoken Dialogue With Conversational AI

从语音对话中解码工作负荷与共识的EEG信号

Lucija Mihić Zidar, Philipp Wicke, Praneel Bhatia, Rosa Lutz, Marius Klug, Thorsten O. Zander

机构 * Chair of Neuroadaptive Human--Computer Interaction(神经适应性人机交互教授席位) Brandenburg Technical University Cottbus--Senftenberg(勃兰登堡技术大学库滕堡-森芬根堡分校) Auryal GmbH(Auryal公司) Cottbus, Germany(库滕堡,德国)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出通过EEG解码语音对话中的工作负荷与共识,验证了现有分类器在对话场景中的迁移能力,并探讨了其应用限制。

Comments Accepted at the 14th International Winter Conference on Brain-Computer Interface

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01443 2026-02-03 cs.AI 57%

SimGym: Traffic-Grounded Browser Agents for Offline A/B Testing in E-Commerce

SimGym:基于流量的浏览器代理用于电子商务中的离线A/B测试

Alberto Castelo, Zahra Zanjani Foumani, Ailin Fan, Keat Yang Koay, Vibhor Malik, Yuanzheng Zhu, Han Li, Meysam Feghhi, Ronie Uliana, Shuang Xie, Zhaoyu Zhang, Angelo Ocana Martins, Mingyu Zhao, Francis Pelland, Jonathan Faerman, Nikolas LeBlanc, Aaron Glazer, Andrew McNamara, Lingyun Wang, Zhong Wu

机构 * Shopify

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 SimGym通过基于流量的合成买家和大型语言模型代理,实现快速离线A/B测试,提升电子商务实验效率

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01382 2026-02-03 cs.CV cs.LG 57%

PromptRL: Prompt Matters in RL for Flow-Based Image Generation

PromptRL: 流基于图像生成中强化学习中提示的重要性

Fu-Yun Wang, Han Zhang, Michael Gharbi, Hongsheng Li, Taesung Park

机构 * The Chinese University of Hong Kong, Hong Kong(香港中文大学) Meta Superintelligence Labs, USA(Meta超智能实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 PromptRL通过整合语言模型作为可训练的提示精修代理,提升流基于图像生成中强化学习的性能和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01335 2026-02-03 cs.CV cs.AI 57%

Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning

超越像素:通过基于模式的代理推理实现视觉隐喻转移

Yu Xu, Yuxin Zhang, Juan Cao, Lin Gao, Chunyu Wang, Oliver Deussen, Tong-Yee Lee, Fan Tang

机构 * University of Chinese Academy of Sciences(中国科学院大学) Tencent Hunyuan(腾讯文脉) University of Konstanz(康斯坦茨大学) National Cheng-Kung University(国立成功大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出基于模式的代理推理框架,实现视觉隐喻转移,通过多代理协作提升隐喻生成的抽象逻辑与视觉创造力。

Comments 11 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01161 2026-02-03 cs.CL 57%

Beyond Training for Cultural Awareness: The Role of Dataset Linguistic Structure in Large Language Models

超越文化意识的训练:大型语言模型中数据集语言结构的作用

Reem I. Masoud, Chen Feng, Shunta Asano, Saied Alshahrani, Philip Colin Treleaven, Miguel R. D. Rodrigues

机构 * University College London(伦敦大学学院) Queen’s University Belfast(贝尔法斯特女王大学) The University of Tokyo(东京大学) University of Bisha(比沙大学) King Abdulaziz University(阿卜杜勒阿齐兹国王大学) AI Centre, University College London(伦敦大学学院人工智能中心)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 本文研究了数据集语言结构对大型语言模型文化表现的影响,发现词汇导向的成分在不同模型中表现更稳定,而语义或多样性极端成分可能产生中性或有害效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00574 2026-02-03 cs.AI 57%

Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddings

通过潜在嵌入学习模态混合的思考链推理

Yifei Shao, Kun Zhou, Ziming Xu, Mohammad Atif Quamar, Shibo Hao, Zhen Wang, Zhiting Hu, Biwei Huang

机构 * University of California, San Diego(加州大学圣地亚哥分校)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出模态混合CoT方法,通过潜在嵌入结合视觉与文本信息,提升多模态推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00446 2026-02-03 cs.LG cs.CR 57%

Towards Building Non-Fine-Tunable Foundation Models

构建不可微调的基础模型

Ziyao Wang, Nizhang Li, Pingzhi Li, Guoheng Sun, Tianlong Chen, Ang Li

机构 * University of Maryland, College Park(马里兰大学 College Park 分校) Macau University of Science and Technology(澳门科学技术大学) The University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 本文提出PMP框架,通过构建不可微调的基础模型,限制未经授权微调的适应性提升,从而提高模型的安全性和稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00361 2026-02-03 cs.LG quant-ph 57%

Quantum Generator Kernels

量子生成核

Philipp Altmann, Maximilian Mansky, Maximilian Zorn, Jonas Stein, Claudia Linnhoff-Popien

机构 * LMU Munich(慕尼黑莱布尼茨大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本文提出量子生成核方法,通过可变生成组实现大规模数据在量子空间中的高效嵌入,提升量子机器学习的分类与投影能力。

Comments 28 pages, 4 figures, 8 tables, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00127 2026-02-03 cs.LG 57%

ALIGN: Aligned Delegation with Performance Guarantees for Multi-Agent LLM Reasoning

ALIGN: 多智能体LLM推理中的对齐委托与性能保证

Tong Zhu, Baiting Chen, Jin Zhou, Hua Zhou, Sriram Sankararaman, Xiaowu Dai

机构 * Department of Biostatistics, UCLA(生物统计学系,加州大学洛杉矶分校) Department of Statistics and Data Science, UCLA(统计学与数据科学系,加州大学洛杉矶分校) Department of Computer Science, UCLA(计算机科学系,加州大学洛杉矶分校) Departments of Statistics and Data Science, and of Biostatistics, UCLA(统计学与数据科学系和生物统计学系,加州大学洛杉矶分校)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 ALIGN通过多智能体对齐委托机制,提升大语言模型在复杂推理任务中的性能保障与表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00110 2026-02-03 cs.CV cs.LG 57%

Observing Health Outcomes Using Remote Sensing Imagery and Geo-Context Guided Visual Transformer

利用遥感图像和地理上下文引导的视觉变换器观察健康结果

Yu Li, Guilherme N. DeSouza, Praveen Rao, Chi-Ren Shyu

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本文提出一种结合地理空间信息的视觉变换器模型,通过引导注意力机制提升遥感图像处理效果,有效预测疾病流行率。

Comments Submitted to IEEE Transactions on Geoscience and Remote Sensing

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09114 2026-02-03 cs.LG 57%

TRACE: Grounding Time Series in Context for Multimodal Embedding and Retrieval

TRACE: 在上下文中对时间序列进行建模以实现多模态嵌入与检索

Jialin Chen, Ziyu Zhao, Gaukhar Nurbek, Aosong Feng, Ali Maatouk, Leandros Tassiulas, Yifeng Gao, Rex Ying

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 TRACE通过在上下文中对时间序列进行建模,实现多模态嵌入与检索,提升下游任务的预测精度和可解释性,同时作为强大的独立编码器优化上下文感知表示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07272 2026-02-03 cs.CL q-bio.GN 57%

GENERator: A Long-Context Generative Genomic Foundation Model

生成器:一种长上下文生成基因组基础模型

Wei Wu, Qiuyi Li, Yuanyuan Zhang, Zhihao Zhan, Ruipu Chen, Mingyang Li, Kun Fu, Junyan Qi, Yongzhou Bao, Chao Wang, Yiheng Zhu, Zhiyun Zhang, Jian Tang, Fuli Feng, Jieping Ye, Yuwen Liu, Hui Xiong, Zheng Wang

机构 * Equal Contribution 12pt Senior Authorship 0.15in 1Alibaba Cloud Computing, Beijing, China 0.05in 2Zhongguancun Academy, Beijing, China 0.05in 3Zhongguancun Institute of Artificial Intelligence, Beijing, China 0.05in 4University of Science Technology of China, Hefei, China 0.05in 5State Key Laboratory of Genome Multi-omics Technologies, Shenzhen Branch, Guangdong Laboratory for Lingnan Modern Agriculture, Key Laboratory of Livestock Poultry Multi-Omics of MARA, Agricultural Genomics Institute at Shenzhen, Chinese Academy of Agricultural Sciences, Shenzhen, China 0.05in 6Innovation Group of Pig Genome Design Breeding, Research Centre for Animal Genome, Agricultural Genomics Institute at Shenzhen, Chinese Academy of Agricultural Sciences, Shenzhen, China 0.05in 7The Hong Kong University of Science Technology (Guangzhou), Guangzhou, China 0.05in 8The Hong Kong University of Science Technology , Hong Kong SAR, China 0.05in 9Mila - Qu\'ebec AI Institute, Montr\'eal, Canada 0.05in 10University of Montr\'eal, Montr\'eal, Canada 0.05in 11HEC Montr\'eal, Montr\'eal, Canada 0.05in 12CIFAR AI Chair, Canada 0.05in 13Carnegie Mellon University, Pittsburgh, USA 0.15in to

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 GENErator是一种长上下文生成基因组基础模型,通过大规模预训练实现高效基因组解读与可编程序列设计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01896 2026-02-03 cond-mat.mtrl-sci 50%

Thermophysical properties of spark plasma sintered UCo: a comparison with machine learning predictions

火花等离子烧结UCo的热物理性质:与机器学习预测的比较

Yifan Sun, Hironobu Nakamura, Masaya Kumagai, Yuji Ohishi, Ken Kurosaki

专题命中 其他安全 :safety(abstract)

AI总结 本研究通过实验验证了机器学习对UCo热导率预测的准确性,并展示了其在筛选高热导率铀化合物中的应用价值。

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10173 2026-02-03 cs.SE 50%

ATLAS: Automated Toolkit for Large-Scale Verified Code Synthesis

ATLAS:大规模验证代码合成的自动化工具包

Mantas Baksys, Stefan Zetzsche, Olivier Bouissou, Remi Delmas, Soonho Kong, Sean B. Holden

专题命中 其他安全 :safety(abstract)

AI总结 ATLAS通过生成验证代码提升LLM在形式验证领域的性能,其在DafnyBench和DafnySynthesis上的表现显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01344 2026-02-03 astro-ph.SR 50%

Machine learning for understanding pulsating stars I: the non-linear phenomenon in δ Scuti stars

利用机器学习理解脉动星I:δ Scuti星的非线性现象

J. R. Rodon, J. Pascual-Granado, M. Lares-Martiz, M. Rodríguez Sánchez, C. Roche

专题命中 其他安全 :alignment(abstract)

AI总结 本研究利用机器学习聚类技术,通过分析频域特征和非线性机制,揭示δ Scuti星的内在子群,挑战基于幅度的传统分类方法。

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00941 2026-02-03 cs.NI 50%

LMTE: Putting the "Reasoning" into WAN Traffic Engineering with Language Models

LMTE:利用语言模型进行WAN流量工程中的推理

Xinyu Yuan, Yan Qiao, Zonghui Wang, Meng Li, Wenzhi Chen

专题命中 其他安全 :alignment(abstract)

AI总结 LMTE利用语言模型进行WAN流量工程,通过高效多模态对齐和轻量级配置生成,实现更高效的流量规划和优化。

Comments Accepted as a conference paper at IEEE INFOCOM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00639 2026-02-03 cs.CV 50%

Diff-PC: Identity-preserving and 3D-aware Controllable Diffusion for Zero-shot Portrait Customization

Diff-PC: 保持身份和3D感知的可控扩散用于零样本肖像定制

Yifang Xu, Benxiang Zhai, Chenyu Zhang, Ming Li, Yang Li, Sidan Du

机构 * School of Electronic Science and Engineering, Nanjing University(南京大学电子科学与工程学院) School of Business, Hohai University(河海大学商学院) School of Artificial Intelligence, Nanjing University of Information Science and Technology(南京信息工程大学人工智能学院)

专题命中 其他安全 :alignment(abstract)

AI总结 Diff-PC通过3D感知和可控扩散方法实现零样本肖像定制,提升身份保持和面部可控性。

Comments Accepted by Information Fusion 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00604 2026-02-03 cs.SD eess.AS 50%

The TMU System for the XACLE Challenge: Training Large Audio Language Models with CLAP Pseudo-Labels

TMU系统用于XACLE挑战:利用CLAP伪标签训练大型音频语言模型

Ayuto Tsutsumi, Kohei Tanaka, Sayaka Shiota

机构 * Tokyo Metropolitan University(东京 Metropolitan 大学)

专题命中 其他安全 :alignment(abstract)

AI总结 TMU系统通过CLAP伪标签预训练,在XACLE挑战中实现0.632的SRCC成绩,优于基线系统并获得第三名。

Comments 3 pages; 2 figures; 2 tables; Accepted at ICASSP 2026 Workshop (SP Grand Challenges, GC-12: XACLE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09058 2026-02-03 cs.RO cs.SY eess.SY 50%

Reach-Avoid-Stabilize Using Admissible Control Sets

利用可接受控制集实现到达-回避-稳定

Zheng Gong, Boyang Li, Sylvia Herbert

机构 * Mechanical and Aerospace Engineering at UC San Diego(机械与航空航天工程系,UC圣地亚哥大学)

专题命中 其他安全 :safety(abstract)

AI总结 本文提出利用可接受控制集的方法,解决可达-回避-稳定问题,确保系统在连续目标和障碍物回避下安全稳定到达感兴趣点。

Comments 7 pages, 5 figures, submitted to 64th IEEE Conference on Decision and Control

Journal ref 2025 IEEE 64th Conference on Decision and Control (CDC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00494 2026-02-03 cs.HC 50%

SRL Proxemics: Spatial Guidelines for Supernumerary Robotic Limbs in Near-Body Interactions

SRL 亲疏距离:近身交互中冗余机械肢体的空间指南

Hongyu Zhou, Chia-An fan, Yihao Dong, Shuto Takashita, Masahiko Inami, Zhanna Sarsenbayeva, Anusha Withana

专题命中 其他安全 :safety(abstract)

AI总结 本文提出SRL Proxemics框架,通过近身交互实验揭示自主性与空间行为校准对安全性和信任的影响。

Comments Accepted to CHI 2026. Version of Record: DOI 10.1145/3772318.3790532

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00123 2026-02-03 cs.HC cs.CV 50%

Visual Affect Analysis: Predicting Emotions of Image Viewers with Vision-Language Models

视觉情感分析:利用视觉-语言模型预测图像观看者的感受

Filip Nowicki, Hubert Marciniak, Jakub Łączkowski, Krzysztof Jassem, Tomasz Górecki, Vimala Balakrishnan, Desmond C. Ong, Maciej Behnke

机构 * Faculty of Mathematics and Computer Science, Adam Mickiewicz University(数学与计算机科学学院,亚当·密茨凯维奇大学) Faculty of Computer Science and Information Technology, Universiti Malaya(计算机科学与信息技术学院,马来亚大学) Department of Computer Science and Engineering, Korea University(计算机科学与工程系,韩国大学) Department of Psychology, University of Texas at Austin(心理学系,德克萨斯大学奥斯汀分校) Cognitive Neuroscience Center, Adam Mickiewicz University(认知神经科学中心,亚当·密茨凯维奇大学)

专题命中 其他安全 :alignment(abstract)

AI总结 本文研究了视觉-语言模型在预测图像观众情感方面的性能,发现其在离散情绪分类上表现良好,但在连续评分预测中存在偏差,揭示了其在情感计算中的潜力与局限。

详情

展开后加载摘要…

URL PDF HTML 收藏