arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-28 至 2026-01-28 共收录 15 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 15 篇

2506.08512 2026-01-28 cs.CV cs.AI 79%

MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for Multi-Modal Video Temporal Grounding

MLVTG: 基于Mamba的特征对齐与LLM驱动的多模态视频时间定位

Zhiyi Zhu, Xiaoyu Wu, Zihao Liu, Linlin Yang

机构 * State Key Laboratory of Media Convergence and Communication(媒体融合与传播国家重点实验室)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

AI总结 MLVTG通过结合MambaAligner和LLMRefiner,实现了多模态视频时间定位的高精度定位与语义净化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19247 2026-01-28 cs.CV 78%

TIGaussian: Disentangle Gaussians for Spatial-Awared Text-Image-3D Alignment

TIGaussian: 解耦高斯分布以实现空间感知的图文-3D对齐

Jiarun Liu, Qifeng Chen, Yiru Zhao, Minghua Liu, Baorui Ma, Sheng Yang

机构 * Unmanned Vehicle Dept., Cainiao Inc., Alibaba Group, Hangzhou, China(阿里巴巴集团 Cainiao 无人机部门,杭州,中国) Hillbot, Sunnyvale, USA(Hillbot,美国 Sunnyvale) Beijing Academy of Artificial Intelligence, Beijing, China(北京人工智能研究院,北京,中国)

专题命中 其他安全 :alignment(title,abstract)

AI总结 TIGaussian通过多分支3DGS分词器和模态特定对齐策略,实现空间感知的图文-3D跨模态对齐,提升3D相关任务的预训练效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15600 2026-01-28 cs.AI cs.CL 62%

Unleashing Scientific Reasoning for Bio-experimental Protocol Generation via Structured Component-based Reward Mechanism

通过结构化组件奖励机制解锁生物实验协议生成中的科学推理

Haoran Sun, Yankai Jiang, Zhenyu Tang, Yaning Pan, Shuang Gu, Zekai Lin, Lilong Wang, Wenjie Lou, Lei Liu, Lei Bai, Xiaosong Wang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 通过结构化组件奖励机制,Thoth在多个基准测试中超越了现有LLMs,实现了更准确的生物实验协议生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15580 2026-01-28 cs.LG cs.CL 62%

Language Models are Symbolic Learners in Arithmetic

语言模型在算术中是符号学习者

Chunyuan Deng, Zhiqi Li, Roy Xie, Ruidi Chang, Hanjie Chen

机构 * Department of Computer Science(计算机科学系) Rice University(里士满大学) College of Computing(计算学院) Georgia Institute of Technology(佐治亚理工学院) Duke University(杜克大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

AI总结 本文研究语言模型在算术运算中通过学习符号捷径而非算法来掌握算术能力。

Comments TMLR 2026. Code at https://github.com/chili-lab/Symbolic-Arithmetic

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19825 2026-01-28 cs.AI cs.DB 57%

Routing End User Queries to Enterprise Databases

将用户查询路由到企业数据库

Saikrishna Sudarshan, Tanay Kulkarni, Manasi Patwardhan, Lovekesh Vig, Ashwin Srinivasan, Tanmay Tulsidas Verlekar

机构 * TCS Research(TCS研究机构) Department of CSIS, BITS Pilani KK Birla Goa Campus(计算机科学与信息系统系,比斯科恩理工学院KK Birla Goa校区)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出了一种基于推理的重新排序策略,通过建模模式覆盖、结构连接性和语义对齐,提升多数据库环境中自然语言查询路由的准确性。

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19170 2026-01-28 cs.AI 57%

Multi-Agent Procedural Graph Extraction with Structural and Logical Refinement

多代理过程图提取与结构逻辑细化

Wangyang Ying, Yanchi Liu, Xujiang Zhao, Wei Cheng, Zhengzhang Chen, Wenchao Yu, Yanjie Fu, Haifeng Chen

机构 * Arizona State University(亚利桑那州立大学) NEC Labs America(NEC美国实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出多代理框架\model{},通过结构和逻辑细化提升过程图提取的准确性和可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06201 2026-01-28 cs.LG 57%

K2-V2: A 360-Open, Reasoning-Enhanced LLM

K2-V2:一种360开放、推理增强的LLM

K2 Team, Zhengzhong Liu, Liping Tang, Linghao Jin, Haonan Li, Nikhil Ranjan, Desai Fan, Shaurya Rohatgi, Richard Fan, Omkar Pangarkar, Huijuan Wang, Zhoujun Cheng, Suqi Sun, Seungwook Han, Bowen Tan, Gurpreet Gosal, Xudong Han, Varad Pimpalkhute, Shibo Hao, Ming Shan Hee, Joel Hestness, Haolong Jia, Liqun Ma, Aaryamonvikram Singh, Daria Soboleva, Natalia Vassilieva, Renxi Wang, Yingquan Wu, Yuekai Sun, Taylor Killian, Alexander Moreno, John Maggs, Hector Ren, Guowei He, Hongyi Wang, Xuezhe Ma, Yuqi Wang, Mikhail Yurochkin, Eric P. Xing

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 K2-V2是一种通过注入领域知识、推理、长上下文和工具使用能力,提升复杂推理任务性能的360开放LLM,具备强大的推理能力和开源训练资源。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24940 2026-01-28 cs.CL 57%

SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens

SemCoT:通过语义对齐的隐式令牌加速链式推理

Yinhan He, Wendy Zheng, Yaochen Zhu, Zaiyi Zheng, Lin Su, Sriram Vasudevan, Qi Guo, Liangjie Hong, Jundong Li

机构 * University of Virginia(弗吉尼亚大学) LinkedIn Inc.(领英公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 SemCoT通过语义对齐的隐式令牌优化,提升链式推理的效率与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06297 2026-01-28 cs.LG 57%

LoaQ: Layer-wise Output Approximation Quantization

LoaQ: 层级输出近似量化

Li Lin, Xiaojun Wan

机构 * Wangxuan Institute of Computer Technology, Peking University(计算机技术研究所,北京大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 LoaQ通过引入输出匹配因素,改进了基于层级的后训练量化方法,提升了模型量化质量并具有简洁的闭式解。

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18939 2026-01-28 cs.LG 57%

A Few Bad Neurons: Isolating and Surgically Correcting Sycophancy

几个坏神经元:隔离并外科手术性纠正阿谀行为

Claire O'Brien, Jessica Seto, Dristi Roy, Aditya Dwivedi, Sunishchal Dev, Kevin Zhu, Sean O'Brien, Ashwinee Panda, Ryan Lagasse

机构 * Algoverse RAND Meta FAIR University of Maryland(马里兰大学) Lockheed Martin AI Center(洛克希德·马丁人工智能中心)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本文提出通过隔离并微调最负责特定行为的神经元来实现LLM行为对齐,展示了在减少阿谀行为任务上的有效性。

Comments Accepted to NeurIPS Workshop on CogInterp and NeurIPS Workshop on Reliable ML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19387 2026-01-28 cs.SE cs.HC 50%

Bridging the Socio-Emotional Gap: The Functional Dimension of Human-AI Collaboration for Software Engineering

弥合社会情感鸿沟:面向软件工程的人机协作的功能维度

Lekshmi Murali Rani, Richard Berntsson Svensson, Robert Feldt

专题命中 其他安全 :alignment(abstract)

AI总结 本文研究了人机协作中社会情感鸿沟问题,提出通过功能等效设计实现有效协作,而非复制人类情感智能特征。

Comments This is the authors accepted manuscript. The final version appears in ACM CHASE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19119 2026-01-28 cs.RO 50%

Agree to Disagree: Consensus-Free Flocking under Constraints

同意却分歧:在约束下无需共识的编队

Peter Travis Jardine, Sidney Givigi

机构 * School of Computing, Queen’s University(女王大学计算机学院)

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出了一种无需共识的编队方法,通过局部观察实现参数协商,以应对智能体间冲突目标和约束条件下的协调运动。

Comments 7 pages. This work has been accepted for publication in the Proceedings of IEEE SYSCON 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05895 2026-01-28 cs.CV 50%

BTCChat: Advancing Remote Sensing Bi-temporal Change Captioning with Multimodal Large Language Model

BTCChat: 通过多模态大语言模型推进遥感双时相变化描述

Yujie Li, Wenjia Xu, Yuanben Zhang, Zhiwei Wei, Mugen Peng

机构 * State Key Laboratory of Networking and Switching Technology(网络与交换技术国家重点实验室) Beijing University of Posts and Telecommunications(北京邮电大学) Aerospace Information Research Institute(航天信息研究所) Chinese Academy of Sciences(中国科学院) School of Geographical Sciences(地理科学学院) Hunan Normal University(湖南师范大学)

专题命中 其他安全 :alignment(abstract)

AI总结 BTCChat通过多模态大语言模型提升遥感双时相变化描述能力,引入变化提取模块和提示增强机制,实现更精确的视觉-语义对齐和更优的性能表现。

Comments 5 pages, 2 figures; Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20497 2026-01-28 physics.soc-ph physics.data-an 50%

Scaling Pedestrian Crossing Analysis to 100 U.S. Cities via AI-based Segmentation of Satellite Imagery

通过基于AI的卫星图像分割将行人过街分析扩展到美国100个城市

Marcel Moran, Arunav Gupta, Jiali Qian, Debra Laefer

专题命中 其他安全 :safety(abstract)

AI总结 通过AI分割卫星图像,研究者准确测量了美国100大城市中行人过街距离,发现其与城市成立年份正相关,揭示了城市设计趋势。

Comments 12 figures, 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18966 2026-01-28 cs.HC 50%

People Can Accurately Predict Behavior of Complex Algorithms That Are Available, Compact, and Aligned

人们可以准确预测可获得、紧凑且对齐的复杂算法的行为

Lindsay Popowski, Helena Vasconcelos, Ignacio Javier Fernandez, Chijioke Chinaza Mgbahurike, Ralf Herbrich, Jeffrey Hancock, Michael S. Bernstein

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出用户能准确预测复杂算法行为的三个条件:认知可用性、概念紧凑性和高对齐度,并通过实验验证了这些条件对预测准确性和心理模型的影响。

Comments 41 pages, 9 figures; this work to appear in PACMHCI V10, N2, April 2026 and be presented at the 29th ACM SIGCHI Conference on Computer-Supported Cooperative Work & Social Computing (CSCW)

详情

展开后加载摘要…

URL PDF HTML 收藏