arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-12-30 至 2025-12-30 共收录 19 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 19 篇

2511.21218 2025-12-30 cs.CL 79%

Can Finetuing LLMs on Small Human Samples Increase Heterogeneity, Alignment, and Belief-Action Coherence?

能否通过在小样本人类数据上微调LLM增加异质性、对齐性和信念-行为一致性?

Steven Wang, Kyle Hunt, Shaojie Tang, Kenneth Joseph

机构 * Department of Management Science and Systems, University at Buffalo(管理科学与系统系,布法罗大学) Department of Computer Science and Engineering, University at Buffalo(计算机科学与工程系,布法罗大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

AI总结 本研究探讨通过微调LLM在小样本人类数据上是否能提高模拟结果的异质性、对齐性和信念-行为一致性,发现虽有提升但无法完全替代人类数据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22741 2025-12-30 cs.CL 79%

Text-Routed Sparse Mixture-of-Experts Model with Explanation and Temporal Alignment for Multi-Modal Sentiment Analysis

基于解释和时序对齐的文本引导稀疏专家混合模型用于多模态情感分析

Dongning Rao, Yunbiao Zeng, Zhihua Jiang, Jujian Lv

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

AI总结 本文提出基于解释和时序对齐的文本引导稀疏专家混合模型,用于提升多模态情感分析的性能。

Comments 9 pages, 9 figures, accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00115 2025-12-30 cs.CL cs.AI 76%

Cognitive Alignment in Personality Reasoning: Leveraging Prototype Theory for MBTI Inference

人格推理中的认知对齐:利用原型理论进行MBTI推断

Haoyuan Li, Yuanbo Tong, Yuchen Li, Zirui Wang, Chunhou Liu, Jiamou Liu

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 ProtoMBTI通过原型理论与认知对齐,提升文本人格推断的准确性和泛化能力。

Comments The authors have decided to withdraw this version to substantially revise and extend the work

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15545 2025-12-30 cs.CL cs.AI 76%

TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs

TokenTiming: 一种适用于通用推测解码模型对的动态对齐方法

Sibo Xiao, Jinyuan Fu, Zhongle Xie, Lidan Shou

机构 * Zhejiang University(浙江大学)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 TokenTiming通过动态对齐方法实现通用推测解码,无需重新训练即可处理不同词汇模型,提升LLM推理效率1.57倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23487 2025-12-30 cs.LG cs.AI stat.ML 73%

ML Compass: Navigating Capability, Cost, and Compliance Trade-offs in AI Model Deployment

ML Compass:在AI模型部署中导航能力、成本和合规性之间的权衡

Vassilis Digalakis, Ramayya Krishnan, Gonzalo Martin Fernandez, Agni Orfanoudaki

机构 * Questrom School of Business, Boston University(波士顿大学Questrom商学院) Centre de Formació Interdisciplinària Superior and Universitat Politècnica de Catalunya(巴塞罗那理工大学跨学科教育中心)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 ML Compass通过系统性框架优化模型选择,考虑能力、成本和合规性之间的权衡,提供部署导向的推荐和排行榜。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08270 2025-12-30 cs.LG cs.AI cs.CL cs.MM 67%

Doctor Sun: A Bilingual Multimodal Large Language Model for Biomedical AI

Doctor Sun: 一种双语多模态大语言模型用于生物医学AI

Dong Xue, Ziyao Shao, Zhaoyang Duan, Fangzhou Liu, Bing Li, Zhongheng Zhang

机构 * Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education East China University of Science and Technology(能源化工过程智能制造重点实验室,东华大学) Research Institute of Intelligent Control and Systems Harbin Institute of Technology(智能控制与系统研究室,哈尔滨工业大学) Department of Emergency Medicine, Sir Run Run Shaw Hospital Zhejiang University School of Medicine(浙江大学医学院急诊医学科) Provincial Key Laboratory of Precise Diagnosis Treatment of Abdominal Infection, Sir Run Run Shaw Hospital Zhejiang University School of Medicine(腹部感染精准诊断治疗省级重点实验室,浙江大学医学院) School of Medicine Shaoxing University(绍兴大学医学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 Doctor Sun是一种双语多模态大语言模型,通过整合预训练视觉编码器和医学LLM,提升生物医学多模态任务的性能,并提供SunMed-VL数据集支持研究进展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04208 2025-12-30 cs.LG cs.AI 62%

Aligning Agents like Large Language Models

像大型语言模型一样对齐代理

Adam Jelley, Yuhan Cao, Dave Bignell, Amos Storkey, Sam Devlin, Tabish Rashid

机构 * School of Informatics, University of Edinburgh, Edinburgh, United Kingdom(爱丁堡大学信息学院) Micosoft Research, Cambridge, United Kingdom(微软研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过像训练大型语言模型一样训练代理,以实现更通用、稳健和对齐的行为,并通过3D视频游戏环境中的概念验证展示方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23422 2025-12-30 cs.CL 57%

Entropy-Guided Token Dropout: Training Autoregressive Language Models with Limited Domain Data

熵引导的标记丢弃:在有限领域数据下训练自回归语言模型

Jiapeng Wang, Yiwen Hu, Yanzipeng Gao, Haoyu Wang, Shuo Wang, Hongyu Lu, Jiaxin Mao, Wayne Xin Zhao, Junyi Li, Xiao Zhang

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Tsinghua University(清华大学) WeChat, Tencent(微信、腾讯) Department of Data Science, City University of Hong Kong(香港城市大学数据科学系)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 本文提出熵引导的标记丢弃方法,通过正则化优化解决有限领域数据下自回归模型的性能退化问题,实验表明其在多轮训练中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22742 2025-12-30 cs.DB cs.AI 57%

Robust LLM-based Column Type Annotation via Prompt Augmentation with LoRA Tuning

基于提示增强与LoRA微调的鲁棒列类型标注

Hanze Meng, Jianhao Cao, Rachel Pottinger

机构 * University of British Columbia(不列颠哥伦比亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出基于提示增强与LoRA微调的鲁棒列类型标注方法,通过减少可训练参数提升模型稳定性与性能。

Comments 13 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22605 2025-12-30 cs.AI cs.CV 57%

Learning Multi-Modal Mobility Dynamics for Generalized Next Location Recommendation

学习多模态移动动态以实现通用的下一站推荐

Junshu Dai, Yu Wang, Tongya Zheng, Wei Ji, Qinghong Guo, Ji Cao, Jie Song, Canghong Jin, Mingli Song

机构 * Zhejiang University(浙江大学) Hangzhou City University(杭州市大学) Nanjing University(南京大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出多模态移动(M^3ob)方法,通过构建统一时空关系图和门控机制,提升位置推荐任务的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22302 2025-12-30 cs.LG 57%

Statistical and Machine Learning Analysis of Traffic Accidents on US 158 in Currituck County: A Comparison with HSM Predictions

对美国158号公路在Currituck县交通事故的统计与机器学习分析:与HSM预测的比较

Jennifer Sawyer, Julian Allagan

机构 * Elizabeth City State University(埃里克森州立大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 本研究利用统计与机器学习方法分析美国158号公路交通事故,通过随机森林预测受伤严重程度并验证热点区域,为提升交通安全性提供方法学支持。

Comments 9 pages,7 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22154 2025-12-30 cs.CR cs.AI cs.MA 57%

Practical challenges of control monitoring in frontier AI deployments

前沿AI部署中的控制监控实际挑战

David Lindner, Charlie Griffin, Tomek Korbak, Roland S. Zimmermann, Geoffrey Irving, Sebastian Farquhar, Alan Cooney

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 本文探讨了前沿AI部署中控制监控面临的实际挑战,分析了同步、半同步和异步监控方案,并通过案例研究探讨了监督、延迟和恢复等关键问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22146 2025-12-30 eess.SP cs.LG cs.SD 57%

EEG-to-Voice Decoding of Spoken and Imagined speech Using Non-Invasive EEG

通过非侵入性EEG解码 spoken 和 imagined 语音

Hanbeot Park, Yunjeong Cho, Hunhee Kim

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本研究提出了一种无需显式时间对齐的EEG-to-Voice方法,通过非侵入性EEG信号直接重建说话和想象语音,结合生成器、语音编码器和自动语音识别模块,实现了稳定的声学重建和可比的语言准确性。

Comments 20 pages, 7 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20368 2025-12-30 cs.AI 57%

AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective Reinforcement Learning

AI-SearchPlanner: 通过帕累托最优多目标强化学习实现模块化代理搜索

Lang Mei, Zhihan Yang, Xiaohan Yu, Huanyao Zhang, Chong Chen

机构 * Huawei Cloud BU, China(华为云业务部,中国) School of Computer Science, Peking University(北京大学计算机学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 AI-SearchPlanner通过帕累托最优多目标强化学习提升冻结QA模型的搜索规划性能,实现模块化代理搜索。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23483 2025-12-30 cs.CV 50%

TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding

TV-RAG:一种具有时间意识和语义熵权的长视频检索与理解框架

Zongsheng Cao, Yangfan He, Anran Liu, Feng Chen, Zepeng Wang, Jun Xie

机构 * Researcher(研究者)

专题命中 其他安全 :alignment(abstract)

AI总结 TV-RAG通过时间衰减检索和熵加权关键帧采样,提升长视频检索与理解性能,无需重新训练即可集成至现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23320 2025-12-30 cs.MM 50%

Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions

多智能体语义情感对齐的音乐到图像生成与音乐衍生标题

Junchang Shi, Gang Li

专题命中 其他安全 :alignment(abstract)

AI总结 MESA MIG通过多智能体协作生成音乐到图像,实现语义和情感对齐,提升生成图像的审美质量和情感一致性。

Comments 10 pages,3 figures.Under review for ICME 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05822 2025-12-30 cs.CV 50%

Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models

通过融合来自LLM的世界知识与视觉基础模型进行视频事件推理与预测

L'ea Dubois, Klaus Schmidt, Chengyu Wang, Ji-Hoon Park, Lin Wang, Santiago Munoz

机构 * INRIA(法国国家信息与自动化研究所) Max Planck Institute for Intelligent Systems(人工智能研究所) San Francisco State University(旧金山州立大学) Seoul AI Institute (SAII)(首尔人工智能研究所) Vision & Robotics Center, Tsinghua University(清华大学视觉与机器人中心) Polytechnic University of Madrid(马德里理工大学)

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出融合视觉基础模型与LLM的世界知识,以提升视频事件推理与预测能力,实现从简单识别到高级认知理解的突破。

Comments 22 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22680 2025-12-30 eess.SY cs.SY 50%

From Electrochemical Energy Storage to Next-Generation Intelligent Battery Technologies for Electric Vehicles: A Survey

从电化学储能到下一代智能电池技术:电动汽车的综述

Abderaouf Bahi, Amel Ourici, Chaima Lagraa, Siham Lameche, Soundess Halimi, Inoussa Mouiche, Ylias Sabri, Waseem Haider, Mohamed Trari

专题命中 其他安全 :safety(abstract)

AI总结 本文综述了电化学储能技术的最新进展,探讨了智能电池管理系统中机器学习和AI的应用,旨在为电动汽车领域提供下一代电池技术的全面理解。

Comments This work was supervised by leading professor in the field (Pr. Mohamed Trari, Pr. Waseem Haider, Pr. Ylias Sabri)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22454 2025-12-30 cs.CV 50%

Comparing Object Detection Models for Electrical Substation Component Mapping

比较电气变电站组件映射的检测模型

Haley Mody, Namish Bansal, Dennies Kiprono Bor, Edward J. Oughton

专题命中 其他安全 :safety(abstract)

AI总结 本文比较了三种检测模型在电气变电站组件映射中的性能,旨在提高变电站基础设施的自动化检测和映射能力。

Comments 26 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏