arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-12-30 至 2025-12-30 共收录 63 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 18 篇

2410.01324 2025-12-30 cs.LG cs.AI 62%

Fair Class-Incremental Learning using Sample Weighting

基于样本加权的公平类增量学习

Jaeyoung Park, Minsu Kim, Steven Euijong Whang

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种基于样本加权的公平类增量学习框架,通过调整训练权重减少敏感群体的遗忘,提升模型公平性。

Comments 30 pages, 30 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22280 2025-12-30 cs.LG cs.AI cs.DB cs.DC 62%

Valori: A Deterministic Memory Substrate for AI Systems

Valori:一种用于人工智能系统的确定性内存基质

Varshith Gudur

机构 * Valori Kernel Project(Valori内核项目)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 Valori通过使用固定点算术和可重放状态机,为人工智能系统提供确定性内存解决方案,解决浮点运算导致的非确定性问题。

Comments 7 pages, 1 figure. systems paper with empirical evaluation and determinism validation experiments. Code available at https://github.com/varshith-Git/Valori-Kernel

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22258 2025-12-30 cs.AI cs.LG cs.LO cs.SC 62%

Logic Sketch Prompting (LSP): A Deterministic and Interpretable Prompting Method

逻辑草图提示(LSP):一种确定性和可解释的提示方法

Satvik Tripathi

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

AI总结 逻辑草图提示(LSP)通过引入类型变量、确定性条件评估器和规则验证器,提升了大语言模型在需要严格规则遵守、确定性和可审计性任务中的性能和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22178 2025-12-30 cs.LG cs.AI 62%

Wireless Traffic Prediction with Large Language Model

基于大语言模型的无线交通预测

Chuanting Zhang, Haixia Zhang, Jingping Qiao, Zongzhang Li, Mohamed-Slim Alouini

机构 * Shandong Key Laboratory of Intelligent Communication and Sensing-Computing Integration, Shandong University(山东大学智能通信与感知-计算整合重点实验室) School of Information Science and Engineering, Shandong Normal University(山东师范大学信息科学与工程学院) China Mobile Communications Group Shandong Co., Ltd(中国移动通信集团山东有限公司) Computer, Electrical and Mathematical Science and Engineering Division, King Abdullah University of Science and Technology (KAUST)(卡布斯大学科学与工程学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出TIDES框架,利用大语言模型和DeepSeek模块,通过空间对齐和个性化模型提升无线交通预测的精度和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23445 2025-12-30 cs.MA cs.LG cs.RO 57%

Assessing behaviour coverage in a multi-agent system simulation for autonomous vehicle testing

评估多智能体系统模拟在自动驾驶测试中的行为覆盖率

Manuel Franco-Vivo

专题命中 安全评测 :safety(abstract);分类 cs.LG

AI总结 本研究提出了一种基于模型预测控制的行人智能体,用于评估多智能体系统模拟在自动驾驶测试中的行为覆盖率,以提升系统安全性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23217 2025-12-30 cs.AI 57%

TCEval: Using Thermal Comfort to Assess Cognitive and Perceptual Abilities of AI

TCEval:利用热舒适性评估人工智能的认知与感知能力

Jingming Li

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 TCEval通过热舒适性场景评估AI的跨模态推理、因果关联和适应性决策能力,揭示当前LLM在热舒适性领域存在基础推理能力但缺乏精确因果理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21580 2025-12-30 cs.CL 57%

Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM

Gamayun的多语言 mastery 之路:一种 1.5B 参数 LLM 的高效训练

Alexander Podolskiy, Semen Molokov, Timofey Gerasin, Maksim Titov, Alexey Rukhovich, Artem Khrapov, Kirill Morozov, Evgeny Tetin, Constantine Korikov, Pavel Efimov, Polina Lazukova, Yuliya Skripkar, Nikita Okhotnikov, Irina Piontkovskaya, Meng Xiaojun, Zou Xueyi, Zhang Zhenhe

机构 * Gamayun Team(Gamayun团队)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 Gamayun 通过两阶段预训练策略,在资源受限环境下高效训练出 1.5B 参数多语言模型,支持 12 种语言,尤其在俄语上取得先进成果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12411 2025-12-30 cs.CL 57%

The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases

大语言模型的文化基因:跨语料库训练对模型价值观与偏见的影响研究

Emanuel Z. Fenech-Borg, Tilen P. Meznaric-Kos, Milica D. Lekovic-Bojovic, Arni J. Hentze-Djurhuus

机构 * Department of Communications(通讯系) University of Malta(马耳他大学) Faculty of Mathematics(数学系) University of Primorska(普里摩尔卡大学) Faculty of Electrical Engineering(电气工程系) University of Montenegro(黑山大学) Faculty of Science & Technology(科学与技术系) University of the Faroe Islands(法罗群岛大学) Department of Computer Science(计算机科学系) San Francisco State University(旧金山州立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 本研究通过跨文化探针数据集分析大语言模型在个人主义-集体主义和权力距离维度上的文化偏见,揭示模型价值观与训练语料文化背景的关联。

Comments 10 pages, 5 figures, IEEE conference format, submitted to [Conference Name]

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04654 2025-12-30 eess.IV cs.LG 57%

Image and Video Quality Assessment using Prompt-Guided Latent Diffusion Models for Cross-Dataset Generalization

基于提示引导的潜在扩散模型的图像和视频质量评估用于跨数据集泛化

Shankhanil Mitra, Diptanu De, Shika Rao, Rajiv Soundararajan

机构 * Samsung Research Institute(三星研究机构) Qualcomm(高通) New York University(纽约大学) Indian Institute of Science(印度科学研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

AI总结 本文提出基于提示引导的潜在扩散模型,通过学习跨注意力图和时间质量调节器,实现图像和视频质量评估的跨数据集泛化。

Comments Accepted to Transactions on Machine Learning Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22629 2025-12-30 cs.AI cs.IR 57%

DICE: Discrete Interpretable Comparative Evaluation with Probabilistic Scoring for Retrieval-Augmented Generation

DICE:基于概率评分的离散可解释比较评估用于检索增强生成

Shiyan Liu, Jian Ma, Rui Qu

机构 * School of Computer Science and Technology(计算机科学与技术学院) Huazhong University of Science and Technology(华中科技大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 DICE提出一种基于概率评分的两阶段框架,提升RAG评估的可解释性和稳健性,通过透明判断和系统性错误诊断,实现高效且可信的评估。

Comments Accepted at ResponsibleFM @ NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12382 2025-12-30 cs.CL 57%

LLM-as-a-Judge: Rapid Evaluation of Legal Document Recommendation for Retrieval-Augmented Generation

LLM-as-a-Judge: 快速评估法律文档推荐用于检索增强生成

Anu Pradhan, Alexandra Ortan, Apurv Verma, Madhavan Seshadri

机构 * Bloomberg(彭博)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 本文提出LLM-as-a-Judge方法,通过改进评估指标和统计检验,实现法律文档推荐系统的自动化评估,提高评估效率和准确性。

Comments Accepted in EARL 25: The 2nd Workshop on Evaluating and Applying Recommender Systems with Large Language Models at RecSys 2025

Journal ref Proceedings of the 2nd Workshop on Evaluating and Applying Recommender Systems with Large Language Models (EARL), RecSys 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23033 2025-12-30 cs.SE cs.CV 50%

Interpretable Gallbladder Ultrasound Diagnosis: A Lightweight Web-Mobile Software Platform with Real-Time XAI

可解释的胆囊超声诊断:一种轻量级的网页-移动软件平台与实时XAI

Fuyad Hasan Bhoyan, Prashanta Sarker, Parsia Noor Ethila, Md. Emon Hossain, Md Kaviul Hossain, Md Humaion Kabir Mehedi

机构 * University of Liberal Arts Bangladesh(乌拉尔大学 Bangladesh) BRAC University(BRAC大学)

专题命中 安全评测 :trustworthy(abstract)

AI总结 本研究提出了一种轻量级网页-移动软件平台,利用混合深度学习模型MobResTaNet实现胆囊疾病的实时可解释诊断,准确率达99.85%。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 2 篇

2512.22725 2025-12-30 cs.CL cs.CY 62%

Mitigating Social Desirability Bias in Random Silicon Sampling

在随机硅采样中缓解社会可取性偏差

Sashank Chapala, Maksym Mironov, Songgaojun Deng

机构 * Eindhoven University of Technology(埃因霍温理工大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

AI总结 本研究通过提示工程方法减轻LLM中的社会可取性偏差,提升硅样本与人类数据的一致性。

Comments 31 pages, 9 figures, and 24 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23430 2025-12-30 cs.CL 57%

C2PO: Diagnosing and Disentangling Bias Shortcuts in LLMs

C2PO:诊断和解构大语言模型中的偏见捷径

Xuan Feng, Bo An, Tianlong Gu, Liang Chang, Fengrui Hao, Peipeng Yu, Shuai Zhao

机构 * Jinan University(济南大学) Nanyang Technological University(南洋理工大学) Engineering Research Center of Trustworthy AI (Ministry of Education)(可信人工智能工程研究中心) Guangxi Key Laboratory of Trusted Software(广西可信软件重点实验室)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 C2PO通过因果对比偏好优化框架,解决大语言模型中的偏见问题,同时保持推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 19 篇

2511.21218 2025-12-30 cs.CL 79%

Can Finetuing LLMs on Small Human Samples Increase Heterogeneity, Alignment, and Belief-Action Coherence?

能否通过在小样本人类数据上微调LLM增加异质性、对齐性和信念-行为一致性?

Steven Wang, Kyle Hunt, Shaojie Tang, Kenneth Joseph

机构 * Department of Management Science and Systems, University at Buffalo(管理科学与系统系,布法罗大学) Department of Computer Science and Engineering, University at Buffalo(计算机科学与工程系,布法罗大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

AI总结 本研究探讨通过微调LLM在小样本人类数据上是否能提高模拟结果的异质性、对齐性和信念-行为一致性,发现虽有提升但无法完全替代人类数据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22741 2025-12-30 cs.CL 79%

Text-Routed Sparse Mixture-of-Experts Model with Explanation and Temporal Alignment for Multi-Modal Sentiment Analysis

基于解释和时序对齐的文本引导稀疏专家混合模型用于多模态情感分析

Dongning Rao, Yunbiao Zeng, Zhihua Jiang, Jujian Lv

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

AI总结 本文提出基于解释和时序对齐的文本引导稀疏专家混合模型,用于提升多模态情感分析的性能。

Comments 9 pages, 9 figures, accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00115 2025-12-30 cs.CL cs.AI 76%

Cognitive Alignment in Personality Reasoning: Leveraging Prototype Theory for MBTI Inference

人格推理中的认知对齐:利用原型理论进行MBTI推断

Haoyuan Li, Yuanbo Tong, Yuchen Li, Zirui Wang, Chunhou Liu, Jiamou Liu

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 ProtoMBTI通过原型理论与认知对齐,提升文本人格推断的准确性和泛化能力。

Comments The authors have decided to withdraw this version to substantially revise and extend the work

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15545 2025-12-30 cs.CL cs.AI 76%

TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs

TokenTiming: 一种适用于通用推测解码模型对的动态对齐方法

Sibo Xiao, Jinyuan Fu, Zhongle Xie, Lidan Shou

机构 * Zhejiang University(浙江大学)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 TokenTiming通过动态对齐方法实现通用推测解码,无需重新训练即可处理不同词汇模型,提升LLM推理效率1.57倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23487 2025-12-30 cs.LG cs.AI stat.ML 73%

ML Compass: Navigating Capability, Cost, and Compliance Trade-offs in AI Model Deployment

ML Compass:在AI模型部署中导航能力、成本和合规性之间的权衡

Vassilis Digalakis, Ramayya Krishnan, Gonzalo Martin Fernandez, Agni Orfanoudaki

机构 * Questrom School of Business, Boston University(波士顿大学Questrom商学院) Centre de Formació Interdisciplinària Superior and Universitat Politècnica de Catalunya(巴塞罗那理工大学跨学科教育中心)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 ML Compass通过系统性框架优化模型选择,考虑能力、成本和合规性之间的权衡,提供部署导向的推荐和排行榜。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08270 2025-12-30 cs.LG cs.AI cs.CL cs.MM 67%

Doctor Sun: A Bilingual Multimodal Large Language Model for Biomedical AI

Doctor Sun: 一种双语多模态大语言模型用于生物医学AI

Dong Xue, Ziyao Shao, Zhaoyang Duan, Fangzhou Liu, Bing Li, Zhongheng Zhang

机构 * Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education East China University of Science and Technology(能源化工过程智能制造重点实验室,东华大学) Research Institute of Intelligent Control and Systems Harbin Institute of Technology(智能控制与系统研究室,哈尔滨工业大学) Department of Emergency Medicine, Sir Run Run Shaw Hospital Zhejiang University School of Medicine(浙江大学医学院急诊医学科) Provincial Key Laboratory of Precise Diagnosis Treatment of Abdominal Infection, Sir Run Run Shaw Hospital Zhejiang University School of Medicine(腹部感染精准诊断治疗省级重点实验室,浙江大学医学院) School of Medicine Shaoxing University(绍兴大学医学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 Doctor Sun是一种双语多模态大语言模型,通过整合预训练视觉编码器和医学LLM,提升生物医学多模态任务的性能,并提供SunMed-VL数据集支持研究进展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04208 2025-12-30 cs.LG cs.AI 62%

Aligning Agents like Large Language Models

像大型语言模型一样对齐代理

Adam Jelley, Yuhan Cao, Dave Bignell, Amos Storkey, Sam Devlin, Tabish Rashid

机构 * School of Informatics, University of Edinburgh, Edinburgh, United Kingdom(爱丁堡大学信息学院) Micosoft Research, Cambridge, United Kingdom(微软研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过像训练大型语言模型一样训练代理,以实现更通用、稳健和对齐的行为,并通过3D视频游戏环境中的概念验证展示方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23422 2025-12-30 cs.CL 57%

Entropy-Guided Token Dropout: Training Autoregressive Language Models with Limited Domain Data

熵引导的标记丢弃:在有限领域数据下训练自回归语言模型

Jiapeng Wang, Yiwen Hu, Yanzipeng Gao, Haoyu Wang, Shuo Wang, Hongyu Lu, Jiaxin Mao, Wayne Xin Zhao, Junyi Li, Xiao Zhang

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Tsinghua University(清华大学) WeChat, Tencent(微信、腾讯) Department of Data Science, City University of Hong Kong(香港城市大学数据科学系)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 本文提出熵引导的标记丢弃方法,通过正则化优化解决有限领域数据下自回归模型的性能退化问题,实验表明其在多轮训练中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22742 2025-12-30 cs.DB cs.AI 57%

Robust LLM-based Column Type Annotation via Prompt Augmentation with LoRA Tuning

基于提示增强与LoRA微调的鲁棒列类型标注

Hanze Meng, Jianhao Cao, Rachel Pottinger

机构 * University of British Columbia(不列颠哥伦比亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出基于提示增强与LoRA微调的鲁棒列类型标注方法,通过减少可训练参数提升模型稳定性与性能。

Comments 13 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22605 2025-12-30 cs.AI cs.CV 57%

Learning Multi-Modal Mobility Dynamics for Generalized Next Location Recommendation

学习多模态移动动态以实现通用的下一站推荐

Junshu Dai, Yu Wang, Tongya Zheng, Wei Ji, Qinghong Guo, Ji Cao, Jie Song, Canghong Jin, Mingli Song

机构 * Zhejiang University(浙江大学) Hangzhou City University(杭州市大学) Nanjing University(南京大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出多模态移动(M^3ob)方法,通过构建统一时空关系图和门控机制,提升位置推荐任务的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22302 2025-12-30 cs.LG 57%

Statistical and Machine Learning Analysis of Traffic Accidents on US 158 in Currituck County: A Comparison with HSM Predictions

对美国158号公路在Currituck县交通事故的统计与机器学习分析:与HSM预测的比较

Jennifer Sawyer, Julian Allagan

机构 * Elizabeth City State University(埃里克森州立大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 本研究利用统计与机器学习方法分析美国158号公路交通事故,通过随机森林预测受伤严重程度并验证热点区域,为提升交通安全性提供方法学支持。

Comments 9 pages,7 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22154 2025-12-30 cs.CR cs.AI cs.MA 57%

Practical challenges of control monitoring in frontier AI deployments

前沿AI部署中的控制监控实际挑战

David Lindner, Charlie Griffin, Tomek Korbak, Roland S. Zimmermann, Geoffrey Irving, Sebastian Farquhar, Alan Cooney

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 本文探讨了前沿AI部署中控制监控面临的实际挑战,分析了同步、半同步和异步监控方案,并通过案例研究探讨了监督、延迟和恢复等关键问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22146 2025-12-30 eess.SP cs.LG cs.SD 57%

EEG-to-Voice Decoding of Spoken and Imagined speech Using Non-Invasive EEG

通过非侵入性EEG解码 spoken 和 imagined 语音

Hanbeot Park, Yunjeong Cho, Hunhee Kim

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本研究提出了一种无需显式时间对齐的EEG-to-Voice方法,通过非侵入性EEG信号直接重建说话和想象语音,结合生成器、语音编码器和自动语音识别模块,实现了稳定的声学重建和可比的语言准确性。

Comments 20 pages, 7 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20368 2025-12-30 cs.AI 57%

AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective Reinforcement Learning

AI-SearchPlanner: 通过帕累托最优多目标强化学习实现模块化代理搜索

Lang Mei, Zhihan Yang, Xiaohan Yu, Huanyao Zhang, Chong Chen

机构 * Huawei Cloud BU, China(华为云业务部,中国) School of Computer Science, Peking University(北京大学计算机学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 AI-SearchPlanner通过帕累托最优多目标强化学习提升冻结QA模型的搜索规划性能,实现模块化代理搜索。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23483 2025-12-30 cs.CV 50%

TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding

TV-RAG:一种具有时间意识和语义熵权的长视频检索与理解框架

Zongsheng Cao, Yangfan He, Anran Liu, Feng Chen, Zepeng Wang, Jun Xie

机构 * Researcher(研究者)

专题命中 其他安全 :alignment(abstract)

AI总结 TV-RAG通过时间衰减检索和熵加权关键帧采样,提升长视频检索与理解性能,无需重新训练即可集成至现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23320 2025-12-30 cs.MM 50%

Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions

多智能体语义情感对齐的音乐到图像生成与音乐衍生标题

Junchang Shi, Gang Li

专题命中 其他安全 :alignment(abstract)

AI总结 MESA MIG通过多智能体协作生成音乐到图像,实现语义和情感对齐,提升生成图像的审美质量和情感一致性。

Comments 10 pages,3 figures.Under review for ICME 2026

详情

展开后加载摘要…

URL PDF HTML 收藏