arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-30 至 2026-01-30 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2510.23596 2026-01-30 cs.CL 70%

Think Twice: Branch-and-Rethink Reasoning Reward Model

再思两次:分支与再思考推理奖励模型

Yizhu Jiao, Jiaqi Zeng, Julien Veron Vialard, Oleksii Kuchaiev, Jiawei Han, Olivier Delalleau

机构 * NVIDIA(NVIDIA公司) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 偏好对齐 :RLHF(abstract);safety(abstract);分类 cs.CL

AI总结 BR-RM通过两轮次的分支与再思考机制,改进奖励模型的判断准确性与敏感性,实现更精确的推理能力。

Comments Source Code: https://github.com/yzjiao/BR-RM. Model Checkpoints: https://huggingface.co/nvidia/Qwen3-Nemotron-14B-BRRM and https://huggingface.co/nvidia/Qwen3-Nemotron-8B-BRRM

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21611 2026-01-30 cs.IR cs.AI cs.CL 62%

Thinking Broad, Acting Fast: Latent Reasoning Distillation from Multi-Perspective Chain-of-Thought for E-Commerce Relevance

深度思考,快速行动:多视角链式推理的潜在推理蒸馏用于电子商务相关性

Baopu Qiu, Hao Chen, Yuanrong Wu, Changtong Zan, Chao Wei, Weiru Zhang, Xiaoyi Zeng

机构 * Alibaba International Digital Commerce Group(阿里巴巴国际数字商业集团) Zhejiang University(浙江大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

AI总结 本文提出多视角链式推理蒸馏方法,通过改进的教师模型和轻量级学生模型提升电子商务搜索相关性建模的准确性和效率。

Comments 12 pages, 6 figures, Accepted by WWW2026 industry track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11311 2026-01-30 cs.AI cs.CY 62%

Prompts to Proxies: Emulating Human Preferences via a Compact LLM Ensemble

通过紧凑的LLM集合模拟人类偏好:提示到代理

Bingchen Wang, Zi-Yu Khoo, Jingtan Wang

机构 * Independent Researcher, China(中国独立研究者) School of Computing, National University of Singapore(新加坡国立大学计算机学院)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 通过紧凑的LLM集合模拟人类偏好,P2P方法在无需微调和敏感数据的情况下,有效重建目标人群偏好并实现高质量预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21740 2026-01-30 cs.MM cs.SD 50%

MIDI-LLaMA: An Instruction-Following Multimodal LLM for Symbolic Music Understanding

MIDI-LLaMA:一种用于符号音乐理解的指令遵循多模态大语言模型

Meng Yang, Jon McCormack, Maria Teresa Llano, Wanchao Su, Chao Lei

机构 * SensiLab, Monash University, Australia(Monash大学) University of Sussex, Brighton, United Kingdom(Sussex大学) School of Computing and Information Systems, The University of Melbourne, Australia(墨尔本大学计算机与信息系统学院)

专题命中 偏好对齐 :alignment(abstract)

AI总结 MIDI-LLaMA通过结合MusicBERT和Llama-3-8B,实现了对符号音乐的指令遵循理解,显著提升了音乐描述和语义对齐能力。

Comments Accepted for publication at International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏