Less is More: Improving LLM Alignment via Preference Data Selection
少即是多:通过偏好数据选择改进大语言模型对齐
Xun Deng, Han Zhong, Rui Ai, Fuli Feng, Zheng Wang, Xiangnan He
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Peking University(北京大学)
;
Massachusetts Institute of Technology(麻省理工学院)
;
Alibaba Cloud Computing(阿里云计算)
;
MoE Key Lab of BIPC, University of Science and Technology of China(中国科学技术大学MoE关键实验室)
机构
*
Meta Superintelligence Labs(Meta超智能实验室)
;
Massachusetts Institute of Technology(麻省理工学院)
;
Arizona State University(亚利桑那州立大学)
;
University of Southern California(南加州大学)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
评分标准作为攻击面:LLM裁判中的隐蔽偏好漂移
Ruomeng Ding, Yifei Pang, He Sun, Yizhong Wang, Zhiwei Steven Wu, Zhun Deng
机构
*
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Yale University(耶鲁大学)
;
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构
*
School of Computer Science and Engineering, Northeastern University, Shenyang, China(东北大学计算机科学与工程学院)
;
Department of Computer Science and Technology, Tsinghua University, Beijing, China(清华大学计算机科学与技术系)
;
Huawei Technologies Co., Ltd(华为技术有限公司)
GRAIL: Goal Recognition Alignment through Imitation Learning
GRAIL:通过模仿学习进行目标识别对齐
Osher Elhadad, Felipe Meneguzzi, Reuth Mirsky
机构
*
Department of Computer Science, Bar Ilan University(巴伊兰大学计算机科学系)
;
Department of Computer Science, Aberdeen University(阿伯丁大学计算机科学系)
;
Department of Computer Science, Tufts University(塔夫茨大学计算机科学系)
机构
*
University College London(伦敦大学学院)
;
Ulsan National Institute of Science and Technology(釜山国立科学技术研究所)
;
University of Basel(巴塞尔大学)
;
University College London AI Centre(伦敦大学学院人工智能中心)
An Agentic AI Control Plane for 6G Network Slice Orchestration, Monitoring, and Trading
面向6G网络切片编排、监控与交易的代理式AI控制平面
Eranga Bandara, Ross Gore, Sachin Shetty, Ravi Mukkamala, Tharaka Hewa, Abdul Rahman, Xueping Liang, Safdar H. Bouk, Amin Hass, Peter Foytik, Ng Wee Keong, Kasun De Zoysa
机构
*
Florida International University, USA(佛罗里达国际大学)
;
Center for Wireless Communications, University of Oulu, Finland(无线通信中心,奥卢大学)
;
Accenture Technology Labs, Arlington, VA, USA(埃森哲技术实验室)
;
Nanyang Technological University, Singapore(南洋理工大学)
Image Can Bring Your Memory Back: A Novel Multi-Modal Guided Attack against Image Generation Model Unlearning
图像能帮你找回记忆:一种新颖的多模态引导攻击对抗图像生成模型去学习
Renyang Liu, Guanlin Li, Tianwei Zhang, See-Kiong Ng
机构
*
Institute of Data Science, National University of Singapore(数据科学研究所,新加坡国立大学)
;
S-Lab, Nanyang Technological University(南洋理工大学S实验室)
;
College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)
机构
*
Department of Computer Science(计算机科学系)
;
Princeton University(普林斯顿大学)
;
Center for Information Technology Policy(信息政策中心)
;
Department of Electrical and Computer Engineering(电气与计算机工程系)
;
NVIDIA(英伟达)
CommentsPaper accepted at the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026), main conference. 21 pages
Directional Concentration Uncertainty: A representational approach to uncertainty quantification for generative models
方向性集中不确定性:一种代表方法用于生成模型的不确定性量化
Souradeep Chattopadhyay, Brendan Kennedy, Sai Munikoti, Soumik Sarkar, Karl Pazdernik
机构
*
Department of Mechanical Engineering, Iowa State University, Ames, IA, USA(机械工程系,爱荷华州立大学)
;
Pacific Northwest National Laboratory, Richland, WA, USA(太平洋西北国家实验室)
;
Department of Statistics, North Carolina State University, Raleigh, NC, USA(统计系,北卡罗来纳州立大学)
机构
*
Department of Statistics and Data Science, Southern University of Science and Technology(统计与数据科学系,南方科技大学)
;
Department of Mathematics, The Chinese University of HongKong(数学系,香港中文大学)
;
School of Software Engineering, South China University of Technology(软件工程学院,华南理工大学)