Contextual Online Uncertainty-Aware Preference Learning for Human Feedback
基于上下文的在线不确定性感知偏好学习用于人类反馈
Nan Lu, Ethan Lee, Ethan X. Fang, Junwei Lu
机构
*
Department of Biostatistics, Harvard T.H. Chan School of Public Health(哈佛大学T.H.陈公共卫生学院生物统计学系)
;
Department of Biostatistics & Bioinformatics, Duke University(杜克大学生物统计学与生物信息学系)
EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training
EstLLM:通过持续预训练和后训练增强多语言大语言模型中的爱沙尼亚能力
Aleksei Dorkin, Taido Purason, Emil Kalbaliyev, Hele-Andra Kuulmets, Marii Ojastu, Mark Fišel, Tanel Alumäe, Eleri Aedmaa, Krister Kruusmaa, Kairit Sirts
机构
*
Institute of Computer Science, University of Tartu(塔尔图大学计算机科学研究院)
;
Department of Software Science, Tallinn University of Technology(塔林技术大学软件科学系)
;
Institute of the Estonian Language, Tallinn, Estonia(爱沙尼亚语言研究院)
;
School of Humanities, Tallinn University(塔林大学人文学院)
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Beijing Institute of AI Safety and Governance(北京人工智能安全与治理研究院)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
机构
*
Qwen Large Model Application Team, Alibaba(阿里巴巴文勤大模型应用团队)
;
Beijing University Of Posts and Telecommunications(北京邮电大学)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
Comments12 pages, 12 figures. Independent research. Code and artifacts: this https URL (https://github.com/Harry-Ashley/adversa-guardrail-degradation). Published in the 2026 IEEE/ACIS 24th International Conference on Software Engineering Research, Management and Applications (SERA), pp. 414-417. Recipient of the Best Paper Award at IEEE/ACIS SERA 2026
MPIB: A Benchmark for Medical Prompt Injection Attacks and Clinical Safety in LLMs
MPIB:用于医疗提示注入攻击和大语言模型临床安全性的基准
Junhyeok Lee, Han Jang, Kyu Sung Choi
机构
*
Seoul National University College of Medicine(首尔国立大学医学院)
;
Seoul National University(首尔国立大学)
;
Seoul National University College of Medicine, Seoul National University Hospital(首尔国立大学医学院、首尔国立大学医院)
CASTLE: A Comprehensive Benchmark for Evaluating Student-Tailored Personalized Safety in Large Language Models
CASTLE:一个评估大语言模型学生定制个性化安全性的综合基准
Rui Jia, Ruiyi Lan, Fengrui Liu, Zhongxiang Dai, Bo Jiang, Jing Shao, Jingyuan Chen, Guandong Xu, Fei Wu, Min Zhang
机构
*
East China Normal University(华东师范大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Zhejiang University(浙江大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
The Education University of Hong Kong(香港教育大学)
CommentsThe experiment has a design defect: the VSA drift trigger was not in the causal path of its own runs, so the results do not test the claimed mechanism. Clean-room replication: content-aware re-injection beats neither random re-injection (p=0.46) nor a fixed timer (p=0.29), while positive controls confirm task signal (p=0.002, p=0.001). The conclusions are unsupported and should not be cited
Comments47 pages (18 pages main text, 29 pages Supporting Information), 9 figures, 2 tables, 70 references. SI contains complete proofs of all theorems, extended experiments, and the verification methodology
Retrieval-aligned Tabular Foundation Models Enable Robust Clinical Risk Prediction in Electronic Health Records Under Real-world Constraints
检索对齐的表格基础模型实现电子健康记录中在现实约束下的稳健临床风险预测
Minh-Khoi Pham, Thang-Long Nguyen Ho, Thao Thi Phuong Dao, Tai Tan Mai, Minh-Triet Tran, Marie E. Ward, Una Geary, Rob Brennan, Nick McDonald, Martin Crane, Marija Bezbradica