Rewarding How Models Think Pedagogically: Integrating Pedagogical Reasoning and Thinking Rewards for LLMs in Education
奖励模型如何思考:在教育中整合教学推理和思考奖励
Unggi Lee, Jiyeong Bae, Jaehyeon Park, Haeun Park, Taejun Park, Younghoon Jeon, Sungmin Cho, Junbo Koh, Yeil Jeong, Gyeonggeon Lee
机构
*
Chosun University(昌原大学)
;
Korea University(韩国大学)
;
Seoul National University(首尔国立大学)
;
Korea Institute for Curriculum and Evaluation(韩国课程与评价研究院)
;
Upstage
;
Indiana University Bloomington(印第安纳大学布卢明顿分校)
;
Nanyang Technological University(南洋理工大学)
TreeWriter: AI-Assisted Hierarchical Planning and Writing for Long-Form Documents
TreeWriter: 人工智能辅助的长篇文档分层规划与写作
Zijian Zhang, Fangshi Du, Xingjian Liu, Pan Chen, Oliver Huang, Runlong Ye, Michael Liut, Alán Aspuru-Guzik
机构
*
Department of Computer Science, University of Toronto, Sandford Fleming Building, 10 King’s College Road, ON M5S 3G4, Toronto, Canada
;
Vector Institute for Artificial Intelligence, 661 University Ave. Suite 710, ON M5G 1M1, Toronto, Canada
;
Department of Chemistry, University of Toronto, Lash Miller Chemical Laboratories, 80 St. George Street, ON M5S 3H6, Toronto, Canada
;
Department of Mathematical
;
Computational Sciences, University of Toronto Mississauga, 3359 Mississauga Road, Deerfield Hall, ON L5L 1C6, Mississauga, Canada
;
Department of Materials Science \& Engineering, University of Toronto, 184 College St., M5S 3E4, Toronto, Canada
;
Department of Chemical Engineering \& Applied Chemistry, University of Toronto, 200 College St. ON M5S 3E5, Toronto, Canada
;
Acceleration Consortium, 700 University Ave., M7A 2S4, Toronto, Canada
;
Canadian Institute for Advanced Research (CIFAR), 661 University Ave., M5G 1M1, Toronto, Canada
;
NVIDIA, 431 King St W \#6th, M5V 1K4, Toronto, Canada
Compartmentalised Agentic Reasoning for Clinical NLI
临床自然语言推理中的 compartmentalised agentic 推理
Maël Jullien, Lei Xu, Marco Valentino, André Freitas
机构
*
Department of Computer Science, University of Manchester, UK(曼彻斯特大学计算机科学系)
;
National Biomarker Centre, CRUK-MI, University of Manchester, UK(曼彻斯特大学国家生物标记中心)
;
Idiap Research Institute, Switzerland(日内瓦IDIAP研究所)
;
School of Computer Science, University of Sheffield, UK(谢菲尔德大学计算机科学系)
;
École Polytechnique Fédérale de Lausanne (EPFL), Switzerland(洛桑联邦理工学院(EPFL))
UserLM-R1: Modeling Human Reasoning in User Language Models with Multi-Reward Reinforcement Learning
UserLM-R1: 通过多奖励强化学习建模人类推理在用户语言模型中的应用
Feng Zhang, Shijia Li, Chunmao Zhang, Zhanyu Ma, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He, Jingwen Xu, Han Liu
机构
*
Meituan(美团)
;
Peking University(北京大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Dalian University of Technology(大连理工大学)
UNCAP: Uncertainty-Guided Neurosymbolic Planning Using Natural Language Communication for Cooperative Autonomous Vehicles
UNCAP:基于自然语言通信的不确定性引导神经符号规划
Neel P. Bhatt, Po-han Li, Kushagra Gupta, Rohan Siva, Daniel Milan, Alexander T. Hogue, Sandeep P. Chinchali, David Fridovich-Keil, Zhangyang Wang, Ufuk Topcu
机构
*
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构
*
Tsinghua University, Beijing, China(清华大学)
;
JD.com, Beijing, China(京东)
;
Faculty of Engineering and Faculty of Business and Economics, The University of Hong Kong, China(香港大学工程学院和商学院)
;
College of Engineering, University of California, Berkeley, CA, USA(加州大学伯克利分校工程学院)
机构
*
School of Computer Science and Engineering, Beihang University, Beijing, China(北京航空航天大学计算机科学与工程学院)
;
School of Software, Beihang University, Beijing, China(北京航空航天大学软件学院)
;
Hangzhou Innovation Institute, Beihang University, Hangzhou, China(北京航空航天大学杭州创新研究院)
Knowledge Distillation for Temporal Knowledge Graph Reasoning with Large Language Models
基于大语言模型的时序知识图谱推理知识蒸馏
Wang Xing, Wei Song, Siyu Lin, Chen Wu, Zhesi Li, Man Wang
机构
*
School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院)
;
School of Computing and Artificial Intelligence, Southwest Jiaotong University(西南交通大学计算机与人工智能学院)
;
School of Information Science and Engineering, Chongqing Jiaotong University(重庆交通大学信息科学与工程学院)
;
School of Information Engineering, Chang’an University(长安大学信息工程学院)
Towards Responsible and Explainable AI Agents with Consensus-Driven Reasoning
迈向负责任和可解释的AI代理:基于共识驱动的推理
Eranga Bandara, Tharaka Hewa, Ross Gore, Sachin Shetty, Ravi Mukkamala, Peter Foytik, Abdul Rahman, Safdar H. Bouk, Xueping Liang, Amin Hass, Sachini Rajapakse, Ng Wee Keong, Kasun De Zoysa, Aruna Withanage, Nilaan Loganathan
机构
*
Old Dominion University, Norfolk, VA, USA(老奥本大学)
;
Center for Wireless Communications, University of Oulu, Finland(无线通信中心,奥卢大学)
;
Florida International University, USA(佛罗里达国际大学)
;
Nanyang Technological University, Singapore(南洋理工大学)
;
University of Colombo, Sri Lanka(科伦坡大学)
;
Accenture Technology Labs, Arlington, VA, USA(Accenture技术实验室)
Trajectory Planning for UAV-Based Smart Farming Using Imitation-Based Triple Deep Q-Learning
基于模仿学习的三重深度Q学习用于无人机智能农业的轨迹规划
Wencan Mao, Quanxi Zhou, Tomas Couso Coddou, Manabu Tsukada, Yunling Liu, Yusheng Ji
机构
*
National Institute of Informatics(国立信息研究所)
;
The University of Tokyo(东京大学)
;
Pontificia Universidad Católica de Chile(智利天主教大学)
;
China Agricultural University(中国农业大学)