机构
*
Department of Computer Science, National University of Singapore(新加坡国立大学计算机科学系)
;
Singapore-MIT Alliance for Research and Technology Centre(新加坡-麻省理工联盟研究技术中心)
;
The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳))
;
CSAIL, Massachusetts Institute of Technology(麻省理工学院计算机科学与人工智能实验室)
;
Institute of Data Science, National University of Singapore(新加坡国立大学数据科学研究院)
RTLC -- Research, Teach-to-Learn, Critique: A three-stage prompting paradigm inspired by the Feynman Learning Technique that lifts LLM-as-judge accuracy on JudgeBench with no fine-tuning
Sitao Cheng, Tianle Li, Xuhan Huang, Xunjian Yin, Difan Zou
机构
*
Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家)
;
School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家)
Learning a Continue-Thinking Token for Enhanced Test-Time Scaling
学习一个持续思考标记以增强测试时扩展
Liran Ringel, Elad Tolochinsky, Yaniv Romano
机构
*
Department of Computer Science, Technion – Israel Institute of Technology(计算机科学系,技术学院–以色列理工学院)
;
Department of Electrical and Computer Engineering, Technion – Israel Institute of Technology(电气与计算机工程系,技术学院–以色列理工学院)
ChatSR: Multimodal Large Language Models for Scientific Formula Discovery
ChatSR:用于科学公式发现的多模态大语言模型
Yanjie Li, Lina Yu, Weijun Li, Min Wu, Liping Zhang, Jingyi Liu, Yusong Deng, Mingzhu Wan, Xin Ning
机构
*
AnnLab, Institute of Semiconductors, Chinese Academy of Sciences, Beijing, China(安 lab,半导体研究所,中国科学院,北京,中国)
;
School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences, Beijing, China(电子、电气与通信工程学院,中国科学院大学,北京,中国)
;
Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)
;
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing 101408, China(先进交叉科学学院,中国科学院大学,北京101408,中国)
;
College of Materials Science and Opto-Electronic Technology, University of Chinese Academy of Sciences, Beijing, 100049, China(材料科学与光电技术学院,中国科学院大学,北京100049,中国)
;
School of Integrated Circuits, University of Chinese Academy of Sciences, Beijing 100049, China(集成电路学院,中国科学院大学,北京100049,中国)
Breaking Contextual Inertia: Reinforcement Learning with Single-Turn Anchors for Stable Multi-Turn Interaction
打破情境惯性:基于单轮锚点的强化学习用于稳定多轮交互
Xingwu Chen, Zhanqiu Zhang, Yiwen Guo, Difan Zou
机构
*
Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家)
;
School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家)
机构
*
Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家)
;
School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家)
;
JP Morgan AI Research, London, UK(摩根大通AI研究,伦敦,英国)
;
JP Morgan AI Research, New York, USA(摩根大通AI研究,纽约,美国)
Ismam Nur Swapnil, Aranya Saha, Tanvir Ahmed Khan, Mohammad Ariful Haque, Ser-Nam Lim
机构
*
Bangladesh University of Engineering and Technology(孟加拉工程与技术大学)
;
University of Maryland, College Park(马里兰大学学院公园分校)
;
Illinois Institute of Technology(伊利诺伊理工学院)
;
University of Central Florida(佛罗里达中央大学)
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
重新思考RLVR中的熵干预:从熵变化视角
Zhezheng Hao, Hong Wang, Haoyang Liu, Jian Luo, Jiarui Yu, Hande Dong, Qiang Lin, Can Wang, Jiawei Chen
机构
*
State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室)
;
Zhejiang University(浙江大学)
;
Tencent(腾讯)
;
Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新技术区(滨江)区块链与数据安全研究院)