EffGen: Enabling Small Language Models as Capable Autonomous Agents
EffGen: 使小型语言模型成为能干的自主智能体
Gaurav Srivastava, Aafiya Hussain, Chi Wang, Yingyan Celine Lin, Xuan Wang
机构
*
Department of Computer Science, Virginia Tech, Blacksburg, VA, USA(弗吉尼亚理工大学计算机科学系)
;
Georgia Institute of Technology, Atlanta, GA, USA(佐治亚理工学院)
;
Google DeepMind, USA(谷歌DeepMind)
CommentsAccepted to ICML 2026. The title has been changed from "From Correction to Mastery: Reinforced Distillation of Large Language Model Agents" to "Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs"; the camera-ready version has been uploaded
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms
哪种RAG范式在规模化场景下表现最优?检索增强生成范式的规模化研究
Pengyu Wang, Benfeng Xu, Shaohan Wang, Mingxuan Du, Xin Zeng, Huarui Wu, Lei Zhang, Licheng Zhang
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Metastone Technology(Metastone科技)
;
Information Technology Research Center, Beijing Academy of Agriculture and Forestry Sciences(北京农林科学院信息技术研究中心)
Toward Efficient Agents: Memory, Tool learning, and Planning
迈向高效智能体:记忆、工具学习与规划
Xiaofang Yang, Lijun Li, Heng Zhou, Tong Zhu, Xiaoye Qu, Yuchen Fan, Qianshan Wei, Rui Ye, Li Kang, Yiran Qin, Daizong Liu, Qi Li, Ning Ding, Siheng Chen, Jing Shao
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Fudan University(复旦大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Jiaotong University(上海交通大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
The Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳))
;
Hong Kong Polytechnic University(香港理工大学)
;
Wuhan University(武汉大学)
;
Tsinghua University(清华大学)
BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment
评估两用生物学环境中的校准拒绝和安全有用性
Edwin H. Wintermute, Harmon Bhasin, Christina M. Agapakis, Dianzhuo Wang, Evan Seeyave, Arjun Banerjee, Daniel Fulop, Matthew C. Watson, Adam J. Meyer, Sandrine Boissel, Jens H. Kuhn, Rishi Jain, Noah D. Taylor, Helena Shomar, Patrick M. Boyle, Kenny Workman
Commentsv4 adds 2 new appendices: app J evaluates a "generative" Bayesian alternative to BLF (which does worse than the "discriminative" approached used by BLF), and app K sketches how to do value of information computation to decide when to stop searching (although this idea has not been tried). v4 also adds a few more recent references
Journal refv3 was published in ICML AI Forecasting workshop 2026 (https://forecasting-workshop.github.io/)
Platonic Representations for Poverty Mapping: Unified Vision-Language Codes or Agent-Induced Novelty?
贫困地图绘制的柏拉图式表示:统一的视觉语言代码还是智能体诱导的新颖性?
Satiyabooshan Murugaboopathy, Connor T. Jerzak, Adel Daoud
机构
*
AI and Global Development Lab(人工智能与全球发展实验室)
;
Institute for Analytical Sociology(分析社会学研究所)
;
Institute of Computer Science(计算机科学研究所)
;
Department of Government(政府系)