SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents
SubtleMemory: 面向长时程AI智能体的细粒度关系记忆辨别基准
Wenxuan Wang, Haoyu Sun, Fukuan Hou, Mingyang Song, Weinan Zhang, Yu Cheng, Yang Yang
机构
*
Harbin Institute of Technology(哈尔滨工业大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Tongji University(同济大学)
;
Xiamen University(厦门大学)
;
Fudan University(复旦大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
The Chinese University of Hong Kong(香港中文大学)
MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval
MCERF:通过增强检索推进工程文档的多模态大语言模型评估
Kiarash Naghavi Khanghah, Hoang Anh Nguyen, Anna C. Doris, Amir Mohammad Vahedi, Daniele Grandi, Faez Ahmed, Hongyi Xu
机构
*
School of Mechanical, Aerospace, and Manufacturing Engineering, University of Connecticut, Storrs, CT 06269(机械、航空航天与制造工程学院,康涅狄格大学,斯托尔斯,CT 06269)
;
Department of Mechanical Engineering, Massachusetts Institute of Technology, Cambridge, MA 02139, USA(机械工程系,麻省理工学院,剑桥,MA 02139,美国)
机构
*
Hangzhou Institute for Advanced Study(杭州高等研究院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Computer Network Information Center(计算机网络信息中心)
;
Chinese Academy of Sciences(中国科学院)
When Contextual Inference Fails: Cancelability in Interactive Instruction Following
当情境推理失效时:交互指令跟随中的可取消性
Natalia Bila, Kata Naszádi, Alexandra Mayn, Christof Monz
机构
*
Language Technology Lab, University of Amsterdam(阿姆斯特丹大学语言技术实验室)
;
Department of Language Science and Technology, Saarland University(萨尔兰大学语言科学与技术系)
机构
*
Qwen Business Unit of Alibaba(阿里巴巴通义千问业务部)
;
ShanghaiTech University(上海科技大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
Institute of Computing Technology(中国科学院计算技术研究所)
;
Southeast University(东南大学)
CommentsThe authors are withdrawing this manuscript due to errors identified in the experimental evaluation and result aggregation, which affect several reported quantitative results and some conclusions. These issues require substantial re-evaluation of the experiments and analysis
Lost in Historical Time? A Polish History Matura Benchmark for Large Language Models
大语言模型通过了历史考试却遗漏了《历史》:波兰高中毕业考试Matura基准测试
Adrian Trzoss, Kacper Dudzic, Wiktor Werner, Marcin Moskalewicz
机构
*
Adam Mickiewicz University(亚当·密茨凯维奇大学)
;
IDEAS Research Institute(IDEAS研究院)
;
AMU Center for Artificial Intelligence(亚当·密茨凯维奇大学人工智能中心)
;
Poznań University of Medical Sciences(波兹南医科大学)
;
Maria Curie-Skłodowska University(玛丽·居里-斯克洛多夫斯卡大学)
;
WSB Merito University(WSB梅里托大学)
PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research
PhysMaster:构建一个自主的AI物理学家用于理论和计算物理学研究
Tingjia Miao, Wenkai Jin, Jinxin Tan, Muhua Zhang, Xianghe Pang, Zexi Liu, Yuwen Du, Tian Jin, Tu Guo, Zhengliang Zhang, Jingkun Liu, Yuelin Hu, Jiejun Zhang, Yunjie Huang, Yuhan Wang, Wenbo Li, Yinuo Gao, Shuo Chen, Rui Ye, Yuzhi Zhang, Linfeng Zhang, Kun Chen, Wei Wang, Weinan E, Siheng Chen
机构
*
School of Artificial Intelligence(人工智能学院)
;
Shanghai Jiao Tong University(上海交通大学)
;
School of Physics and Astronomy(物理天文学院)
;
State Key Laboratory of Dark Matter Physics(暗物质物理国家重点实验室)
;
Tsung-Dao Lee Institute(李政道研究所)
;
Zhiyuan College(智源学院)
;
Institute of Theoretical Physics(理论物理研究所)
;
Chinese Academy of Sciences(中国科学院)
;
DP Technology(DP技术)
MOSAIC: Unveiling the Moral, Social and Individual Dimensions of Large Language Models
MOSAIC:揭示大型语言模型的道德、社会和个体维度
Erica Coppolillo, Emilio Ferrara
机构
*
University of Southern California(南加州大学)
;
University of Calabria and ICAR-CNR(卡利亚里大学和ICAR-CNR)
;
Thomas Lord Department of Computer Science(托马斯·劳德计算机科学系)