MDGYM: Benchmarking AI Agents on Molecular Simulations
MDGYM:在分子模拟上评估AI代理的基准测试
Vinay Kumar, Satyendra Rajput, Mausam, N. M. Anoop Krishnan
机构
*
Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi(印度理工学院德里人工智能学院)
;
Department of Computer Science and Engineering, Indian Institute of Technology Delhi(印度理工学院德里计算机科学与工程系)
;
Department of Civil and Environmental Engineering, Indian Institute of Technology Delhi(印度理工学院德里土木与环境工程系)
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs
Forge:面向NP难优化问题的质量感知强化学习
Xiaozhe Li, Xinyu Fang, Shengyuan Ding, Yang Li, Linyang Li, Haodong Duan, Qingwen Liu, Kai Chen
机构
*
Tongji University(同济大学)
;
Shanghai AI Lab(上海人工智能实验室)
;
Zhejiang University(浙江大学)
;
Fudan University(复旦大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Independent(独立)
Human-LLM Dialogue Improves Diagnostic Accuracy in Emergency Care
人类与大语言模型对话提高急诊护理的诊断准确性
Burcu Sayin, Ngoc Vo Hong, Ipek Baris Schlicht, Jacopo Staiano, Pasquale Minervini, Sara Allievi, Nicola Susca, Nicola Osti, Alberto Maino, Vito Racanelli, Andrea Passerini
机构
*
Department of Information Engineering and Computer Science, University of Trento(特伦托大学信息工程与计算机科学系)
;
Department of Medicine, Azienda Sanitaria Universitaria Integrata del Trentino (ASUIT)(特伦托大学整合卫生机构医学部)
;
Center for Medical Sciences (CISMed), University of Trento(特伦托大学医学科学中心)
;
Universitat Politècnica de València(瓦伦西亚理工大学)
;
The University of Edinburgh(爱丁堡大学)
机构
*
University of Notre Dame(诺丁汉大学)
;
LMU Munich(慕尼黑大学)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
;
University of Pennsylvania(宾夕法尼亚大学)
;
Lehigh University(莱恩大学)
;
Massachusetts Institute of Technology(麻省理工学院)
;
Bake AI UC Santa Barbara(加州大学圣芭芭拉分校)
;
University of Washington(华盛顿大学)
SpectraLLM: Uncovering the Ability of LLMs for Molecular Structure Elucidation from Multi-Spectral Data
SpectraLLM:从多光谱数据中揭示LLMs对分子结构解析的能力
Yunyue Su, Jiahui Chen, Zao Jiang, Zhenyi Zhong, Liang Wang, Qiang Liu, Zhaoxiang Zhang
机构
*
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences (CASIA)(模式识别新技术实验室,自动化研究所,中国科学院)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Tianjin University Peiyangyuan Campus(天津大学-pieceyang Campus)
Document Retrieval Augmented Fine-Tuning (DRAFT) for safety-critical software assessments
用于安全关键软件评估的文档检索增强微调(DRAFT)
Regan Bolton, Mohammadreza Sheikhfathollahi, Simon Parkinson, Vanessa Vulovic, Gary Bamford, Dan Basher, Howard Parkinson
机构
*
Digital Transit Limited, 3M Buckley Innovation Centre, UK, HD1 3BD(数字交通有限公司,3M Buckley创新中心,英国,HD1 3BD)
;
Department of Computer Science, University of Huddersfield, UK, HD1 3DH(计算机科学系,赫德瑟菲尔德大学,英国,HD1 3DH)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
SWE Atlas:超越问题解决的编码代理基准测试
Mohit Raghavendra, Soham Dan, Miguel Romero Calvo, Yannis Yiming He, Johannes Baptist Mols, Gautam Anand, Cole McCollum, Edgar Arakelyan, Vijay Bharadwaj, Andrew Park, Jeff Da, MohammadHossein Rezaei, Bing Liu, Brad Kenstler, Yunzhong He
专题命中
推理评测
:reasoning(abstract);分类 cs.LG
AI总结
SWE Atlas通过涵盖代码库问答、测试编写和重构三个软件工程流程,提供了一种新的基准测试套件,评估编码代理的正确性和工程质量。
LithoBench: Benchmarking Large Multimodal Models for Remote-Sensing Lithology Interpretation
LithoBench:用于遥感岩石学解释的大型多模态模型基准测试
Jun Wang, Fengpeng Li, Hang Dong, Tianjin Huang, Wei Han
机构
*
School of Computer Science, China University of Geosciences(中国地质大学(北京)计算机学院)
;
PRADA Lab, King Abdullah University of Science and Technology(科廷大学PRADA实验室)
;
Department of Computer Science, University of Exeter(埃克塞特大学计算机科学系)
PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection
PRPO:用于视觉-语言深度伪造检测的段级策略优化
Tuan Nguyen, Naseem Khan, Khang Tran, NhatHai Phan, Issa Khalil
机构
*
Qatar Computing Research Institute, Hamad Bin Khalifa University, Doha, Qatar(卡塔尔计算研究所,哈马德·本·哈利法大学,多哈)
;
New Jersey Institute of Technology, NJ, USA(新泽西理工学院)
Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study
基于LLM的软件工程社区心理安全定性编码的提示工程策略:一项受控实证研究
Moaath Alshaikh, Tasneem Alshaher, Ricardo Vieira, Beatriz Santana, Clelio Xavier, Jose Amancio, Glauco Carneiro, Julio Leite, Savio Freire, Manoel Mendonca
机构
*
Federal University of Bahia(巴伊亚联邦大学)
;
State University of Feira de Santana(费拉德桑塔纳州立大学)
;
Federal University of Sergipe(塞格皮联邦大学)
;
Federal Institute of Ceara(塞阿拉联邦理工学院)
Comments9 pages, 5 figures. Accepted at the 1st International Workshop on Prompt Engineering for Software Engineering (PROMPT-SE 2026), co-located with the 30th International Conference on Evaluation and Assessment in Software Engineering (EASE 2026), Glasgow, Scotland, United Kingdom, June 9--12, 2026
CommentsThe authors have identified an issue in the evaluation protocol in Section 5.1.3. Feature extraction and semantic matching used to compute P-EHR require correction and re-validation, as they may not have been applied consistently across all generated explanations and baselines. This may affect part of the reported quantitative results and analysis, so the authors withdraw this version