BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law
BenGER:德国法律中基于归入的法律推理的LLM系统基准测试
Sebastian Nagl, Ann-Kristin Mayrhofer, Martin Heidebach, Aleyna Koçak, Anne Zettelmeier, Elly Breu, Angelina Greiner, Sofija Milijas, Matthias Grabmair
机构
*
Technical University of Munich (TUM)(慕尼黑技术大学)
;
Ludwig Maximilian University of Munich (LMU)(慕尼黑路德维希-马克西米利安大学)
;
University of Konstanz(康斯坦茨大学)
;
University of Saarbrücken(萨尔布吕肯大学)
FlowPipe: LLM-Enhanced Conditional Generative Flow Networks for Data Preparation Pipeline Construction
FlowPipe: 基于条件生成流网络的LLM增强数据准备流水线构建
Kunyu Ni, Lei Cao, Jie He, Xiaotong Zhang, Jianfeng Jin, Junyu Dong, Yanwei Yu
机构
*
Ocean University of China(中国海洋大学)
;
University of Arizona(亚利桑那大学)
;
University of Science and Technology Beijing(北京科技大学)
;
Northeastern University(东北大学)
BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems
BELLS-O:评估LLM监督系统的运营权衡
Leonhard Waibl, Felix Michalak, Hadrien Mariaccia
机构
*
University of Graz, Graz, Austria(格拉茨大学)
;
Supervised Program for Alignment Research (SPAR)(对齐研究监督计划 (SPAR))
;
Centre pour la Sécurité de l'IA (CeSIA), Paris, France(人工智能安全研究中心 (CeSIA),巴黎,法国)
Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection
我们是否仍然需要人在回路中?比较主动学习中用于敌意检测的人类与LLM标注
Ahmad Dawar Hakimi, Lea Hirlimann, Isabelle Augenstein, Hinrich Schütze
机构
*
Center for Information and Language Processing, LMU Munich(慕尼黑大学信息与语言处理中心)
;
Department of Computer Science, University of Copenhagen(哥本哈根大学计算机科学系)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation
风险感知的LLM智能体用于地理空间数据检索:设计与初步对抗性评估
Kyle Gao, Joel Cumming, Jonathan Li, Linlin Xu, David A. Clausi
机构
*
Dept. of Systems Design Engineering, University of Waterloo(滑铁卢大学系统设计工程系)
;
SkyWatch
;
Dept. of Geography and Environmental Management, University of Waterloo(滑铁卢大学地理与环境管理系)
;
Dept. of Geomatics Engineering, University of Calgary(卡尔加里大学测绘工程系)
CommentsAccepted for publication in the International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences (ISPRS Archives), ISPRS Congress 2026
LLM-Powered Personalized Glycemic Assessment in Type 2 Diabetes with Wearable Sensor Data
基于可穿戴传感器数据的2型糖尿病个性化血糖评估:LLM驱动方法
Yifan Gao, Yanmin Gong, Yun Shi, Yuanxiong Guo
机构
*
Department of Information Systems and Cybersecurity, The University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校信息系统与网络安全系)
;
School of Engineering Medicine, Texas A&M University(德克萨斯农工大学工程医学院)
;
Department of Family and Community Medicine, The University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校家庭与社区医学系)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
Automated Creativity Evaluation of Language Models Across Open-Ended Tasks
语言模型在开放式任务中的自动化创造力评估
Min Sen Tan, Zachary Kit Chun Choy, Syed Ali Redha Alsagoff, Nadya Yuki Wangsajaya, Mohor Banerjee, Swaagat Bikash Saikia, Alvin Chan
机构
*
Raffles Institution(莱佛士书院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
;
Lee Kong Chian School of Medicine, Nanyang Technological University(南洋理工大学李光前医学院)
;
Centre of AI in Medicine (C-AIM), Nanyang Technological University(南洋理工大学人工智能医学中心)
专题命中
评测与基准
:LLM(summary_cn,abstract);language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI