clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
Chalamalasetti Kranti, Sherzod Hakimov, David Schlangen
机构
*
Computational Linguistics, Department of Linguistics University of Potsdam(普腾多夫大学语言学系)
;
German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)
专题命中
评测与基准
:LLM(title);large language model(abstract);language model(abstract);prompting(abstract)
Beyond Single Models: Enhancing LLM Detection of Ambiguity in Requests through Debate
Ana Davila, Jacinto Colan, Yasuhisa Hasegawa
机构
*
Institutes of Innovation for Future Society(未来社会创新研究所)
;
Nagoya University(名古屋大学)
;
Department of Micro-Nano Mechanical Science and Engineering(微纳米机械科学与工程系)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL
CommentsAccepted at the 2025 SICE Festival with Annual Conference (SICE FES)
Journal ref2025 SICE Festival with Annual Conference (SICE FES)
LLM Agents for Bargaining with Utility-based Feedback
Jihwan Oh
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG
CommentsarXiv admin comment: This version has been removed by arXiv administrators as the submitter did not have the rights to agree to the license at the time of submission
机构
*
Mohamed bin Zayed University of Artificial Intelligence(莫扎伊德大学人工智能学院)
;
Monash University(莫纳什大学)
;
University of Copenhagen(哥本哈根大学)
;
The University of Melbourne(墨尔本大学)
;
LibrAI(LibrAI公司)
;
Institute of Foundation Models(基础模型研究所)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL
CheckEmbed: Effective Verification of LLM Solutions to Open-Ended Tasks
Maciej Besta, Lorenzo Paleari, Marcin Copik, Robert Gerstenberger, Ales Kubicek, Piotr Nyczyk, Patrick Iff, Eric Schreiber, Tanja Srindran, Tomasz Lehmann, Hubert Niewiadomski, Torsten Hoefler
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL
机构
*
Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)(数据科学与分析 thrust,香港科学与技术大学(广州))
;
Department of Industrial Engineering and Decision Analytics, The Hong Kong University of Science and Technology(工业工程与决策分析系,香港科学与技术大学)
;
MOVENSYS Inc.(MOVENSYS公司)
;
University of Cologne(科隆大学)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
CommentsIEEE CASE 2025 Best Student Paper Finalists
机构
*
Indian Institute of Technology, Kharagpur(印度理工学院,克哈格布尔分校)
;
Indian Statistical Institute, Kolkata(印度统计研究所,科契)
;
Accenture Labs, Bangalore(埃森哲实验室,班加罗尔)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL
CommentsAccepted to appear at the ACL 2025 findings
机构
*
China Telecom Research Institute(中国电信研究院)
;
The Conversational Artificial Intelligence (CoAI) Group, Tsinghua University(清华大学对话人工智能(CoAI)小组)
;
University of Electronic Science and Technology of China(电子科技大学)
;
The Knowledge Engineering Group (KEG), Tsinghua University(清华大学知识工程小组)
;
Zhipu AI(智谱AI)
专题命中
评测与基准
:language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL
Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores
Robert E. Blackwell, Jon Barry, Anthony G. Cohn
机构
*
The Alan Turing Institute(艾伦·图灵研究所)
;
The Centre for Environment Fisheries and Aquaculture Science(环境渔业与水产科学研究中心)
;
School of Computer Science, University of Leeds(利兹大学计算机科学学院)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL