CommentsPublished at AAAI/ACM AIES 2025. Presented at NeurIPS 2025 Workshop on LLM Evaluation and the International Monetary Fund's 12th Statistical Forum. GermanPartiesQA Benchmark under https://github.com/janbatzner/germanpartiesqa
Journal refProceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 8(1), 2025, pp. 330-342
Building Trustworthy AI for Materials Discovery: From Autonomous Laboratories to Z-scores
构建可信的人工智能用于材料发现:从自主实验室到Z分数
Benhour Amirian, Ashley S. Dale, Sergei Kalinin, Jason Hattrick-Simpers
机构
*
University of Toronto(多伦多大学)
;
University of Tennessee(田纳西大学)
;
Vector Institute for Artificial Intelligence(人工智能矢量研究所)
;
Schwartz Reisman Institute for Technology and Society(技术与社会斯瓦茨-雷曼研究所)
Afsah Sharaf Khan, Falong Fan, Doohwan DH Kim, Abdurrahman Alshareef, Dong Chen, Justin Kim, Ernest Carter, Bo Liu, Jerzy W. Rozenblit, Bernard Zeigler
Confident RAG: Enhancing the Performance of LLMs for Mathematics Question Answering through Multi-Embedding and Confidence Scoring
Confident RAG: 通过多嵌入和置信度评分提升LLM在数学问题回答中的性能
Shiting Chen, Zijian Zhao, Jinsong Chen
机构
*
Faculty of Education, The University of Hong Kong(香港大学教育学院)
;
Department of Civil and Environmental Engineering, The Hong Kong University of Science and Technology(香港科学与技术大学土木与环境工程系)
机构
*
Morgan Stanley(摩根士丹利)
;
Clemson University(克莱姆森大学)
;
Arizona State University(亚利桑那州立大学)
;
Washington University in St. Louis(圣路易斯华盛顿大学)
;
University of Arizona(亚利桑那大学)
;
University of Notre Dame(圣母大学)
Comments9 pages, 4 figures, 9 tables. Study on diagnostic prompting for multimodal LLM-based visual complexity assessment of Amazon search result pages