WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Comments Link: https://hf.co/spaces/allenai/WildBench
作者
Natural Language Processing
Comments Link: https://hf.co/spaces/allenai/WildBench
Comments ICLR 2024 version , 31 pages
Comments Accepted at Conference on Language Modeling (COLM), 2024
Comments AIES 2024, 34 pages, 4 figures, 23 tables
Comments COLM 2024 camera-ready, code available at https://github.com/alisawuffles/proxy-tuning
Comments ICML 2024
Comments NAACL 2024
Comments Accepted to ACL 2024 Main Conference; Camera Ready. Project website: https://allenai.github.io/lumos/
Comments 26 pages
Comments ICLR 2024 Camera Ready version. With respect to the original submission, we added text generation experiments, plots of entire accuracy distributions for each task + stdev computations, and prompt length correlation with spread analysis
Comments 2024 ICLR Spotlight. The dataset and code can be found at https://confaide.github.io
Comments Accepted as a long paper to ACL 2024 Main
Comments link: https://hf.co/spaces/WildVision/vision-arena
Comments Accepted to ACL Findings 2024
Comments Github: https://github.com/VIM-Bench/VIM_TOOL, Model and Data: https://huggingface.co/VIM-Bench
Comments 44 pages, 19 figures, 12 tables