arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

International Conference on Learning Representations · 会议 · Machine Learning

2025-12-05 至 2025-12-05 共收录 2
2510.10062 2025-12-05 cs.CL

HUME: Measuring the Human-Model Performance Gap in Text Embedding Tasks

HUME:文本嵌入任务中人类-模型性能差距的测量

Adnan El Assadi, Isaac Chung, Roman Solomatin, Niklas Muennighoff, Kenneth Enevoldsen

机构 * Carleton University(卡尔顿大学) Zendesk(Zendesk公司) Stanford University(斯坦福大学) Aarhus University(阿arhus大学)

AI总结 HUME通过测量人类与模型在文本嵌入任务中的性能差距,揭示了模型与人类在不同语言资源下的表现差异,并提供了一个可扩展的评估框架。

Comments Submitted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17773 2025-12-05 cs.CV cs.AI cs.CL cs.LG

KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal Models

KiVA:儿童启发的视觉类比用于测试大多模态模型

Eunice Yiu, Maan Qraitem, Anisa Noor Majhi, Charlie Wong, Yutong Bai, Shiry Ginosar, Alison Gopnik, Kate Saenko

机构 * University of California, Berkeley(加州大学伯克利分校) Boston University(波士顿大学) Google DeepMind(谷歌DeepMind) Toyota Technological Institute at Chicago(芝加哥丰田技术研究所)

AI总结 KiVA通过4300个日常物体视觉转换测试大模型的类比推理能力,发现儿童和成人表现优于现有模型,尤其在复杂任务上存在显著差距。

Comments 10 pages. Project website: https://ey242.github.io/kiva.github.io/. Benchmark and code: https://github.com/ey242/KiVA

Journal ref The Thirteenth International Conference on Learning Representations (ICLR), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏