TREC 2025 百万大语言模型赛道概述
Overview of the TREC 2025 Million Large Language Models track
- University of Amsterdam(阿姆斯特丹大学)
- Carnegie Mellon University(卡内基梅隆大学)
- RMIT(皇家墨尔本理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该论文介绍TREC 2025百万LLM赛道,提出基于检索的专家选择范式,通过模型行为动态推断专长,并构建首个大规模专长检索基准。
AI中文摘要:
智能体人工智能设想了一个由智能体组成的生态系统,这些智能体以最少的人工干预协作解决复杂任务。在这样的生态系统中,每个智能体都拥有专业化的专长,因此有效的专家选择对整体系统性能至关重要。虽然当前大多数方法假设存在少量文档完善、记录详尽的模型,但现实世界中的专长远为多样,无法通过静态元数据或手工编写的描述充分捕捉。我们预见未来将出现数百万个专业化语言模型(LLM),每个模型在不同领域或问题类型中表现出色。我们不依赖预定义的能力声明,而是提出一种基于检索的范式,在该范式中,助手智能体通过检查模型的可观察行为来动态推断其专长。收到用户查询后,助手根据展示出的能力对候选LLM进行排序,从而实现高效且自适应的专家选择。TREC百万LLM赛道通过将检索目标从文档转向专家LLM来实践这一范式。参与者获得一个发现集,其中包含来自一千多个LLM的查询、答案和日志概率,并被挑战为每个模型推断有意义的专长表示。面对未见过的测试查询,系统必须根据预期性能对LLM进行排序,从而为智能体人工智能中的专长检索提供首个大规模基准。
英文摘要:
Agentic AI envisions ecosystems of intelligent agents collaboratively solving complex tasks with minimal human intervention. In such ecosystems, each agent possesses specialized expertise, making effective expert selection central to overall system performance. While most current approaches assume a small number of well-documented models, real-world expertise is far more diverse and cannot be adequately captured through static metadata or hand-written descriptions. We anticipate a future with millions of specialized language models (LLMs), each excelling in different domains or problem types. Rather than relying on predefined capability statements, we propose a retrieval-based paradigm in which an assistant agent infers expertise dynamically by examining models' observable behavior. Upon receiving a user query, the assistant ranks candidate LLMs based on demonstrated competence, enabling efficient and adaptive expert selection. The TREC Million LLM Track operationalizes this paradigm by shifting the retrieval target from documents to expert LLMs. Participants are given a discovery set consisting of queries, answers, and log-probabilities from more than one thousand LLMs and are challenged to infer meaningful expertise representations for each model. Given an unseen test query, systems must then rank the LLMs according to their expected performance, providing the first large-scale benchmark for expertise retrieval in agentic AI.