LEGIT:可信AI智能体市场的认证协议
LEGIT: Credentialing Protocol for Trustworthy AI Agent Marketplaces
- University of Calgary(卡尔加里大学)
- University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对智能体市场买家难以验证智能体性能的问题,提出LEGIT认证协议,通过签名记录绑定质量与成本,结合声誉机制,实现可信评估与分配。
AI中文摘要:
智能体市场正在兴起,其中具有不同能力的AI智能体自主为买家完成专业任务。此类市场的一个主要挑战是,买家难以确定哪个智能体在其任务上表现最佳。报告的基准分数可能难以跨任务、软件和预算进行验证或比较。我们提出了LEGIT,一种连接认证、声誉和提议的市场分配的认证协议。认证通过签名记录将每个已解决任务的测量质量和成本绑定到智能体配置、任务领域、评估预算和证据上。声誉将过去任务结果的记录关联到同一身份,但受限于所报告反馈的可靠性。买家和智能体可以验证认证记录,并检查可选的视觉档案。评估揭示了具有相似观察任务成功率的智能体配置之间的成本差异,并表明比较结果取决于评估预算。这些结果支持将性能测量绑定到已测试的配置和资源限制。一项补充分析量化了在所述女巫攻击模型下,进行声誉操纵所需的存款和费用。
英文摘要:
Agentic marketplaces are emerging where AI agents with varying capabilities autonomously complete specialized tasks for buyers. A major challenge of such marketplaces is that buyers cannot easily determine which agent will perform best on their tasks. Reported benchmark scores may be difficult to verify or compare across tasks, software, and budgets. We introduce LEGIT, a credentialing protocol connecting certification, reputation, and proposed marketplace allocation. Certification binds measured quality and cost per solved task to an agent configuration, task domain, evaluation budget, and evidence through a signed record. Reputation links records of past task outcomes to the same identity, subject to the reliability of the reported feedback. Buyers and agents can verify credential records and inspect optional visual profiles. Evaluations reveal cost differences between agent configurations with similar observed task success, and show that comparisons depend on the evaluation budget. These results support binding performance measurements to the tested configuration and resource limits. A complementary analysis quantifies the deposits and fees required for reputation manipulation under a stated Sybil attack model.