发表机构
Thumbtack(Thumbtack)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出自动研究循环,每次生成一种职业的分类体系,已部署于美国大型消费者服务市场,通过迭代优化与多角色评判生成标签集,并将传统表单问答映射至分类体系,实现AI原生匹配。
AI 中文摘要
双边服务市场正从确定性的请求表单录入转向由大语言模型(LLMs)驱动的AI原生概率匹配,LLMs可从自然语言中推断用户意图、偏好及潜在约束。依赖推断出的意图而非固定表单字段,迫使这些平台重新生成支撑匹配、搜索和定价的供给方偏好分类体系——该体系需同时满足服务提供者可理解、且能为市场决策提供有效信号的要求。我们提出一种自动研究循环,每次生成一种职业的分类体系,该循环自2026年4月起已部署于美国某大型消费者服务市场的生产环境,覆盖132种职业。该循环不采用单一全局层级,而是将每种职业视为独立生成问题,运行迭代的“提议-评估-保留”优化周期。每个候选标签集由重新校准的六准则LLM评判框架评分,且由包含不同角色的7人评判小组对调整后分数施加加权惩罚,无硬性否决机制。另有一个 parity映射阶段,将传统请求表单的问答对映射回生成的分类体系,该阶段先推断每个传统问题旨在测量的供给方属性,而非逐字将问题翻译为标签,最终同时得到覆盖度信号和人工质量保障接口。
英文摘要
Two-sided service marketplaces are moving from deterministic request-form intake to AI-native probabilistic matching, enabled by large language models (LLMs) that infer intent, preferences, and latent constraints from natural language. Relying on inferred intent rather than fixed-form fields forces these platforms to regenerate the provider-side preference taxonomy underwriting matching, search, and pricing: attributes interpretable to service providers while remaining a useful signal for marketplace decisions. We present an autoresearch loop that generates this taxonomy, one occupation at a time, and has been deployed in production at a major U.S. consumer services marketplace since April 2026, spanning 132 occupations. Instead of one global hierarchy, the loop treats each occupation as an independent generation problem and runs iterative propose-evaluate-keep refinement cycles. Each candidate tag set is scored by a recalibrated six-rubric LLM-as-judge framework, and a 7-critic panel of distinct personas contributes weighted penalties to an adjusted score, with no hard vetoes. A separate parity-mapping stage maps legacy request-form Q&A pairs back to the generated taxonomy, yielding both a coverage signal and an interface for human quality assurance; it does so by first inferring the provider attribute each legacy question was meant to measure, rather than translating questions to tags literally.