arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30023cs.IRcs.CL

生成式引擎优化的需求侧测量:构建与验证百万角色意图标注的买家语料库

Demand-Side Measurement for Generative Engine Optimization: Constructing and Validating a Million-Persona, Intent-Annotated Buyer Corpus

Dmitrij Żatuchin, Daniil Dzemesjuk

首次发表
浏览论文内容

中文总结 AI 辅助

该研究构建并验证了含百万级角色的PersonaGen-1M买家语料库,带意图标注与来源偏好,旨在为生成式引擎优化的需求侧测量提供数据支撑,其分层子集已公开发布。

中文摘要 AI 辅助

ChatGPT、Gemini和Perplexity等生成式引擎会直接回答买家问题,并在答案中列出少量品牌。研究品牌如何进入或无法进入该短名单,需要需求侧数据:某品类买家的提问内容、所需信息以及信任的来源。现有大型角色语料库旨在训练数据多样性,且既没有阶段化搜索意图标签,也没有首选来源字段,因此无法与供给侧推荐测量结合。我们构建并验证了PersonaGen-1M,这是一个包含1031732个合成买家角色的语料库,涵盖511个行业标签和4种市场场景,包含19416821个结构化行为属性,其中5160046个是搜索查询。每个角色都带有一个primary_intent主意图标签,覆盖其查询集(78.3%为信息类,17.4%为商业类,4.3%为交易类),以及一个preferred_sources首选来源字段,标明该买家会信任的来源类型。该语料库基于从四个公共数据集提取的约4000万原始角色描述,通过GPU加速的MinHash LSH(局部敏感哈希)加语义去重构建,随后按固定模式丰富。意图字段选择其查询驱动推荐的商业评估角色,首选来源字段与引用来源数据配对;这种配对是主要预期用途,其受控实证估计属于未来工作。在2026年8月调查的百万级角色语料库中,另有一个带有来源偏好属性,为六值媒体渠道枚举;PersonaGen-1M将每个角色的命名来源列表与阶段化商业搜索意图标签及附加查询集配对。完整语料库可应要求用于非商业研究;分层子集已公开发布,以便无需请求即可检查和复用协议、模式及验证过程。

英文摘要

Generative engines such as ChatGPT, Gemini, and Perplexity answer buyer questions directly and name a shortlist of brands inside the answer. Studying how brands enter or fail to enter that shortlist requires demand-side data: what buyers in a category ask, what information they need, and which sources they trust. Existing large persona corpora are built for training-data diversity and carry neither a staged search-intent label nor a preferred-sources field, so they cannot be joined to supply-side recommendation measurements. We built and validated PersonaGen-1M, a corpus of 1,031,732 synthetic buyer personas spanning 511 industry labels and 4 market contexts, carrying 19,416,821 structured behavioral attributes, 5,160,046 of them search queries. Each persona carries a single primary_intent label covering its query set (78.3% informational, 17.4% commercial, 4.3% transactional) and a preferred_sources field naming the source types that buyer would trust. The corpus was built from roughly 40 million raw persona descriptions drawn from four public datasets through GPU-accelerated MinHash LSH plus semantic deduplication, then enriched to a fixed schema. The intent field selects the commercial-evaluation personas whose queries drive recommendation, and the preferred_sources field pairs against citation-provenance data; that join is the primary intended use, and its controlled empirical estimate is future work. Among million-scale persona corpora surveyed in August 2026, one other carries a source-preference attribute, as a six-value media-channel enum; PersonaGen-1M pairs named per-persona source lists with a staged commercial search-intent label and an attached query set. The full corpus is shared on request for non-commercial research; a stratified subset is published openly so the protocol, the schema and the validation can be inspected and reused without asking us.

发表机构

  • Estonian Entrepreneurship University of Applied Sciences (EUAS)(爱沙尼亚应用创业大学)
  • Rankfor.AI OÜ(Rankfor.AI公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑