发表机构
Amazon.com, Inc.(亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出两阶段大语言模型流水线,先发现购买区分属性模式,再用带超并行解码的Qwen3-4B提取属性值,实现85%准确率且推理成本降低92%,构建产品知识库。
AI 中文摘要
客户依赖特定的产品属性来比较产品并做出购买决策,但电子商务目录杂乱且非结构化,这使得识别哪些属性最为关键并大规模提取它们变得困难。标准的属性值提取(AVE)系统对所有属性一视同仁,产生庞大且不一致的属性集合,未能反映消费者用于区分产品的因素。我们引入了一个两阶段的大语言模型(LLM)流水线:首先为每个产品类别发现一个紧凑的、按购买区分度排序的属性模式(schema),然后使用经过微调的紧凑型大语言模型(Qwen3-4B)结合超并行解码(HPD)从目录文本中提取属性值。该流水线实现了85%的提取准确率,与其蒸馏所基于的基础大语言模型相当,同时将推理成本相比基础大语言模型降低了92%,从而支持产品发现和目录增强的生产级应用。由此产生的类别级结构化表示有效地构成了自动构建的产品知识库,为不同产品类别提供一致、可比较的属性,这些属性可为下游知识密集型应用提供基础。
英文摘要
Customers rely on specific product attributes to compare products and make purchasing decisions, but e-commerce catalogs are messy and unstructured, making it difficult to identify which attributes matter most and extract them at scale. Standard Attribute Value Extraction (AVE) systems treat all attributes equally, producing large, inconsistent attribute sets that do not reflect the factors consumers use to differentiate products. We introduce a two-stage LLM pipeline that first discovers a compact, ranked schema of purchase-discriminative attributes for each product category, then extracts their values from catalog text using a fine-tuned compact LLM (Qwen3-4B) with Hyper-Parallel Decoding (HPD). This pipeline achieves 85% extraction accuracy, on par with the foundational LLM it was distilled from, while reducing inference costs by 92% over foundational LLMs, enabling production-scale use for product discovery and catalog enrichment. The resulting category-level structured representations effectively constitute automatically constructed product knowledge bases, providing consistent, comparable attributes across varied product categories that can ground downstream knowledge-intensive applications.
CommentsAccepted to 11th Workshop on Automated Knowledge Base Construction