arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12390cs.LGcs.CL

用于预测特征的长文本:基于可执行程序搜索的大语言模型引导分块特征工程

Long Text to Predictive Features: LLM-Guided Blockwise Feature Engineering via Executable Program Search

Ziming Dai, Dabiao Ma, Ziheng Guo, Jack Dong, Zimu Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出LLM-BlockFE框架,通过可执行程序搜索将长文本转为特征程序,避免在线推理调用LLM,在多数据集上提升了风控模型的AUC与KS指标。

中文摘要 AI 辅助

工业风控系统通常依赖结构化数据模型进行高效预测,但大量有价值信息仍嵌入在非结构化长文本中。通过手动特征工程提取这些信息耗时费力,而要求大语言模型(LLM)处理每一条实时输入可能无法满足实际部署需求。为应对这一挑战,本文提出LLM-BlockFE,一种LLM引导的离线特征构建框架,可将长文本转换为可执行特征程序,从而避免在线推理阶段的LLM调用。LLM-BlockFE通过逐步追加不可变代码块来构建特征程序,并利用下游模型评估候选特征。针对传统贪心搜索易陷入次优解的问题,本文方法引入基于深度校准信用分配的分块级回滚机制,并以交错方式推进多条独立搜索轨迹,通过共享每条轨迹探索方向的固定描述减少冗余探索。搜索完成后,生成的程序被冻结并部署,以提取结构化特征供下游预测模型使用。在两个公开数据集和两个私有数据集上,LLM-BlockFE在全数据集对比中,相较于各数据集最强基线实现了0.0069至0.0358的绝对AUC提升;在五个已部署的金融风控应用中进行的上线后监控显示,相较于现有手动设计策略,实现了0.02至1.56个百分点的绝对KS提升。

英文摘要

Industrial risk-control systems typically rely on structured-data models for efficient prediction, yet substantial valuable information remains embedded in unstructured long text. Extracting this information through manual feature engineering is labor-intensive, while requiring a large language model (LLM) to process every real-time input may not meet practical deployment requirements. To address this challenge, we propose LLM-BlockFE, an LLM-guided offline feature construction framework that converts long text into executable feature programs, thereby avoiding LLM calls during online inference. LLM-BlockFE constructs feature programs by incrementally appending immutable code blocks and evaluates candidate features using a downstream model. To address the tendency of conventional greedy search to become trapped in suboptimal solutions, our method introduces a block-level rollback mechanism based on depth-calibrated credit allocation and advances multiple independent search trajectories in an interleaved manner, reducing redundant exploration by sharing fixed descriptions of each trajectory's exploration direction. After the search, the resulting programs are frozen and deployed to extract structured features for downstream prediction models. Across two public and two private datasets, LLM-BlockFE achieves absolute AUC improvements of 0.0069 to 0.0358 over the strongest baseline on each dataset in the full-dataset comparison. Post-launch monitoring across five deployed financial risk-control applications shows absolute KS improvements of 0.02 to 1.56 percentage points over the existing manually designed strategy.

发表机构

  • City University of Hong Kong(香港城市大学)
  • Qfin Holdings, Inc.(Qfin控股公司)
  • Tianjin University(天津大学)
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

↑