arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

边际响应面启发用于零标签表格学习

Marginal Response Surface Elicitation for Zero-Label Tabular Learning

Liangyu Teng, Yicheng Ding, Jing Liu, Hengsong Liu, Juncen Guo, Hongru Li, Jingyu Zhang, Liang Song

arXiv 2609.39639首次发表:更新:

发表机构

Fudan Institute on Networking Systems of AI(复旦大学人工智能网络系统研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MARS方法,利用LLM从特征级先验构建零样本表格分类器,在八个基准上取得最优平均AUC和AP,分别超越直接提示1.97和6.21个百分点,并降低端到端成本。

AI 中文摘要

表格学习利用结构化数据来预测目标结果。传统上,这一过程依赖于标注数据。然而,大型语言模型(LLMs)可基于任务描述和特征语义来启发领域先验,从而在没有标注数据的情况下实现预测。我们提出了边际响应面启发(MARS),一种将特征级LLM先验转化为可复用的零样本表格分类器的方法。为构建此分类器,MARS从未标注数据中为每个特征选择代表性值,并提示LLM提供相应的类别支持分数和特征权重。然后,它使用中位数聚合多个响应以构建特征响应函数,并通过其加权和进行预测,无需进一步的LLM查询。在八个表格基准任务中,MARS取得了最高的平均AUC和AP,分别比直接提示高出1.97和6.21个百分点,同时大幅降低了端到端成本。使用不同规模的LLM进行的评估进一步证明了其相对于直接提示的预测优势。

英文摘要

Tabular learning uses structured data to predict target outcomes. Traditionally, this process has relied on labeled data. However, large language models (LLMs) can be used to elicit domain priors based on the task description and feature semantics, thereby enabling predictions without labeled data. We propose Marginal Response Surface Elicitation (MARS), a method that transforms feature-level LLM priors into a reusable, zero-shot tabular classifier. To construct this classifier, MARS selects representative values for each feature from unlabeled data and prompts the LLM to provide corresponding class support scores and feature weights. It then aggregates multiple responses using the median to construct feature response functions, and makes predictions through their weighted sum without further LLM queries. Across eight tabular benchmark tasks, MARS achieves the highest average AUC and AP, outperforming direct prompting by 1.97 and 6.21 percentage points respectively, while substantially reducing end-to-end costs. Evaluations with LLMs of different sizes further demonstrate its predictive advantage over direct prompting.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑