arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22149cs.CLcs.LG

超越拼接假设:基于语义量化的多模态合成数据评估统一框架

Beyond the Stitching Assumption: A Unified Framework for Multimodal Synthetic Data Evaluation via Semantic Quantization

  • Santa Clara University(圣克拉拉大学)
  • eBay Inc.(eBay公司)

机构由 AI 辅助整理,请以论文原文为准。

Yefeng Yuan, Zhan Shi, Liang Cheng, Yuhong Liu

AI总结:

针对多模态合成数据被分别评估且拼接破坏后指标仍高的问题,提出基于语义量化的投影评估器,用置换对照检测跨模态依赖,实验验证其有效性并支持显式跨模态评估。

AI中文摘要:

多模态合成数据集将结构化属性与自由文本相结合,但通常被分别评估。当表格-文本配对被打乱时,此类指标仍可保持较高水平。我们提出了一种用于表格-文本合成数据的基于投影的评估器。一个固定的句子编码器将文本映射为嵌入,k均值将其转换为簇状态,表格变量则表示为分类或分位数状态。真实与合成的列联表通过Jensen-Shannon散度(JSD)、归一化互信息(NMI)、条件JSD(cJSD)和联合状态熵进行比较。我们还报告了文本到属性(T2A)效用以及一个基于保留集校准的邻近标志率(PFR)作为表示层面的诊断指标。一个文本置换对照在保持两个边缘分布的同时破坏其配对。在Amazon Reviews、Kiva Loans和Employment Scam Aegean数据集上的实验表明,在该对照下,模态特定分数仍然较高。当真实投影依赖超过置换基线时,投影诊断能够检测到破坏,但对于弱或稀疏投影则信息量较少。一些有条件的LLM基线也表现出比相应真实数据投影更强的测量依赖。这些结果支持采用置换基线和覆盖报告的显式跨模态评估。

英文摘要:

Multimodal synthetic datasets combine structured attributes with free text, but are often evaluated separately. Such metrics can remain high after tabular--text pairings are disrupted. We present a projection-based evaluator for tabular--text synthetic data. A fixed sentence encoder maps text to embeddings, \(k\)-means converts them to cluster states, and tabular variables are represented as categorical or quantile-binned states. Real and synthetic contingency tables are compared using Jensen--Shannon divergence (JSD), normalized mutual information (NMI), conditional JSD (cJSD), and joint-state entropy. We also report text-to-attribute (T2A) utility and a holdout-calibrated proximity flag rate (PFR) as a representation-level diagnostic. A text-permutation control preserves both marginal distributions while disrupting their pairing. Experiments on Amazon Reviews, Kiva Loans, and the Employment Scam Aegean Dataset show that modality-specific scores remain high under this control. The projection diagnostics detect disruption when real projected dependence exceeds a permutation baseline, but are less informative for weak or sparse projections. Some conditioned LLM baselines also exhibit stronger measured dependence than the corresponding real-data projections. These results support explicit cross-modal evaluation with permutation baselines and coverage reporting.

补充信息

↑