arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

意大利面架构师:一种抗污染、自带标签、多语言代码数据集生成器

Spaghetti Architect: A Contamination-Resistant, By-Construction-Labelled, Multi-Language Code Dataset Generator

Yuxiang Ji

arXiv 2607.18642首次发表:更新:

发表机构

Sunway College(双威学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对挖掘代码语料库的问题,提出意大利面架构师工具,通过反优化转译器生成多语言代码数据集,能抵抗污染、自带标签并沿难度轴标记,给出结构效度证据及基线,表明其在代码数据集生成上的有效性和优势。

AI 中文摘要

挖掘出的代码语料库丰富但无控制:代码片段的语义、表面“杂乱程度”和难度都是随机的;没有已知的最优参考标准;任何公共样本可能已在模型训练集中。我们提出了意大利面架构师,一种能生成有控制的代码数据集的工具。一个反优化转译器将干净的、语言无关的JSON中间表示映射为五种语言(Python、JavaScript、Go、Java、C++)中故意冗余、完全扁平化的程序;每个程序都经过编译、运行并与参考预言机核对,所以每个实例在构建时就是正确的。干净的IR是已知的最优参考,通过严格嵌套的反模式配置文件调节杂乱程度,每个实例沿着两个正交的难度轴(内在的问题规模和偶然的固定语义下的呈现方式)进行标记,通过从私有的保留种子生成新变体来抵抗污染。我们给出了结构效度证据,表明质量顺序符合既定的复杂度和可读性指标,并报告了一个四模型开放阶梯的基线:精确匹配随规模上升,内在旋钮使即使最强模型的算术聚合准确率降至零。此外,开发集分数与新生成的保留对应物在$|\Delta|\le 0.012$(理解)和$\le 0.011$(重构)内相等;在相同程序上,重构等效性($0.73 \rightarrow 0.99$)是规模不变的而输出预测会崩溃;去除生成器的自我注释表明它们使最弱模型提升的幅度比最强模型大一个数量级($-0.173$对$-0.017$):带注释的阶梯解决了三个相邻对中的一个,而无注释的解决了所有三个。开源(MIT),无依赖,通过持久的DOI存档。

英文摘要

Mined code corpora are abundant but uncontrolled: a snippet's semantics, surface "messiness," and difficulty are whatever the wild contained; there is no known-optimal reference to grade against; and any public sample may already sit in a model's training set. We present Spaghetti Architect, a tool that mints code datasets with the control such corpora lack. An anti-optimization transpiler maps a clean, language-agnostic JSON intermediate representation to deliberately redundant, fully-flattened programs in five languages (Python, JavaScript, Go, Java, C++); every program is compiled, run, and checked against a reference oracle, so each instance is correct by construction. The clean IR is a known-optimal reference, messiness is dialed by strictly-nested anti-pattern profiles, each instance is labelled along two orthogonal difficulty axes, intrinsic (problem size) and incidental (presentation at fixed semantics), and contamination is resisted by minting fresh variants from a private held-out seed. We give construct-validity evidence that the quality order moves established complexity and readability metrics, and report baselines on a four-model open ladder: exact match rises with scale, and the intrinsic knob collapses arithmetic-aggregation accuracy of even the strongest model to zero. Further, development-set scores equal freshly re-minted held-out counterparts within $|Δ|\le 0.012$ (comprehension) and $\le 0.011$ (refactoring); on identical programs, refactoring equivalence ($0.73 \rightarrow 0.99$) is scale-invariant while output prediction collapses; and ablating the generator's self-annotations shows they inflate the weakest model an order of magnitude more than the strongest ($-0.173$ vs $-0.017$): the annotated ladder resolves one of three adjacent pairs where the unannotated resolves all three. Open source (MIT), dependency-free, archived under a persistent DOI.

CommentsCode: https://github.com/KurathSec/Spaghetti-Architect (artifact archived at doi:10.5281/zenodo.21033174)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑