AI 中文总结
本文为 Croissant 数据集描述符补充了可评估的决策语义,通过可加性配置文件定义五运算符决策程序,实现基于描述符的自动门控,并以低开销验证了其与原生及 ODRL 策略的一致性。
AI 中文摘要
Croissant 是机器学习数据集的事实标准机器可读描述符:基于 this http URL 的 JSON-LD。自 1.1 版本起,它还承载了数据使用条件,并推荐使用 DUO 和 ODRL 来表示这些条件。但任何版本都没有规定这些条件如何被评估:没有决策程序,没有评估成本的界限,没有针对实现无法评估的条件时的结果,没有所检查内容的记录,也没有关于与调用方权限组合的说明。我们补充了这一半。一个可加性的配置文件允许数据集声明其允许的操作以及允许这些操作的条件,基于一个封闭的五种运算符集合,其决策程序被完整给出,因此一个门控仅根据描述符本身即可做出决策,并记录其所检查的内容。两个语料库对其进行了评估,且它们的证据被分开保存。三个描述符门控了一个真实的 nf-core 流水线,给出了部署结果:来自配置文件文档的决策与门控的原生描述符逐条记录匹配,剥离该层后仍是一个有效的 Croissant 文档,并且增加的代价是 11.7 微秒,相对于 119 微秒的决策时间。一个从配置文件语法生成的语料库提供了广度,覆盖了每个运算符、拒绝类别和一致性条款。在其有效案例中,552 条完整的决策记录在三个方面一致:原生描述符、配置文件术语,以及 usageInfo 中作为 ODRL 的相同策略。因此,载体本身不是贡献;评估语义才是。最后,调用方绑定和数据绑定策略的范围覆盖不重叠的状态空间,因此两者的允许集合互不包含。
英文摘要
Croissant is the de facto machine-readable descriptor for ML datasets: JSON-LD over schema.org. Since version 1.1 it also carries data use conditions, recommending DUO and ODRL for them. What no version specifies is how any of them is evaluated: no decision procedure, no bound on evaluation cost, no outcome for a condition an implementation cannot evaluate, no record of what was checked, and nothing on composition with caller-side authority. We supply that half. An additive profile lets a dataset declare the operations it admits and the conditions under which it admits them, over a closed set of five operators whose decision procedure is given in full, so a gate decides from the descriptor alone and records what it checked. Two corpora evaluate it and their evidence is kept apart. Three descriptors that gated a real nf-core pipeline give the deployment result: decisions from a profile document match the gate's native descriptor record for record, stripping the layer leaves a valid Croissant document, and the added cost is 11.7 $μ$s against a 119 $μ$s decision. A corpus generated from the profile's grammar gives the breadth, covering every operator, refusal class and conformance clause. Across its valid cases, 552 complete decision records agree three ways -- native descriptor, profile terms, and the same policy as ODRL in usageInfo. The carrier is therefore not the contribution; the evaluation semantics is. Finally, caller-bound and data-bound policies range over non-overlapping state spaces, so neither permit set contains the other.
Comments23 pages. Reference implementation and conformance corpus archived at doi:10.5281/zenodo.22018156 and doi:10.5281/zenodo.22016112