发表机构
Blossom AI Labs(Blossom AI 实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出生成溯源基底,用于在合成语音数据中先确定数据来源再归因行为,并通过护理交接案例验证其必要性与局限性。
AI 中文摘要
将模型行为归因于合成训练数据,需要在估计每个训练项所导致的结果之前,先了解每个训练项是由什么产生的。波形-标签对并不能保留这一知识。我们提出了一种生成溯源基底,在该基底中,一个合成研究对象将源规格、生成内容、波形、目标、事实要求、质量信号、审查谱系和不可变清单身份绑定在一起。生产者与选择机制决定了证据意义;存储位置和变量名则不然。我们在一个私有的日本护理交接流程中审计了这一基底。一个包含113个资产的审查群体包含六个场景家族中的1.552小时合成语音;所有项目都链接了音频、转录文本、候选笔记和事实检查清单,但人工证据具有选择性和来源特异性。两个仅忠实的清单在场景种子方面不相交,并且不可变地版本化,而精确的上游归因仍因浮动的生成器别名、缺失的每片段TTS和代码戳记,以及未版本化的检查提示而受阻。我们认为,生成溯源对于行为归因是必要的,但并非充分条件:它定义了候选因果图和审计单元,而贡献性归因仍需要冻结的训练运行以及干预或影响证据。本文贡献了一个紧凑的溯源契约、一个审计协议和一个有界的合成数据归因案例研究;可能提供受控的研究访问,但我们不主张因果训练数据归因、临床有效性或不受限制的公开发布。
英文摘要
Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused. A waveform-label pair does not preserve this knowledge. We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and immutable manifest identity. Producer and selection mechanism determine evidentiary meaning; storage location and variable name do not. We audit this substrate in a private Japanese care-handoff pipeline. A 113-asset review population contains 1.552 hours of synthetic speech across six scenario families; all items have linked audio, transcripts, candidate notes, and fact checklists, but human evidence is selective and source-specific. Two faithful-only manifests are scenario-seed-disjoint and immutably versioned, while exact upstream attribution remains blocked by floating generator aliases, missing per-clip TTS and code stamps, and an unversioned checking prompt. We argue that generation provenance is necessary but not sufficient for behavior attribution: it defines the candidate causal graph and audit units, whereas contributive attribution still requires frozen training runs and intervention or influence evidence. The paper contributes a compact provenance contract, an audit protocol, and a bounded case study for synthetic-data attribution; controlled research access may be offered, but we do not claim causal training-data attribution, clinical validity, or unrestricted public release.
CommentsAccepted to the Third NeurIPS Workshop on Attributing Model Behavior at Scale: Data Attribution and Provenance. 4 pages, 0 figures, 1 table. An aggregate reproducibility package is available from the authors on request!