Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects
Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused. A waveform-label pair does not preserve this knowledge. We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage,…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.