Bearings: Self-Supervised Soundfield Embeddings from First-Order Ambisonics
Recently proposed self-supervised audio encoders learn powerful general-purpose representations of sound scenes, yet they are spatially blind. To supply the missing spatial representation of sound scenes, we introduce Bearings. Bearings is a self-supervised framework that learns soundfield embeddings from unlabeled first-order Ambisonics. We pre-train a masked auto-encoder paired with a decoder…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.
