PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation
Text-to-spatial audio generation, such as text-to-First-Order Ambisonics (FOA), provides a convenient way to create spatial audio for billion-dollar gaming and film industries. However, existing text-to-FOA methods are largely data-driven and may produce audio that violates acoustic relations between source direction and distance. They also separate descriptive and parametric control, forcing…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.