Urgent.News

What's breaking now, across thousands of outlets.

AI

I told an AI "put me in the rain with a bee circling my head" — and heard it orbit my ears

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend Spatial Audio Sandbox — describe a scene, hear it in 3D Flutter + Rust HRTF engine · Gemma-3-1B on Render · ElevenLabs · head-tracked on Android + macOS What I Built Spatial (3D) Audio Sandbox — describe a place in plain words and the app builds it as a live binaural 3D scene on your headphones. "I'm in the rain and…

In a recent Hacktoberfest Weekend Challenge, a developer built a Spatial Audio Sandbox that transforms a textual description of a place into a live binaural 3D audio experience on headphones. The app takes a user's description, such as "I'm in the rain and a bee is circling my head," and builds it into a 3D scene that plays around the user's ears.

The pipeline involves Gemma, a model hosted on Render, which generates a screenplay JSON with sources, positions, movement verbs, and timing cues. The SceneCompiler, written in Dart, refines the hallucinated sources and repairs the JSON before ElevenLabs generates each audio clip. A Rust HRTF mixer renders the sounds in real-time, anchored by head tracking to maintain the scene as the user turns their head.

The demo showcases the entire process, with a radar displaying moving sources, stereo meters indicating ear energy, and a timeline outlining cue durations. The project emphasizes open innovation, with open-weight model weights, open-source code, and a keyless demo path.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Making voice-agent failures reproducible with Python

A voice agent can answer correctly and still make a conversation frustrating. It can wait too long before replying, keep speaking after an interruption, or confirm the wrong appointment time after a…

  • Python-based framework "voice-evals" created to evaluate voice-agent failures.
  • Provides deterministic scoring, per-stage latency budgets, and CI gating capabilities.

More from Monday 5 October →