Urgent.News

What's breaking now, across thousands of outlets.

AI

Build real-time voice applications with vLLM-Omni on SageMaker AI – Part 1

Deploy a text-to-speech model on Amazon SageMaker AI with the AWS vLLM-Omni Deep Learning Container and stream generated speech over a persistent bidirectional connection. This Part 1 tutorial deploys Qwen3-TTS and streams speech through a Gradio application.

In today's digital age, voice applications are becoming increasingly prevalent, requiring real-time responses without lengthy pauses. To address this challenge, this tutorial demonstrates how to deploy a text-to-speech (TTS) model on Amazon SageMaker AI that can begin playing speech before it has finished generating the entire response.

By leveraging the AWS vLLM-Omni Deep Learning Container (DLC), the tutorial utilizes Qwen3-TTS to stream text and audio over a persistent bidirectional connection, and showcases the process through a Gradio application. AWS Deep Learning Containers provide preconfigured Docker images with the necessary frameworks and dependencies for both training and inference on AWS.

This article serves as Part 1 of a series focused on specialized DLCs, including vLLM-Omni, WhisperX, and llama.cpp, with an emphasis on streamed speech for real-time voice applications. Part 2 will explore the application of the vLLM-Omni DLC to image and video generation. The vLLM-Omni project extends vLLM's capabilities beyond text generation to encompass models that process or generate text, audio, images, and video.

AWS vLLM-Omni DLC packages the tracked vLLM-Omni releases within AWS images and incorporates routing middleware for SageMaker AI. Utilizing SageMaker bidirectional streaming, the tutorial sends text and receives audio chunks over a persistent connection, enabling the seamless exchange of data between the client and the vLLM-Omni model.

The tutorial begins by cloning a code sample repository that includes the deployment script, shared streaming transport, and Gradio client. SageMaker bidirectional streaming facilitates a full-duplex WebSocket connection over HTTP/2, allowing for seamless communication between the client, SageMaker Runtime endpoint, inference sidecar, and vLLM-Omni DLC.

Audio events and 24 kHz PCM chunks are transmitted over the same connection, ensuring efficient data transfer. The vLLM-Omni v1.5 DLC adds bidirectional streaming capabilities and the necessary routing middleware for SageMaker AI. The sample utilizes the native WebSocket route, v1/audio/speech/stream, to establish communication between the components.

Endpoint configuration takes advantage of SageMaker instance pools, which define a priority-ordered list of instance types compatible with the production variant. SageMaker automatically selects an instance from the pool, starting with ml.g6.xlarge, and falls back to other instance types if capacity is limited. It is crucial to ensure sufficient quota for each instance type included in the pool, as insufficient quota may result in endpoint provisioning issues.

Costs may also vary depending on the selected instance types. The tutorial assumes the use of an AWS account with AWS Command Line Interface (AWS CLI) or SDK credentials configured, Python 3.12 or newer, SageMaker AI execution role, and appropriate permissions for creating and invoking SageMaker AI endpoints. Additionally, Boto3 version 1.40.0 or newer and the SageMaker Runtime HTTP/2 Python client version 0.4.0 are required.

The walkthrough assumes a US East (N. Virginia) region and builds the DLC image URI and runtime endpoint using the AWS_REGION variable. By following these steps, users can successfully deploy a streaming speech endpoint capable of providing real-time TTS responses.

Written by urgent.news from AWS Machine Learning's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at aws.amazon.com →

More in AI

How I connected an AI agent to GitHub with Nango and MCP (without touching a single OAuth token) published: false tags: ai, mcp, python, tutorial

Every time you connect an AI agent to an external API, you inherit the boring, risky part: OAuth flows, token storage, token refresh, and making sure nothing leaks.

  • AI agent accesses GitHub via MCP server without OAuth tokens
  • Nango handles OAuth and token management for GitHub connection
  • Python MCP server provides listmyrepos, listopenissues, and createissue tools

Building IncidentMind: An AI Incident Investigation Assistant That Learns From the Past

Building IncidentMind: An AI Incident Investigation Assistant That Learns From the Past Incident response is rarely difficult because engineers lack the ability to investigate a problem.

  • IncidentMind AI assistant learns from past incidents to aid future investigations
  • Structured investigation reports include evidence, root causes, and recommended actions
  • Confirmed root cause of authentication API issue was signing-key version inconsistency

More from Monday 28 September →