{
  "id": 10488302,
  "title": "Build real-time voice applications with vLLM-Omni on SageMaker AI – Part 1",
  "url": "https://urgent.news/2026/09/28/build-real-time-voice-applications-with-vllm-omni-on-sagemaker-ai",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-28T16:15:46.000Z",
  "source": {
    "name": "AWS Machine Learning",
    "slug": "aws-machine-learning",
    "url": "https://aws.amazon.com/blogs/machine-learning/build-real-time-voice-applications-with-vllm-omni-on-sagemaker-ai-part-1/"
  },
  "original_language": "en",
  "account": "In today's digital age, voice applications are becoming increasingly prevalent, requiring real-time responses without lengthy pauses. To address this challenge, this tutorial demonstrates how to deploy a text-to-speech (TTS) model on Amazon SageMaker AI that can begin playing speech before it has finished generating the entire response. By leveraging the AWS vLLM-Omni Deep Learning Container (DLC), the tutorial utilizes Qwen3-TTS to stream text and audio over a persistent bidirectional connection, and showcases the process through a Gradio application. AWS Deep Learning Containers provide preconfigured Docker images with the necessary frameworks and dependencies for both training and inference on AWS. This article serves as Part 1 of a series focused on specialized DLCs, including vLLM-Omni, WhisperX, and llama.cpp, with an emphasis on streamed speech for real-time voice applications. Part 2 will explore the application of the vLLM-Omni DLC to image and video generation. The vLLM-Omni project extends vLLM's capabilities beyond text generation to encompass models that process or generate text, audio, images, and video. AWS vLLM-Omni DLC packages the tracked vLLM-Omni releases within AWS images and incorporates routing middleware for SageMaker AI. Utilizing SageMaker bidirectional streaming, the tutorial sends text and receives audio chunks over a persistent connection, enabling the seamless exchange of data between the client and the vLLM-Omni model. The tutorial begins by cloning a code sample repository that includes the deployment script, shared streaming transport, and Gradio client. SageMaker bidirectional streaming facilitates a full-duplex WebSocket connection over HTTP/2, allowing for seamless communication between the client, SageMaker Runtime endpoint, inference sidecar, and vLLM-Omni DLC. Audio events and 24 kHz PCM chunks are transmitted over the same connection, ensuring efficient data transfer. The vLLM-Omni v1.5 DLC adds bidirectional streaming capabilities and the necessary routing middleware for SageMaker AI. The sample utilizes the native WebSocket route, v1/audio/speech/stream, to establish communication between the components. Endpoint configuration takes advantage of SageMaker instance pools, which define a priority-ordered list of instance types compatible with the production variant. SageMaker automatically selects an instance from the pool, starting with ml.g6.xlarge, and falls back to other instance types if capacity is limited. It is crucial to ensure sufficient quota for each instance type included in the pool, as insufficient quota may result in endpoint provisioning issues. Costs may also vary depending on the selected instance types. The tutorial assumes the use of an AWS account with AWS Command Line Interface (AWS CLI) or SDK credentials configured, Python 3.12 or newer, SageMaker AI execution role, and appropriate permissions for creating and invoking SageMaker AI endpoints. Additionally, Boto3 version 1.40.0 or newer and the SageMaker Runtime HTTP/2 Python client version 0.4.0 are required. The walkthrough assumes a US East (N. Virginia) region and builds the DLC image URI and runtime endpoint using the AWS_REGION variable. By following these steps, users can successfully deploy a streaming speech endpoint capable of providing real-time TTS responses.",
  "summary": "Deploy a text-to-speech model on Amazon SageMaker AI with the AWS vLLM-Omni Deep Learning Container and stream generated speech over a persistent bidirectional connection. This Part 1 tutorial deploys Qwen3-TTS and streams speech through a Gradio application.",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "AWS Machine Learning",
        "title": "Generate images and video with vLLM-Omni on SageMaker AI – Part 2",
        "url": "https://urgent.news/2026/09/28/generate-images-and-video-with-vllm-omni-on-sagemaker-ai-part-2",
        "published": "2026-09-28T16:15:17.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}