{
  "id": 10488303,
  "title": "Generate images and video with vLLM-Omni on SageMaker AI – Part 2",
  "url": "https://urgent.news/2026/09/28/generate-images-and-video-with-vllm-omni-on-sagemaker-ai-part-2",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-28T16:15:17.000Z",
  "source": {
    "name": "AWS Machine Learning",
    "slug": "aws-machine-learning",
    "url": "https://aws.amazon.com/blogs/machine-learning/generate-images-and-video-with-vllm-omni-on-sagemaker-ai-part-2/"
  },
  "original_language": "en",
  "account": "This article provides instructions on using the vLLM-Omni Deep Learning Container (DLC) on Amazon SageMaker AI to generate images and videos from text prompts. The process involves deploying two separate endpoints for image and video generation, respectively. The image generation uses the FLUX.2-klein-4B model, while the video generation utilizes the Wan2.1-VACE-1.3B model. The workflow starts by sending a text prompt to generate a still image, and then passes the generated image along with a motion prompt to the video generation endpoint. The resulting MP4 video is stored in Amazon Simple Storage Service (Amazon S3). The AWS vLLM-Omni DLC packages the necessary frameworks and dependencies for training and inference on AWS, and adds routing middleware for SageMaker AI. It extends vLLM to support processing and generation of text, audio, images, and video through OpenAI-compatible APIs. The article continues a series about AWS DLCs, building upon the previous parts that covered real-time speech generation. The code sample in the article demonstrates how to clone the provided repository, deploy the image and video endpoints, and send text prompts to generate images and videos. The image endpoint returns a base64-encoded PNG directly to the application, while the video endpoint returns an asynchronous output location. The generated video is then retrieved from Amazon S3 after completion. The solution emphasizes the benefits of separating image and video generation, allowing for customized instance types and inference options tailored to each model's workload. It also highlights the use of SageMaker AI's real-time endpoint for the image generation and asynchronous inference endpoint for the video generation.",
  "summary": "Deploy two generative media models from one AWS vLLM-Omni Deep Learning Container on Amazon SageMaker AI. Generate an image with FLUX.2-klein through real-time inference, then animate it into video with Wan2.1-VACE through asynchronous inference, and retrieve the MP4 from Amazon S3.",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "AWS Machine Learning",
        "title": "Build real-time voice applications with vLLM-Omni on SageMaker AI – Part 1",
        "url": "https://urgent.news/2026/09/28/build-real-time-voice-applications-with-vllm-omni-on-sagemaker-ai",
        "published": "2026-09-28T16:15:46.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}