Urgent.News

What's breaking now, across thousands of outlets.

AI

Generate images and video with vLLM-Omni on SageMaker AI – Part 2

Deploy two generative media models from one AWS vLLM-Omni Deep Learning Container on Amazon SageMaker AI. Generate an image with FLUX.2-klein through real-time inference, then animate it into video with Wan2.1-VACE through asynchronous inference, and retrieve the MP4 from Amazon S3.

This article provides instructions on using the vLLM-Omni Deep Learning Container (DLC) on Amazon SageMaker AI to generate images and videos from text prompts. The process involves deploying two separate endpoints for image and video generation, respectively. The image generation uses the FLUX.2-klein-4B model, while the video generation utilizes the Wan2.1-VACE-1.3B model.

The workflow starts by sending a text prompt to generate a still image, and then passes the generated image along with a motion prompt to the video generation endpoint. The resulting MP4 video is stored in Amazon Simple Storage Service (Amazon S3). The AWS vLLM-Omni DLC packages the necessary frameworks and dependencies for training and inference on AWS, and adds routing middleware for SageMaker AI.

It extends vLLM to support processing and generation of text, audio, images, and video through OpenAI-compatible APIs. The article continues a series about AWS DLCs, building upon the previous parts that covered real-time speech generation. The code sample in the article demonstrates how to clone the provided repository, deploy the image and video endpoints, and send text prompts to generate images and videos.

The image endpoint returns a base64-encoded PNG directly to the application, while the video endpoint returns an asynchronous output location. The generated video is then retrieved from Amazon S3 after completion. The solution emphasizes the benefits of separating image and video generation, allowing for customized instance types and inference options tailored to each model's workload.

It also highlights the use of SageMaker AI's real-time endpoint for the image generation and asynchronous inference endpoint for the video generation.

Written by urgent.news from AWS Machine Learning's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at aws.amazon.com →

More in AI

How I connected an AI agent to GitHub with Nango and MCP (without touching a single OAuth token) published: false tags: ai, mcp, python, tutorial

Every time you connect an AI agent to an external API, you inherit the boring, risky part: OAuth flows, token storage, token refresh, and making sure nothing leaks.

  • AI agent accesses GitHub via MCP server without OAuth tokens
  • Nango handles OAuth and token management for GitHub connection
  • Python MCP server provides listmyrepos, listopenissues, and createissue tools

Building IncidentMind: An AI Incident Investigation Assistant That Learns From the Past

Building IncidentMind: An AI Incident Investigation Assistant That Learns From the Past Incident response is rarely difficult because engineers lack the ability to investigate a problem.

  • IncidentMind AI assistant learns from past incidents to aid future investigations
  • Structured investigation reports include evidence, root causes, and recommended actions
  • Confirmed root cause of authentication API issue was signing-key version inconsistency

More from Monday 28 September →