← Back
AWS ML Blog

Generate images and video with vLLM-Omni on SageMaker AI – Part 2

•9 min read•
#amazon#inference
✦TL;DR

The article demonstrates how a single AWS vLLM‑Omni Deep Learning Container on Amazon SageMaker AI can host both the FLUX.2‑klein image generation model and the Wan2.1‑VACE video animation model, enabling a two‑step pipeline that first produces an image in real‑time and then animates it into an MP4 asynchronously. By deploying the container as a SageMaker endpoint, the image is returned immediately while the video job is queued, and the resulting MP4 is stored in Amazon S3 for later retrieval. This approach showcases how to multiplex multiple generative media models within one container, reducing deployment overhead and simplifying inference orchestration. The tradeoff highlighted is the need to balance real‑time image inference against the longer latency of asynchronous video rendering.

⚡ Key Takeaways

  • The vLLM‑Omni container supports simultaneous deployment of FLUX.2‑klein and Wan2.1‑VACE, eliminating separate model servers.
  • Real‑time inference is achieved by invoking the endpoint with the FLUX.2‑klein prompt, while Wan2.1‑VACE is triggered asynchronously via SageMaker’s batch transform.
  • The resulting MP4 is automatically uploaded to a specified Amazon S3 bucket, allowing downstream services to consume the video without additional I/O steps.
  • Integration requires configuring the SageMaker endpoint with the vLLM‑Omni image and setting the correct IAM role for S3 access; the endpoint can be called using the SageMaker SDK’s `predict` method for image and `transform

Want the full story? Read the original article.

Read on AWS ML Blog ↗

More like this

With a feel for physics, AI models simulate a wider range of real-world scenarios

MIT News AI•#llm

How Condé Nast built multimodal video discovery with Amazon Bedrock

AWS ML Blog•#bedrock

A decade of mathematical certainty: Reflections on the Automated Reasoning Group

Amazon Science•#inference

Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS

Hugging Face Blog•#inference

EXPLORE AI NEWS

Daily hand-picked stories on LLMs, RAG, agents and production AI — curated for engineers who ship.

BROWSE NEWS

GET THE WEEKLY DIGEST

Join engineers getting the Monday signal-over-noise AI breakdown. No spam, unsubscribe anytime.

LEARN AI ENGINEERING

Curated courses, research papers, repos and tutorials built for engineers leveling up in AI.

START LEARNING