Generate images and video with vLLM-Omni on SageMaker AI – Part 2
The article demonstrates how a single AWS vLLM‑Omni Deep Learning Container on Amazon SageMaker AI can host both the FLUX.2‑klein image generation model and the Wan2.1‑VACE video animation model, enabling a two‑step pipeline that first produces an image in real‑time and then animates it into an MP4 asynchronously. By deploying the container as a SageMaker endpoint, the image is returned immediately while the video job is queued, and the resulting MP4 is stored in Amazon S3 for later retrieval. This approach showcases how to multiplex multiple generative media models within one container, reducing deployment overhead and simplifying inference orchestration. The tradeoff highlighted is the need to balance real‑time image inference against the longer latency of asynchronous video rendering.
⚡ Key Takeaways
- The vLLM‑Omni container supports simultaneous deployment of FLUX.2‑klein and Wan2.1‑VACE, eliminating separate model servers.
- Real‑time inference is achieved by invoking the endpoint with the FLUX.2‑klein prompt, while Wan2.1‑VACE is triggered asynchronously via SageMaker’s batch transform.
- The resulting MP4 is automatically uploaded to a specified Amazon S3 bucket, allowing downstream services to consume the video without additional I/O steps.
- Integration requires configuring the SageMaker endpoint with the vLLM‑Omni image and setting the correct IAM role for S3 access; the endpoint can be called using the SageMaker SDK’s `predict` method for image and `transform
Want the full story? Read the original article.
Read on AWS ML Blog ↗