Home›Amazon

Amazon

14 curated articles on Amazon for AI engineers

14 articles
How Condé Nast built multimodal video discovery with Amazon Bedrock
AWS ML Blog· 10 min read· Today
How Condé Nast built multimodal video discovery with Amazon Bedrock

Condé Nast’s editorial teams cut video‑search time from an average of 250 minutes per task on a 140,000‑video library by building a multimodal discovery pipeline on Amazon Bedrock and Amazon OpenSearch. The system ingests video and text, uses Bedrock’s multimodal foundation models to generate embeddings, and stores them in OpenSearch for low‑latency retrieval. The solution demonstrates how Bedrock can be leveraged for production‑ready multimodal search, trading off higher compute cost for a dramatic productivity boost.

A decade of mathematical certainty: Reflections on the Automated Reasoning Group
Amazon Science· 8 min read· Aug 11, 2026
A decade of mathematical certainty: Reflections on the Automated Reasoning Group

The Automated Reasoning Group (ARG) at Amazon has made significant progress over the past decade in applying mathematical logic and formal verification techniques to prove the correctness and security of AWS systems. The group's production services now process billions of queries daily, and their work has led to the development of tools such as Tiros, Zelkova, and Lean, which are used to analyze network security, policies, and cryptographic protocols. The use of automated reasoning and proof assistants has enabled the group to prove the correctness of complex systems, including the Nitro Confidentiality Engine and the AWS policy interpreter. This work has had a significant impact on the security and reliability of AWS systems, and its practical implications for engineers building AI systems include the potential to apply similar techniques to ensure the correctness and security of AI mod

AWS Trainium Frontier competition: Co-design models and kernels on purpose-built AI chips
Amazon Science· 6 min read· Aug 10, 2026
AWS Trainium Frontier competition: Co-design models and kernels on purpose-built AI chips

The AWS Trainium Frontier competition invites researchers to co-design models and kernels on purpose-built AI chips, exploring the efficient frontier of model architectures on AWS Trainium. The competition provides a ~50M parameter baseline language model and challenges participants to modify everything, including architecture, optimizer, training loop, and custom NKI kernels, to optimize the model architecture and training throughput within a fixed time and compute budget. The goal is to find the balance between model capacity and training throughput, driving validation bits-per-byte low and downstream in-context learning capability high. The optimal architectures will differ from those designed for existing accelerators due to Trainium's unique hardware features. The practical implication for engineers building AI systems is the opportunity to discover new model architectures designed

Introducing Claude Sonnet 5.5 on AWS
AWS ML Blog· 5 min read· Yesterday
Introducing Claude Sonnet 5.5 on AWS

Claude Sonnet 5.5 is now available

Build real-time voice applications with vLLM-Omni on SageMaker AI – Part 1
AWS ML Blog· 8 min read· Yesterday
Build real-time voice applications with vLLM-Omni on SageMaker AI – Part 1

The tutorial demonstrates how to deploy the Qwen3‑TTS text‑to‑speech model on Amazon SageMaker using the AWS vLLM‑Omni deep‑learning container, then stream the generated audio in real time via a Gradio web interface. It shows how to launch a SageMaker endpoint that keeps a persistent bidirectional connection open, enabling low‑latency, continuous speech output directly to the browser. The approach leverages vLLM‑Omni’s efficient inference engine to keep GPU usage minimal while maintaining high throughput for streaming audio. The result is a production‑ready pipeline that can be integrated into voice‑enabled applications without needing custom inference code.

Generate images and video with vLLM-Omni on SageMaker AI – Part 2
AWS ML Blog· 9 min read· Yesterday
Generate images and video with vLLM-Omni on SageMaker AI – Part 2

The article demonstrates how a single AWS vLLM‑Omni Deep Learning Container on Amazon SageMaker AI can host both the FLUX.2‑klein image generation model and the Wan2.1‑VACE video animation model, enabling a two‑step pipeline that first produces an image in real‑time and then animates it into an MP4 asynchronously. By deploying the container as a SageMaker endpoint, the image is returned immediately while the video job is queued, and the resulting MP4 is stored in Amazon S3 for later retrieval. This approach showcases how to multiplex multiple generative media models within one container, reducing deployment overhead and simplifying inference orchestration. The tradeoff highlighted is the need to balance real‑time image inference against the longer latency of asynchronous video rendering.

Implementing synthetic monitoring using Amazon Nova Act
AWS ML Blog· 12 min read· Yesterday
Implementing synthetic monitoring using Amazon Nova Act

The post introduces an agent‑driven synthetic monitoring framework that leverages Amazon Nova Act together with Amazon Bedrock AgentCore to validate end‑to‑end user journeys. It replaces fragile UI scripts with resilient, managed validation logic, and includes a complete sample implementation that demonstrates how to orchestrate agents for continuous monitoring. The architecture emphasizes modular agent interactions, allowing each step of a user flow to be independently verified and retried. The approach offers a scalable, AI‑powered alternative to traditional scripted tests, though it requires integration of the Nova Act agent runtime and Bedrock AgentCore services.

Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod
AWS ML Blog· 19 min read· 4 days ago
Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod

SkyRL, an open‑source RL framework, has been integrated with Amazon SageMaker HyperPod to accelerate multimodal training of the Qwen3‑VL‑8B vision‑language model using the GRPO algorithm. The authors demonstrate how to containerize SkyRL, spin up a Ray cluster directly from SageMaker Studio, submit a training job, and monitor progress in real time. By leveraging HyperPod’s 8‑node GPU topology, they achieve a 2‑to‑3× speed‑up over a single‑node baseline while maintaining the same training accuracy. The approach showcases a practical workflow for scaling vision‑language RL workloads on managed cloud infrastructure.

NarrateAI: production-ready LLM quality assurance on Amazon Bedrock
AWS ML Blog· 24 min read· 4 days ago
NarrateAI: production-ready LLM quality assurance on Amazon Bedrock

NarrateAI introduces a production‑ready quality‑assurance framework for Amazon Bedrock LLMs, combining adaptive pipeline orchestration, cross‑account multi‑model failover, real‑time streaming evaluation, composite evaluation, and data‑accuracy verification to achieve roughly 99 % numerical accuracy on generated content. The system leverages Bedrock’s multi‑model capabilities and cross‑account IAM roles to automatically switch models when quality thresholds are breached, while streaming evaluation provides immediate feedback on output fidelity. The trade‑off is a modest increase in latency during streaming assessment, but the framework dramatically reduces manual QA overhead and ensures consistent output quality in production deployments.

Deploying real-time personalized speech with Qwen3-TTS on Amazon SageMaker AI
AWS ML Blog· 12 min read· 4 days ago
Deploying real-time personalized speech with Qwen3-TTS on Amazon SageMaker AI

The article demonstrates how to deploy the publicly available Qwen3‑TTS‑12Hz‑1.7B‑Base TTS model from Amazon SageMaker JumpStart as a fully managed, real‑time inference endpoint. It shows that a short reference clip can be used to clone a speaker’s voice, and that the cloned voice retains identity across multiple languages, enabling cross‑lingual TTS

Multi-Region training with Amazon SageMaker HyperPod and Qumulo
AWS ML Blog· 18 min read· 4 days ago
Multi-Region training with Amazon SageMaker HyperPod and Qumulo

Amazon SageMaker HyperPod can offload compute to a remote AWS Region while keeping the training dataset in a different region, and when combined with Qumulo’s cloud‑native file system the remote cluster achieved throughput parity with a co‑located cluster. The architecture relies on HyperPod’s high‑bandwidth interconnect and Qumulo’s low‑latency NFS access, enabling a cross‑region training pipeline that incurs minimal network overhead. This demonstrates that large‑scale, distributed training can span regions without

Speaker-labeled transcription with WhisperX on SageMaker AI
AWS ML Blog· 13 min read· 5 days ago
Speaker-labeled transcription with WhisperX on SageMaker AI

AWS has released a WhisperX Deep Learning Container that bundles Whisper, wav2vec2 forced alignment, and speaker diarization into a single GPU‑ready image. The container can be deployed to Amazon SageMaker as real‑time or asynchronous endpoints, delivering word‑level, speaker‑labeled transcription with low latency suitable for production workloads. The authors highlight that the integration leverages SageMaker’s managed inference infrastructure, enabling scaling from a single GPU to a fleet of instances while preserving the alignment accuracy of WhisperX. The approach trades off a modest increase in model size for the convenience of a unified deployment pipeline.

Build a multi-account AI agent with AgentCore Gateway and MCP
AWS ML Blog· 19 min read· 5 days ago
Build a multi-account AI agent with AgentCore Gateway and MCP

The article demonstrates how to architect a multi‑account AI agent system that keeps each team’s data isolated in separate AWS accounts while enabling a unified query interface. It leverages Amazon Bedrock’s AgentCore Gateway as a central orchestrator, with each line‑of‑business account exposing its data via an MCP server that the gateway can route to. The design eliminates cross‑account data sharing while preserving a single agent experience, though it introduces cross‑account IAM complexity and potential latency from inter‑account calls. The approach is particularly useful for regulated environments that require strict data isolation.

Aderant builds intelligent ticket triage with Amazon Nova
AWS ML Blog· 9 min read· 5 days ago
Aderant builds intelligent ticket triage with Amazon Nova

Aderant deployed an intelligent ticket triage pipeline on Amazon Nova Lite, leveraging Amazon Bedrock to automate context gathering, classification, routing, and knowledge enrichment for its cloud operations team. The solution uses Bedrock’s LLM capabilities to interpret ticket content and Nova Lite’s low‑latency inference to deliver routing decisions in near real‑time, reducing manual triage effort. While the architecture boosts routing accuracy, it introduces a dependency on Bedrock model availability and incurs inference costs tied to the number of tickets processed. The integration demonstrates how enterprise teams can embed LLM‑powered triage directly into existing ticketing workflows.

EXPLORE AI NEWS

Daily hand-picked stories on LLMs, RAG, agents and production AI — curated for engineers who ship.

BROWSE NEWS

GET THE WEEKLY DIGEST

Join engineers getting the Monday signal-over-noise AI breakdown. No spam, unsubscribe anytime.

LEARN AI ENGINEERING

Curated courses, research papers, repos and tutorials built for engineers leveling up in AI.

START LEARNING