HomeAmazon

Amazon

19 curated articles on Amazon for AI engineers

19 articles
Custom reward functions for multi-turn reinforcement learning with Amazon Nova Forge
AWS ML Blog· 16 min read· Yesterday
Custom reward functions for multi-turn reinforcement learning with Amazon Nova Forge

Amazon Nova Forge enables multi-turn reinforcement learning with custom reward functions, allowing for more precise control over model learning. The platform's Bring Your Own Orchestration (BYOO) capability and serverless option provide flexibility in deploying and managing custom reward logic. By designing a well-crafted reward function, engineers can teach models to learn specific behaviors through iterative feedback, optimizing cumulative reward across entire trajectories. This approach has been shown to improve out-of-distribution (OOD) generalization, with reinforcement fine-tuning (RFT) outperforming supervised fine-tuning (SFT) in certain scenarios. For engineers building AI systems, this means that careful consideration of reward function design is crucial for effective model training.

A decade of mathematical certainty: Reflections on the Automated Reasoning Group
Amazon Science· 8 min read· 4 days ago
A decade of mathematical certainty: Reflections on the Automated Reasoning Group

The Automated Reasoning Group (ARG) at Amazon has made significant progress over the past decade in applying mathematical logic and formal verification techniques to prove the correctness and security of AWS systems. The group's production services now process billions of queries daily, and their work has led to the development of tools such as Tiros, Zelkova, and Lean, which are used to analyze network security, policies, and cryptographic protocols. The use of automated reasoning and proof assistants has enabled the group to prove the correctness of complex systems, including the Nitro Confidentiality Engine and the AWS policy interpreter. This work has had a significant impact on the security and reliability of AWS systems, and its practical implications for engineers building AI systems include the potential to apply similar techniques to ensure the correctness and security of AI mod

Building agentic workflows with SageMaker AI and Bedrock AgentCore
AWS ML Blog· 9 min read· Yesterday
Building agentic workflows with SageMaker AI and Bedrock AgentCore

This article presents a technical solution for building agentic workflows by combining Amazon SageMaker AI with Amazon Bedrock AgentCore runtime, enabling the integration of managed foundation models with custom models. The architecture connects three model-hosting paths through a single Amazon Bedrock AgentCore container, utilizing models such as Qwen 3.5 9B on SageMaker AI and Claude Haiku 4.5 on Bedrock. This integration provides cost optimization, data residency, and model flexibility in a single production-ready architecture. The practical implication for engineers building AI systems is the ability to deploy specialized agents that collaborate on complex tasks while using the most suitable models for each task.

AWS Trainium Frontier competition: Co-design models and kernels on purpose-built AI chips
Amazon Science· 6 min read· 4 days ago
AWS Trainium Frontier competition: Co-design models and kernels on purpose-built AI chips

The AWS Trainium Frontier competition invites researchers to co-design models and kernels on purpose-built AI chips, exploring the efficient frontier of model architectures on AWS Trainium. The competition provides a ~50M parameter baseline language model and challenges participants to modify everything, including architecture, optimizer, training loop, and custom NKI kernels, to optimize the model architecture and training throughput within a fixed time and compute budget. The goal is to find the balance between model capacity and training throughput, driving validation bits-per-byte low and downstream in-context learning capability high. The optimal architectures will differ from those designed for existing accelerators due to Trainium's unique hardware features. The practical implication for engineers building AI systems is the opportunity to discover new model architectures designed

From Hugging Face to Amazon SageMaker Studio in one click
Hugging Face Blog· 5 min read· Jul 7, 2026
From Hugging Face to Amazon SageMaker Studio in one click

Not mentioned. The title suggests a connection between Hugging Face and Amazon SageMaker Studio, but details are not provided. This could potentially simplify the deployment process for AI models. The practical implication for engineers building AI systems is not mentioned.

Amazon is investing in the Lean Focused Research Organization
Amazon Science· 5 min read· Jul 26, 2026
Amazon is investing in the Lean Focused Research Organization

Amazon is investing in the Lean Focused Research Organization (FRO) to support the development of Lean, a programming language that enables mathematical proof and correctness guarantees for AI systems. Lean has already been used to verify the correctness of AI agents and systems, such as Policy in Amazon Bedrock AgentCore and AWS Neuron. The investment aims to make proof accessible to every developer, enabling the creation of verified, trustworthy AI agents. This has significant implications for engineers building AI systems, as it provides a way to ensure the correctness and safety of AI decision-making.

Accelerating M&A due diligence with Amazon Bedrock AgentCore
AWS ML Blog· 13 min read· 2 days ago
Accelerating M&A due diligence with Amazon Bedrock AgentCore

Amazon Bedrock AgentCore can accelerate M&A due diligence by orchestrating AI agents that handle data gathering, analysis, and compliance checks autonomously. The platform allows for the building, connection, and optimization of agents at scale, with any framework or model. By leveraging AI agents, due diligence processes can be transformed, reducing the time required for analysis from weeks to hours. The practical implication for engineers building AI systems is the potential to significantly improve the efficiency and effectiveness of M&A due diligence processes.

Amazon Quick for Microsoft 365: Agentic AI where you work
AWS ML Blog· 12 min read· 2 days ago
Amazon Quick for Microsoft 365: Agentic AI where you work

Amazon Quick is now available as an AI assistant directly inside Microsoft 365 apps, including Word, Excel, PowerPoint, and Outlook, allowing users to access and edit data without leaving their familiar productivity tools. The assistant is agentic, meaning it takes action directly within documents, spreadsheets, presentations, and email messages, and is grounded in the user's data, including Amazon Quick Sight dashboards, Spaces, AWS data sources, and third-party integrations. This integration enables users to automate tasks, such as drafting customer request for proposal responses, and customize customer presentations, saving time and increasing productivity. The practical implication for engineers building AI systems is that they can now embed AI into workflows where business decisions happen, leveraging the breadth of connected data to drive more informed decision-making.

Part 2: Amazon Bedrock cost attribution with Amazon Athena and CUDOS
AWS ML Blog· 13 min read· 2 days ago
Part 2: Amazon Bedrock cost attribution with Amazon Athena and CUDOS

Amazon Bedrock's granular cost attribution feature allows for per-user and per-application visibility, and can be visualized and analyzed using Amazon Athena queries and CUDOS dashboards. The process involves setting up a Cost and Usage Report (CUR) 2.0 data export with IAM principal data, which can then be queried using Amazon Athena for analysis. CUDOS dashboards provide pre-built visuals tailored to an organization's specific structure, offering a more streamlined approach to cost and usage analysis. The practical implication for engineers building AI systems is the ability to track and manage costs at a granular level, enabling more efficient resource allocation and cost optimization. With this approach, engineers can typically track usage at the granularity they want for any Bedrock-powered service or application.

The fuel of the future is already here: Why TRISO matters
Amazon Science· 5 min read· Jun 24, 2026
The fuel of the future is already here: Why TRISO matters

Amazon is investing in next-generation nuclear technology, specifically tristructural isotropic (TRISO) fuel particles, to meet the rising energy demands of AI infrastructure and cloud computing. TRISO particles have a ceramic shell with three layers, providing exceptional mechanical integrity and thermal resilience, with a failure fraction of ≤ 6.6 × 10⁻⁵ at 1600°C. This technology offers greater flexibility in fuel form and reactor design, enabling new operational modes and potentially reducing waste. The practical implication for engineers building AI systems is the potential for more efficient and sustainable energy sources to power their infrastructure.

How OneAdvanced deployed over 50 AI agents on UK-sovereign AWS
AWS ML Blog· 13 min read· 3 days ago
How OneAdvanced deployed over 50 AI agents on UK-sovereign AWS

OneAdvanced, a UK-based enterprise software provider, successfully deployed over 50 AI agents on a UK-sovereign AWS architecture, ensuring data sovereignty and compliance with strict regulations. The solution utilizes Llama 4 Maverick and Llama Guard 4 models, self-hosted on Amazon SageMaker AI, and pairs a Retrieval Augmented Generation (RAG) pipeline with Amazon Aurora PostgreSQL-Compatible Edition and the pgvector extension. The architecture supports rapid agent deployment and content moderation, while maintaining control over model serving infrastructure. This approach enables OneAdvanced to meet the sovereignty requirements of their customers, particularly in highly regulated industries.

Pay with confidence: How Solv Labs built verifiable, auditable agent payments on Amazon Bedrock AgentCore payments
AWS ML Blog· 12 min read· 3 days ago
Pay with confidence: How Solv Labs built verifiable, auditable agent payments on Amazon Bedrock AgentCore payments

Solv Labs engineered a fully governed payment workflow for Amazon Bedrock AgentCore, ensuring each agent transaction is authorized, attested inside an AWS Nitro Enclave, priced according to risk, and anchored to a public blockchain before settlement. By combining Bedrock’s native payment API with enclave‑based attestation and blockchain anchoring, the pattern delivers a tamper‑evident, auditable ledger that satisfies enterprise compliance needs. The approach

Tiered KV cache for large LLMs on Amazon SageMaker HyperPod with Curvine
AWS ML Blog· 31 min read· 3 days ago
Tiered KV cache for large LLMs on Amazon SageMaker HyperPod with Curvine

The article introduces a tiered key‑value (KV) cache for large language models running on Amazon SageMaker HyperPod, leveraging Curvine’s distributed NVMe pool to extend cache capacity beyond GPU memory. By layering GPU‑resident cache with a shared NVMe tier, the approach cuts first‑token latency while keeping memory footprints manageable, enabling higher throughput without scaling to oversized GPU instances. The design trades off modest storage costs for significant speed gains, making it attractive for production deployments that demand low‑latency inference at scale.

How ONESTRUCTION built the Ishigaki-IDS foundation model with AWS GenAIIC
AWS ML Blog· 9 min read· 4 days ago
How ONESTRUCTION built the Ishigaki-IDS foundation model with AWS GenAIIC

ONESTRUCTION built the Ishigaki-IDS foundation model with technical advisory from AWS GenAIIC, addressing data scarcity and specialized knowledge requirements in the construction industry. The model was trained using a three-stage pipeline (CPT, SFT, RLVR) and synthetic data generation to overcome data scarcity. The model's performance was improved by injecting an IFC vocabulary and using verifiable rewards for structured output generation. This approach has practical implications for engineers building AI systems in data-scarce domains, enabling them to develop specialized models with limited training data.

How Pixieset achieved 35% AI feature adoption by solving the right problem with Amazon Bedrock
AWS ML Blog· 9 min read· 4 days ago
How Pixieset achieved 35% AI feature adoption by solving the right problem with Amazon Bedrock

Pixieset, a photography business service, achieved 35% AI feature adoption by solving the problem of generating alt text for images, a task that pulls photographers away from their craft. Using Amazon Bedrock, they launched an AI image alt text generator in 4 months, which generated alt text for over 750,000 photos in the first week. The key to their success was identifying a real problem that photographers face and applying generative AI to alleviate that friction. This approach led to significant subscription upgrades and sustained feature adoption. The practical implication for engineers building AI systems is to focus on solving specific, high-impact problems that users face, rather than trying to force AI into every aspect of their workflow.

First Orion accelerates QA automation using Amazon Nova Act
AWS ML Blog· 13 min read· 4 days ago
First Orion accelerates QA automation using Amazon Nova Act

First Orion, a branded communications company, accelerated its QA automation using Amazon Nova Act, shifting from script-based test automation to AI-driven agents that understand web interfaces like humans. This change enabled the company to keep pace with its rapid development velocity, improving release quality and reducing engineering time spent on regressions. The new architecture built around Amazon Nova Act allowed First Orion to test its web applications more efficiently, despite the growing number of device form-factors and browser versions. As a result, the company can now deliver higher-quality releases faster, with practical implications for engineers building AI-powered QA systems.

Run interactive IDEs on Amazon EKS with SageMaker AI to power up your AI workflows
AWS ML Blog· 16 min read· 5 days ago
Run interactive IDEs on Amazon EKS with SageMaker AI to power up your AI workflows

The Amazon SageMaker AI Spaces add-on for Amazon EKS enables data scientists to run interactive IDEs like JupyterLab and Code Editor on the same cluster as their pipelines, eliminating the need to switch to a standalone JupyterHub deployment or local laptop. This solution can increase GPU utilization by up to 30 percent and reduce costs by avoiding the need for an always-on GPU environment. The add-on can be set up in about 5 minutes, compared to the 3-5 days it typically takes to stand up a standalone JupyterHub environment. The solution runs on a single EKS cluster in three layers: network and access, cluster routing, and compute and storage. For engineers building AI systems, this means they can streamline their workflow and improve productivity by having all their tools and resources in one place.

How nOps shipped FinOps agents 75% faster with Amazon Bedrock AgentCore
AWS ML Blog· 10 min read· 5 days ago
How nOps shipped FinOps agents 75% faster with Amazon Bedrock AgentCore

nOps, an AI-powered cloud optimization solution, has successfully transitioned its FinOps analytics capabilities to Amazon Bedrock AgentCore, resulting in a 75% faster shipping of FinOps agents. The new architecture, centered on Bedrock AgentCore, Databricks Metric Views, and Databricks Lakebase, has improved response quality, reduced operational complexity, and enabled the team to focus on domain logic rather than infrastructure. This transition has allowed nOps to better serve its customers, who manage over $4 billion in cloud spend. The practical implication for engineers building AI systems is that using a purpose-built architecture like Amazon Bedrock AgentCore can significantly accelerate product delivery and improve system reliability.

How TReNDS automates root-cause analysis with Amazon Bedrock
AWS ML Blog· 14 min read· Aug 7, 2026
How TReNDS automates root-cause analysis with Amazon Bedrock

The TReNDS Center at Georgia State University has developed an architecture that automates root-cause analysis using Amazon Bedrock, Amazon CloudWatch subscription filters, AWS Lambda, and the Strands Agents SDK. This system detects errors in real-time, enriches them with log context and source code from GitHub, and delivers AI-powered root-cause analysis to the team, reducing investigation time from 15-30 minutes to near real-time. The core of the system is Amazon Bedrock, which does the actual reasoning about errors, code, and root causes. The practical implication for engineers building AI systems is that they can leverage similar architectures to automate incident response and reduce downtime.

EXPLORE AI NEWS

Daily hand-picked stories on LLMs, RAG, agents and production AI — curated for engineers who ship.

BROWSE NEWS

GET THE WEEKLY DIGEST

Join engineers getting the Monday signal-over-noise AI breakdown. No spam, unsubscribe anytime.

LEARN AI ENGINEERING

Curated courses, research papers, repos and tutorials built for engineers leveling up in AI.

START LEARNING