HomeCompute

Compute

AI infrastructure and compute: GPU availability, cloud pricing, hardware releases, and how compute constraints shape model architecture decisions.

11 articles

11 articles
With a feel for physics, AI models simulate a wider range of real-world scenarios
MIT News AI· 5 min read· Aug 10, 2026
With a feel for physics, AI models simulate a wider range of real-world scenarios

Researchers at MIT's CSAIL and Tsinghua University have developed a new pre-training approach called GeoPT, which enables simulation models to learn physics in a broader and more efficient way, allowing them to model the real world more accurately and train on up to 60% less data. GeoPT uses synthetic dynamics, a series of interactions between small particles and complex 3D shapes, to give models a sense of how physics works. This approach can help engineers predict how vehicles, everyday items, and robots respond to various physical elements, such as wind, water, and collisions. The practical implication for engineers building AI systems is that GeoPT can accelerate the development of more realistic and accurate simulations, enabling the creation of more reliable and efficient AI models.

AWS Trainium Frontier competition: Co-design models and kernels on purpose-built AI chips
Amazon Science· 6 min read· Aug 10, 2026
AWS Trainium Frontier competition: Co-design models and kernels on purpose-built AI chips

The AWS Trainium Frontier competition invites researchers to co-design models and kernels on purpose-built AI chips, exploring the efficient frontier of model architectures on AWS Trainium. The competition provides a ~50M parameter baseline language model and challenges participants to modify everything, including architecture, optimizer, training loop, and custom NKI kernels, to optimize the model architecture and training throughput within a fixed time and compute budget. The goal is to find the balance between model capacity and training throughput, driving validation bits-per-byte low and downstream in-context learning capability high. The optimal architectures will differ from those designed for existing accelerators due to Trainium's unique hardware features. The practical implication for engineers building AI systems is the opportunity to discover new model architectures designed

Learn Vectorized Thinking in Python Through Examples
Machine Learning Mastery· Aug 26, 2026
Learn Vectorized Thinking in Python Through Examples

The article introduces vectorized thinking in Python, demonstrating how to replace slow Python loops with efficient NumPy array operations. It focuses on leveraging NumPy’s broadcasting and array‑level functions to perform computations in bulk, thereby reducing code complexity and improving runtime performance. The core lesson is that thinking in terms of whole‑array operations can lead to cleaner, faster, and more scalable data‑processing code.

Daniela Rus receives Bavarian Minister-President's High-Tech Prize
MIT News AI· 3 min read· Jul 30, 2026
Daniela Rus receives Bavarian Minister-President's High-Tech Prize

Daniela Rus, director of MIT's Computer Science and Artificial Intelligence Laboratory, has received the 2026 High-Tech Prize of the Bavarian Minister-President for her contributions to robotics, artificial intelligence, and autonomous systems, recognizing her 30-year effort to build machines that can operate outside the lab. Her work includes self-organizing robot collectives, soft robotics, autonomous mobility, and brain-inspired artificial intelligence. Rus' research has led to the development of innovative solutions such as ingestible origami robots and liquid neural networks. The practical implication for engineers building AI systems is the potential to create more efficient and adaptable machines that can operate in real-world environments.

Disaggregation Is a Thousand-GPU Problem
Towards Data Science· 5 days ago
Disaggregation Is a Thousand-GPU Problem

The article explains that splitting the prefill and decode phases of large‑language‑model inference only yields gains when three specific conditions are met; otherwise, the simpler chunked‑prefill strategy remains the better default. It highlights that the prefill‑decode split is most effective only at very large GPU scales, whereas chunked prefill keeps latency low and resource usage predictable for smaller clusters. The discussion clarifies that engineers should not blindly enable the split but should first verify that the workload satisfies the outlined criteria.

Professor Emeritus Dimitri Bertsekas, influential computer scientist and prolific author, dies at 83
MIT News AI· 9 min read· Jul 22, 2026
Professor Emeritus Dimitri Bertsekas, influential computer scientist and prolific author, dies at 83

Dimitri Bertsekas, a renowned computer scientist and author, passed away at 83, leaving a lasting impact on fields including optimization, control, large-scale computation, reinforcement learning, and artificial intelligence. His research and teachings have influenced numerous students, colleagues, and institutions. Bertsekas authored over 20 influential books and monographs, and his work continues to shape the foundations of these fields. His legacy will have a lasting impact on engineers and researchers building AI systems, particularly in the areas of optimization and reinforcement learning.

A better way to turn 2D designs into 3D models for rapid prototyping
MIT News AI· 5 min read· Jul 16, 2026
A better way to turn 2D designs into 3D models for rapid prototyping

Researchers from MIT and elsewhere have developed a system that can teach a vision-language model to automatically convert 2D designs into CAD programs, generating more accurate and functional 3D models while using only a fraction of the computation. The system uses a process known as data augmentation to create new data based on the model's abilities and corrects the model's failures, incorporating them into a dataset to teach the model how to fix specific mistakes. This technique could streamline the rapid prototyping process, reduce costs, and help engineers identify beneficial design choices. The researchers are working toward building vision-language models for CAD generation, which take a 2D image and some descriptive text, and output Python code that can be executed in a CAD software program to generate a 3D model.

Build a Physical AI model factory with NVIDIA Cosmos 3 on SageMaker HyperPod
AWS ML Blog· 38 min read· 5 days ago
Build a Physical AI model factory with NVIDIA Cosmos 3 on SageMaker HyperPod

The article demonstrates how to orchestrate a continuous AI model factory—covering synthetic data generation, post‑training optimization, and closed‑loop evaluation—using NVIDIA Cosmos 3 on a persistent, resilient Amazon SageMaker HyperPod cluster deployed via Amazon EKS. By leveraging Cosmos 3 for real‑time evaluation and SageMaker HyperPod’s high‑throughput compute, the pipeline eliminates the need

Securing the Infrastructure of Intelligence
NVIDIA Blog· 5 min read· Aug 17, 2026
Securing the Infrastructure of Intelligence

AI factories are emerging as the core infrastructure of the AI era, with compute becoming a direct revenue driver. These factories depend on a tightly coupled stack of advanced chips, specialized packaging, and high‑density memory to convert raw energy and data into actionable intelligence. The article highlights that the efficiency of this stack determines both performance and cost, underscoring the need for careful design trade‑offs in heat management and power delivery. Ultimately, businesses that can optimize this infrastructure

Universitas Gadjah Mada, Indosat and NVIDIA Open Indonesia’s First University AI Center to Develop Local AI Talent
NVIDIA Blog· 4 min read· Aug 14, 2026
Universitas Gadjah Mada, Indosat and NVIDIA Open Indonesia’s First University AI Center to Develop Local AI Talent

The Universitas Gadjah Mada, Indosat, and NVIDIA have launched the UGM Indosat NVIDIA AI Technology Center (NVAITC) in Yogyakarta, Indonesia's first university-based AI technology center, to develop local AI talent and address Indonesia's most urgent national priorities. The center is powered by NVIDIA's full-stack AI platform and GPU Merdeka, Indosat's sovereign GPU-as-a-service platform, providing access to enterprise-grade accelerated computing, AI software, and technical mentorship. The initial projects focus on healthcare, agriculture, and natural disaster response, aiming to drive real change and innovation with local and global impact. This initiative has the potential to equip Indonesian talent with the necessary tools and expertise to turn their potential into innovation.

NVIDIA AI Factory Compute Is Becoming an Investable Asset Class
NVIDIA Blog· 6 min read· Aug 12, 2026
NVIDIA AI Factory Compute Is Becoming an Investable Asset Class

NVIDIA has announced partnerships with major financial institutions to establish independent financing platforms for AI infrastructure, aiming to mobilize over $500 billion in third-party capital. This development marks a significant milestone in the AI industry, as AI factories can now be financed as productive infrastructure, with repeatable platforms and long-term institutional capital. The NVIDIA AI factory platform, including accelerated computing, networking, systems software, and AI frameworks, can run a broad range of AI models and is built on a globally adopted architecture. This flexibility and fungibility, combined with the continuous improvement of CUDA, make NVIDIA compute a valuable and investable asset. The practical implication for engineers building AI systems is that they can now access scalable and flexible infrastructure to support their production needs.

EXPLORE AI NEWS

Daily hand-picked stories on LLMs, RAG, agents and production AI — curated for engineers who ship.

BROWSE NEWS

GET THE WEEKLY DIGEST

Join engineers getting the Monday signal-over-noise AI breakdown. No spam, unsubscribe anytime.

LEARN AI ENGINEERING

Curated courses, research papers, repos and tutorials built for engineers leveling up in AI.

START LEARNING