HomeCompute

Compute

AI infrastructure and compute: GPU availability, cloud pricing, hardware releases, and how compute constraints shape model architecture decisions.

23 articles

23 articles
GLM-5.3 is here with advanced cyber capabilities — and reportedly already found a 'serious vulnerability' in Cursor
VentureBeat AI· 8 min read· Today
GLM-5.3 is here with advanced cyber capabilities — and reportedly already found a 'serious vulnerability' in Cursor

The Chinese AI startup Z.ai has released GLM-5.3, a language model with substantial gains in long-horizon coding and advanced cybersecurity capabilities, which has already found a potentially serious vulnerability in Cursor. GLM-5.3 builds on the 743-billion-parameter-scale base model of GLM-5.2, with improvements coming from scaling post-training across more environments and tasks. The model achieves sizable generation-over-generation improvements on various benchmarks, including Terminal-Bench 3.0, DeepSWE v1.1, and AutomationBench. The practical implication for engineers building AI systems is that GLM-5.3 demonstrates the potential for significant improvements in language models through post-training scaling, rather than relying on expensive pretraining cycles.

Universitas Gadjah Mada, Indosat and NVIDIA Open Indonesia’s First University AI Center to Develop Local AI Talent
NVIDIA Blog· 4 min read· Today
Universitas Gadjah Mada, Indosat and NVIDIA Open Indonesia’s First University AI Center to Develop Local AI Talent

The Universitas Gadjah Mada, Indosat, and NVIDIA have launched the UGM Indosat NVIDIA AI Technology Center (NVAITC) in Yogyakarta, Indonesia's first university-based AI technology center, to develop local AI talent and address Indonesia's most urgent national priorities. The center is powered by NVIDIA's full-stack AI platform and GPU Merdeka, Indosat's sovereign GPU-as-a-service platform, providing access to enterprise-grade accelerated computing, AI software, and technical mentorship. The initial projects focus on healthcare, agriculture, and natural disaster response, aiming to drive real change and innovation with local and global impact. This initiative has the potential to equip Indonesian talent with the necessary tools and expertise to turn their potential into innovation.

With a feel for physics, AI models simulate a wider range of real-world scenarios
MIT News AI· 5 min read· 4 days ago
With a feel for physics, AI models simulate a wider range of real-world scenarios

Researchers at MIT's CSAIL and Tsinghua University have developed a new pre-training approach called GeoPT, which enables simulation models to learn physics in a broader and more efficient way, allowing them to model the real world more accurately and train on up to 60% less data. GeoPT uses synthetic dynamics, a series of interactions between small particles and complex 3D shapes, to give models a sense of how physics works. This approach can help engineers predict how vehicles, everyday items, and robots respond to various physical elements, such as wind, water, and collisions. The practical implication for engineers building AI systems is that GeoPT can accelerate the development of more realistic and accurate simulations, enabling the creation of more reliable and efficient AI models.

NVIDIA AI Factory Compute Is Becoming an Investable Asset Class
NVIDIA Blog· 6 min read· 3 days ago
NVIDIA AI Factory Compute Is Becoming an Investable Asset Class

NVIDIA has announced partnerships with major financial institutions to establish independent financing platforms for AI infrastructure, aiming to mobilize over $500 billion in third-party capital. This development marks a significant milestone in the AI industry, as AI factories can now be financed as productive infrastructure, with repeatable platforms and long-term institutional capital. The NVIDIA AI factory platform, including accelerated computing, networking, systems software, and AI frameworks, can run a broad range of AI models and is built on a globally adopted architecture. This flexibility and fungibility, combined with the continuous improvement of CUDA, make NVIDIA compute a valuable and investable asset. The practical implication for engineers building AI systems is that they can now access scalable and flexible infrastructure to support their production needs.

Nebius shares jump 34% on continued AI infrastructure demand
SiliconANGLE AI· 2 days ago
Nebius shares jump 34% on continued AI infrastructure demand

Nebius Group NV's shares jumped 34% after reporting second-quarter earnings that exceeded expectations, driven by strong demand for its AI infrastructure cloud platform. The company's revenue surged 454% due to its optimized cloud platform for artificial intelligence workloads. This growth indicates a significant increase in the adoption of AI solutions, which has a practical implication for engineers building AI systems to meet the rising demand for scalable and efficient infrastructure. As a result, engineers should focus on developing and optimizing AI infrastructure to support the growing needs of businesses.

AWS Trainium Frontier competition: Co-design models and kernels on purpose-built AI chips
Amazon Science· 6 min read· 4 days ago
AWS Trainium Frontier competition: Co-design models and kernels on purpose-built AI chips

The AWS Trainium Frontier competition invites researchers to co-design models and kernels on purpose-built AI chips, exploring the efficient frontier of model architectures on AWS Trainium. The competition provides a ~50M parameter baseline language model and challenges participants to modify everything, including architecture, optimizer, training loop, and custom NKI kernels, to optimize the model architecture and training throughput within a fixed time and compute budget. The goal is to find the balance between model capacity and training throughput, driving validation bits-per-byte low and downstream in-context learning capability high. The optimal architectures will differ from those designed for existing accelerators due to Trainium's unique hardware features. The practical implication for engineers building AI systems is the opportunity to discover new model architectures designed

The Pulse: Interesting AI coding stats from Cursor
Pragmatic Engineer· 6 min read· Jul 9, 2026
The Pulse: Interesting AI coding stats from Cursor

A recent report from Cursor reveals that power users generate 10x as many lines of code as the median, with the top 1% of users creating around 30-40K lines of code per week. The report also shows that Cursor consumes 10x more input tokens than it generates in output tokens, with 90% of token usage being input tokens. This highlights the importance of caching context to reduce token costs, with Cursor's caching mechanism reducing token costs by 10x. The practical implication for engineers building AI systems is to prioritize context reuse and caching to improve efficiency.

Why Scaling AI Compute Performance Requires a New Power Architecture
NVIDIA Blog· 4 min read· 4 days ago
Why Scaling AI Compute Performance Requires a New Power Architecture

The increasing demand for AI compute performance requires a new power architecture, with 800 VDC simplifying the power delivery path and reducing inefficiencies. NVIDIA, Google, and Microsoft have developed the 800 VDC architecture through the Open Compute Project, publishing a joint white paper and specification. The new architecture provides a roadmap for AI factories to scale, with on-ramps at every stage of growth, including hybrid-compatible power racks, row power centers, and DC power blocks. This development has significant implications for engineers building AI systems, as it enables higher compute density and more efficient power distribution.

Daniela Rus receives Bavarian Minister-President's High-Tech Prize
MIT News AI· 3 min read· Jul 30, 2026
Daniela Rus receives Bavarian Minister-President's High-Tech Prize

Daniela Rus, director of MIT's Computer Science and Artificial Intelligence Laboratory, has received the 2026 High-Tech Prize of the Bavarian Minister-President for her contributions to robotics, artificial intelligence, and autonomous systems, recognizing her 30-year effort to build machines that can operate outside the lab. Her work includes self-organizing robot collectives, soft robotics, autonomous mobility, and brain-inspired artificial intelligence. Rus' research has led to the development of innovative solutions such as ingestible origami robots and liquid neural networks. The practical implication for engineers building AI systems is the potential to create more efficient and adaptable machines that can operate in real-world environments.

Professor Emeritus Dimitri Bertsekas, influential computer scientist and prolific author, dies at 83
MIT News AI· 9 min read· Jul 22, 2026
Professor Emeritus Dimitri Bertsekas, influential computer scientist and prolific author, dies at 83

Dimitri Bertsekas, a renowned computer scientist and author, passed away at 83, leaving a lasting impact on fields including optimization, control, large-scale computation, reinforcement learning, and artificial intelligence. His research and teachings have influenced numerous students, colleagues, and institutions. Bertsekas authored over 20 influential books and monographs, and his work continues to shape the foundations of these fields. His legacy will have a lasting impact on engineers and researchers building AI systems, particularly in the areas of optimization and reinforcement learning.

Writer says its new Palmyra X6 model cuts AI agent costs by 52% as token spending surges
VentureBeat AI· 12 min read· 2 days ago
Writer says its new Palmyra X6 model cuts AI agent costs by 52% as token spending surges

Writer’s new Palmyra X6 model slashes AI agent costs by 52% amid rising token consumption, while the company simultaneously unveiled a redesigned agent orchestration harness and enhanced governance tools to curb runaway token usage. The updated harness introduces a modular pipeline for orchestrating multi‑step agents, and the governance suite exposes fine‑grained token‑budget controls to IT leaders. Together, these changes aim to keep large‑scale agent deployments within budget while preserving performance.

Firebird Launches CIS Region’s Largest AI Factory in Armenia
NVIDIA Blog· 4 min read· Aug 8, 2026
Firebird Launches CIS Region’s Largest AI Factory in Armenia

Firebird has launched the CIS region's largest AI factory in Armenia, powered by NVIDIA accelerated computing and Dell Technologies high-performance AI infrastructure, with plans to deploy over 70,000 NVIDIA GPUs and 300 megawatts of AI infrastructure capacity by 2027. The AI factory is designed to provide computing capacity for training, fine-tuning, and deploying AI models at scale, and is expected to accelerate Armenia's development as a center for AI research and innovation. With a focus on energy efficiency, the AI factory integrates accelerated computing, networking, power, and cooling as one codesigned system, allowing it to run up to 40% more GPUs on the same footprint. This launch has significant implications for engineers building AI systems, as it provides a large-scale infrastructure for developing and deploying AI models.

Measuring Performance of Transformer Inference
Machine Learning Mastery· Aug 4, 2026
Measuring Performance of Transformer Inference

This chapter outlines a systematic approach to quantifying transformer inference performance, covering everything from per-request latency to multi‑GPU scaling and cost‑per‑token analysis. It introduces practical measurement techniques such as CUDA event timing for GPU workload, memory profiling to capture peak usage, and warm‑up strategies to stabilize latency estimates. The guide also discusses concurrent request handling and how to aggregate metrics across multiple machines, providing a clear path to evaluate both speed and cost efficiency. By applying these methods, engineers can pinpoint bottlenecks and make data‑driven decisions on model deployment.

A better way to turn 2D designs into 3D models for rapid prototyping
MIT News AI· 5 min read· Jul 16, 2026
A better way to turn 2D designs into 3D models for rapid prototyping

Researchers from MIT and elsewhere have developed a system that can teach a vision-language model to automatically convert 2D designs into CAD programs, generating more accurate and functional 3D models while using only a fraction of the computation. The system uses a process known as data augmentation to create new data based on the model's abilities and corrects the model's failures, incorporating them into a dataset to teach the model how to fix specific mistakes. This technique could streamline the rapid prototyping process, reduce costs, and help engineers identify beneficial design choices. The researchers are working toward building vision-language models for CAD generation, which take a 2D image and some descriptive text, and output Python code that can be executed in a CAD software program to generate a 3D model.

Tiered KV cache for large LLMs on Amazon SageMaker HyperPod with Curvine
AWS ML Blog· 31 min read· 3 days ago
Tiered KV cache for large LLMs on Amazon SageMaker HyperPod with Curvine

The article introduces a tiered key‑value (KV) cache for large language models running on Amazon SageMaker HyperPod, leveraging Curvine’s distributed NVMe pool to extend cache capacity beyond GPU memory. By layering GPU‑resident cache with a shared NVMe tier, the approach cuts first‑token latency while keeping memory footprints manageable, enabling higher throughput without scaling to oversized GPU instances. The design trades off modest storage costs for significant speed gains, making it attractive for production deployments that demand low‑latency inference at scale.

How ONESTRUCTION built the Ishigaki-IDS foundation model with AWS GenAIIC
AWS ML Blog· 9 min read· 4 days ago
How ONESTRUCTION built the Ishigaki-IDS foundation model with AWS GenAIIC

ONESTRUCTION built the Ishigaki-IDS foundation model with technical advisory from AWS GenAIIC, addressing data scarcity and specialized knowledge requirements in the construction industry. The model was trained using a three-stage pipeline (CPT, SFT, RLVR) and synthetic data generation to overcome data scarcity. The model's performance was improved by injecting an IFC vocabulary and using verifiable rewards for structured output generation. This approach has practical implications for engineers building AI systems in data-scarce domains, enabling them to develop specialized models with limited training data.

Determining playoff clinching scenarios in the NHL using constraint programming
AWS ML Blog· 6 min read· Aug 7, 2026
Determining playoff clinching scenarios in the NHL using constraint programming

The AWS Generative AI Innovation Center developed an automated system to determine NHL playoff clinching scenarios using constraint programming and custom tree search. The system consists of a 0-day solver that checks if a team has already clinched the playoffs and an n-day lookahead solver that generates scenarios for teams that could clinch based on upcoming games. The approach accounts for the NHL's complex tie-breaking rules and was validated against official NHL results. The practical implication for engineers building AI systems is the application of constraint programming to solve complex combinatorial challenges in real-world domains.

NVIDIA Joins NSF State and Regional AI Hubs Program to Expand AI Research and Education Across the US
NVIDIA Blog· 4 min read· Aug 4, 2026
NVIDIA Joins NSF State and Regional AI Hubs Program to Expand AI Research and Education Across the US

NVIDIA has joined the U.S. National Science Foundation's State and Regional Artificial Intelligence Infrastructure Hubs program to expand access to AI research and education across the US. The program aims to support state and multistate groups of colleges and universities in strengthening America's AI ecosystem by providing advanced computing, data, software, and expertise. With a focus on sharing AI computing resources and accelerating scientific discovery, the regional hubs will help prepare students to participate in the AI economy. The initiative is expected to have a significant impact on the development of AI infrastructure and workforce development, with NVIDIA's partnership serving as a model for other institutions.

As AI Increases Demands on Memory, Storage Steps Up
NVIDIA Blog· 5 min read· Aug 4, 2026
As AI Increases Demands on Memory, Storage Steps Up

The increasing demands of AI on memory and storage have led to the need for more efficient and secure storage architectures, with NVIDIA unveiling new storage advancements at the Future of Memory and Storage (FMS) conference. The NVIDIA Vera CPU delivers up to 3.21x higher throughput than an x86 CPU in a two-stage compression and encryption pipeline, enabling storage platforms to absorb AI data more efficiently. The open sourcing of NVIDIA cuFile APIs enables interoperability for storage solutions, allowing GPUs to read from and write to storage directly. This development has significant implications for engineers building AI systems, as it enables faster and more secure access to data and storage.

Powerful Compute So Compact, It’s Clutch — Build AI Anywhere With NVIDIA Jetson
NVIDIA Blog· 6 min read· Jul 28, 2026
Powerful Compute So Compact, It’s Clutch — Build AI Anywhere With NVIDIA Jetson

The NVIDIA Jetson platform provides a compact and powerful solution for building AI anywhere, with modules and developer kits that can fit in a handbag. The Jetson Orin Nano Super, in particular, offers 67 trillion operations per second (TOPS) of AI performance, making it ideal for building a first AI robot. This platform enables developers to build, learn, and launch the next generation of intelligent robots, with applications in classrooms, labs, and makerspaces. The practical implication for engineers building AI systems is that they can now develop and deploy AI models in a more portable and efficient manner.

Toward a future that preserves benefits of neurotechnology for all
MIT News AI· 4 min read· Jul 6, 2026
Toward a future that preserves benefits of neurotechnology for all

The Envisioning the Future of Computing Prize, presented by the Social and Ethical Responsibilities of Computing, has awarded Rachel Sava for her submission "Superintelligence, Superintimate", which explores the potential risks and benefits of neurotechnology, particularly neural implants. Sava's work highlights the need for guardrails on protected usage as advanced medical technology hits consumer markets. The prize aims to encourage students to consider the societal benefits and costs of technological advancements from the outset. The practical implication for engineers building AI systems is to prioritize ethical considerations and responsible innovation in their work.

3 Questions: Beyond data-driven aesthetics
MIT News AI· 5 min read· Jun 29, 2026
3 Questions: Beyond data-driven aesthetics

The "Beyond Data-Driven Aesthetics" exhibition, led by MIT Architecture alumnus Alexandros Haridis, explores the intersection of computation, aesthetics, and design, translating algorithms and machine-learning systems into physical installations and interactive visualizations. The exhibition draws on research in design computation, shape grammars, and aesthetic theories, examining the relationships between human insight and computation. With a focus on making computational systems more tangible and interpretable, the exhibition aims to capture the salient ideas of research papers and books in a visual, spatial, and experiential format. The practical implication for engineers building AI systems is to consider the potential of design and visualization techniques to interpret and communicate complex computational concepts.

Improving the speed and energy-efficiency of AI agents
MIT News AI· 5 min read· Jun 25, 2026
Improving the speed and energy-efficiency of AI agents

Researchers from MIT and Microsoft have developed an intelligent system that streamlines the process of designing agentic workflows, automatically optimizing the implementation and reducing computational units, energy requirements, and costs. The system allows developers to describe the desired workflow in plain language, without needing to specify all details in advance, and adjusts configurations on the fly based on user priorities. This approach has been shown to significantly cut energy requirements and costs compared to traditional approaches without hampering performance. The practical implication for engineers building AI systems is that they can now design and deploy more efficient agentic workflows, reducing waste and improving overall system performance.

EXPLORE AI NEWS

Daily hand-picked stories on LLMs, RAG, agents and production AI — curated for engineers who ship.

BROWSE NEWS

GET THE WEEKLY DIGEST

Join engineers getting the Monday signal-over-noise AI breakdown. No spam, unsubscribe anytime.

LEARN AI ENGINEERING

Curated courses, research papers, repos and tutorials built for engineers leveling up in AI.

START LEARNING