AWS HPC Blog
Category: Artificial Intelligence
Resilient HPC and ML on AWS: Running Tightly Coupled Workloads on Spot Instances
This post references AWS ParallelCluster. Check out AWS Parallel Computing Service (AWS PCS), our new managed Slurm service for running HPC and AI workloads on AWS. This post was contributed by Santosh Kumar, Bhagyaraju Kasina, Dr. Sandeep Sovani and Dr. Max Starr Researchers and engineering teams running High Performance Computing (HPC) jobs face a constant […]
Transforming HPC Operations with Intelligent Workload Orchestration on AWS
This post was contributed by Manu Pillai, Gloria Macia and Natalia Jimenez, PhD Organizations running high-performance computing (HPC) workloads today operate largely as they have for decades: users manually specify required compute specifications for each of their jobs. Users spend valuable time analyzing workload requirements, selecting instance types, and troubleshooting infrastructure issues – time that […]
Accelerating CFD development from years to weeks with agentic AI and AWS
Agentic AI is revolutionizing computational fluid dynamics (CFD) simulations, enabling experienced engineers to focus on physics, innovation, and engineering judgment rather than tedious coding and debugging. Our latest blog explores how this transformative technology can help your team deliver complex projects more rapidly while maintaining scientific rigor.
Running NVIDIA Cosmos world foundation models on AWS
Running NVIDIA Cosmos world foundation models on AWS provides powerful physical AI capabilities at scale. This blog covers two production-ready architectures, each optimized for different organizational needs and constraints.
Meet the Advanced Computing team of AWS at SC25 in St. Louis
This post references AWS ParallelCluster. Check out AWS Parallel Computing Service (AWS PCS), our new managed Slurm service for running HPC and AI workloads on AWS. We want to empower every scientist and engineer to solve hard problems by giving them access to the compute and analytical tools they need, when they need them. Cloud […]
AWS re:Invent 2025: Your Complete Guide to High Performance Computing Sessions
This post references AWS ParallelCluster. Check out AWS Parallel Computing Service (AWS PCS), our new managed Slurm service for running HPC and AI workloads on AWS. AWS re:Invent 2025 returns to Las Vegas, Nevada on December 1, uniting AWS builders, customers, partners, and IT professionals from across the globe. This year’s event offers you exclusive […]
Leveraging LLMs as an Augmentation to Traditional Hyperparameter Tuning
When seeking to improve machine learning model performance, hyperparameter tuning is often the go-to recommendation. However, this approach faces significant limitations, particularly for complex models requiring extensive training times. In this post, we’ll explore a novel approach that combines gradient norm analysis with Large Language Model (LLM) guidance to intelligently redesign neural network architectures. This […]
Enhanced Performance for Whisper Audio Transcription on AWS Batch and AWS Inferentia
In this post, we’ll review the key optimizations and performance gains for our Whisper audio transcription solution powered by AWS Batch and AWS Inferentia.
Engineering at the speed of thought: Accelerating complex processes with multi-agent AI and Synera
In this post, we’ll examine how this multi-agent approach works, the architecture behind it, and the efficiency improvements it enables. While the focus is on an engineering use case, the principles apply broadly to any organization facing the challenge of coordinating specialized expertise to deliver faster, more consistent results.
Scale Reinforcement Learning with AWS Batch Multi-Node Parallel Jobs
Autonomous robots are increasingly used across industries, from warehouses to space exploration. While developing these robots requires complex simulation and reinforcement learning (RL), setting up training environments can be challenging and time-consuming. AWS Batch multi-node parallel (MNP) infrastructure, combined with NVIDIA Isaac Lab, offers a solution by providing scalable, cost-effective robot training capabilities for sophisticated behaviors and complex tasks.








