Artificial Intelligence

Category: Amazon SageMaker

Generate images and video with vLLM-Omni on SageMaker AI - Part 2

Generate images and video with vLLM-Omni on SageMaker AI – Part 2

Deploy two generative media models from one AWS vLLM-Omni Deep Learning Container on Amazon SageMaker AI. Generate an image with FLUX.2-klein through real-time inference, then animate it into video with Wan2.1-VACE through asynchronous inference, and retrieve the MP4 from Amazon S3.

Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod

Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod

Learn how to run SkyRL, an open-source reinforcement learning framework, on Amazon SageMaker HyperPod to post-train a Qwen3-VL-8B vision-language model with GRPO. This walkthrough covers building the container image, launching a Ray cluster from SageMaker Studio, submitting and monitoring the job, and hosting the trained LoRA adapter for inference.

Multi-Region training with Amazon SageMaker HyperPod and Qumulo

Multi-Region training with Amazon SageMaker HyperPod and Qumulo

Amazon SageMaker HyperPod and Cloud Native Qumulo let you place training compute in one AWS Region while keeping your dataset in another. This post shares the architecture and validation results from a cross-Region training run, where a remote cluster matched a co-located cluster’s throughput after a brief NeuralCache warmup.

Speaker-labeled transcription with WhisperX on SageMaker AI

Speaker-labeled transcription with WhisperX on SageMaker AI

The AWS WhisperX Deep Learning Container packages Whisper, wav2vec2 forced alignment, and speaker diarization into a GPU-ready image. Learn how to deploy it to Amazon SageMaker AI real-time and asynchronous endpoints for word-level, speaker-labeled transcription, plus the production details that matter: the GPU AMI pin, scaling, and cost controls.

Right-size generative AI endpoints with concurrency sweeps on Amazon SageMaker AI

Right-size generative AI endpoints with concurrency sweeps on Amazon SageMaker AI

Concurrency sweeps help you right-size a generative AI endpoint on Amazon SageMaker AI by systematically benchmarking it at increasing load levels. This post walks through deploying a model, running automated concurrency sweeps with the CreateAIBenchmarkJob API, and using the results to make data-driven capacity decisions about fleet size.

How Tata Elxsi detects industrial safety risks in seconds on AWS

How Tata Elxsi detects industrial safety risks in seconds on AWS

Learn how Tata Elxsi built IRIS, a real-time industrial safety platform on AWS. IRIS filters camera video at the edge, streams metadata through Amazon Kinesis, runs computer vision on Amazon SageMaker AI, and correlates detections into high-confidence alerts, detecting unsafe conditions in seconds instead of minutes.

Run Positron on Amazon SageMaker AI for data science workflows

Run Positron on Amazon SageMaker AI for data science workflows

Positron, Posit’s IDE for data science, now runs on Amazon SageMaker AI. This post shows how a data scientist explores an Amazon Athena table, validates features in R, trains an XGBoost model in Python, deploys a real-time SageMaker AI endpoint, and reports results with Quarto, all in one governed SageMaker Studio Space.

Reducing medical claims review time with AI on AWS: The EXL Medical IDP solution

Reducing medical claims review time with AI on AWS: The EXL Medical IDP solution

EXL built an AI-powered Medical intelligent document processing (IDP) solution on AWS, combining IDP with domain-specific large language models on Amazon SageMaker and Amazon Bedrock to extract, summarize, and query medical records at enterprise scale and cut claims review time from over 100 minutes per case.