Model customization on Amazon SageMaker AI
Customize models with your data using the broadest set of techniques. From serverless to dedicated clusters, on infrastructure you trust.
Teams customizing models on SageMaker AI
The future belongs to those with custom models.
lower cost per AI query at Stratasys, after fine-tuning a smaller model on SageMaker AI
shorter experimentation cycles at Collinear AI, with no training cluster to stand up
Start customizing a popular model in minutes
SageMaker AI helps you ship models that know your business
Every technique, your model
Match the model and the technique to your task, and keep the weights when you are done.
- You own the weights, and can deploy them to a SageMaker endpoint
- SageMaker endpointTrain with PyTorch or TensorFlow, and track runs in MLflow or TensorBoard
- Pick the technique that fits the task: supervised fine-tuning, DPO, RLVR or RLAIF, or multi-turn RL for agents. Choose how much of the model to update with LoRA or full fine-tuning, chain techniques together with continuous customization, and run classical or predictive ML in the same place.
Ship custom models in days, not months
Describe the task. The agent builds the.workflow.
- Describe your use case and the agent handles data prep, technique choice, training, evaluation, and deployment
- Export editable notebooks and code, ready to reproduce and automate
- Install open-source skills in the IDE you already use, including Visual Studio and Cursor
- Work with the coding agent you prefer, including Kiro, Claude Code, and Copilot
Governance from customization to production
The same controls cover training and serving, so security and evaluation are part of the workflow rather than a later step.
- Train on data in your own S3 buckets, in the Region you choose, with access governed by IAM
- Your data and the customized weights are not used to train any other model
- Built-in evaluation for accuracy, bias, and safety, and CloudWatch observability once it ships
Proven in production. Chosen by leaders.
With SageMaker AI's serverless model customization, we cut experimentation cycles from weeks to days, and focus on building better training data, not infrastructure.
Soumyadeep Bakshi
Co-founder
With serverless model customization in SageMaker AI, we can rapidly experiment with advanced techniques like reinforcement learning with verifiable rewards in just days.
Diana Mincu
Director of ResearchWe launch parallel tuning jobs across models like Nova and Qwen with different datasets and hyperparameters, without reserving or managing GPU instances.
Fabian Moerchen
Sr. Principal Applied ScientistStart with serverless. Step up when the workload asks.
Start here with serverless
Recipe-driven, no MLOps required
- No instance types to select and no quotas to negotiate
- Pay per token or per job
Step up to Training Jobs
Full control over containers & frameworks
- Select your own instance type, billed per second
- AWS-tested recipes, or bring your own framework and libraries
Dedicated clusters
Dedicated clusters
- Persistent clusters with fault recovery, on EKS or Slurm
- Recipe-driven distributed training for the largest workloads
Common questions about model customization
FAQs
Open allUse retrieval when the model needs to ground answers in facts that change often, since you update the source without retraining. Customize when the model itself must absorb your domain's language and rules, which is also what lets you run a much smaller model for the same result. They are not exclusive, and plenty of production systems use both. The signal that you need customization is a retrieval system that is well tuned and still not good enough on a business-critical workflow.
In SageMaker Studio, open the Models browser, select a model, and choose Customize with UI . Pick your technique, point at your dataset, and run serverless training. You watch metrics as it runs and deploy from the same screen. Or describe your use case in natural language and the agent walks you through the same workflow, then hands back editable notebooks and code so nothing is trapped in a UI. There are no instances to select and no quotas to negotiate. SageMaker provisions GPU capacity from the model size and the job, then releases it when the job finishes.
Your datasets stay in your own Amazon S3 buckets, in the Region you choose, and access is enforced through IAM. Neither your data nor the resulting weights are used to train or improve any AWS or third-party model. With SageMaker Training Jobs you can also run training inside your VPC with no internet egress. Evaluation for accuracy, bias, and safety is part of the workflow rather than added afterward. This is why teams in financial services and healthcare customize here: Region-level data residency, IAM, CloudWatch audit trails, and HIPAA eligibility are inherited from the account you already run in, rather than rebuilt for a separate tuning vendor.
You own it. Customizing an open-weight model produces weights you control, and they are never shared or used to improve other models. Deploy to Amazon Bedrock for serverless inference or to SageMaker endpoints when you want dedicated capacity and control over throughput and tail latency. Deployment is part of the same workflow, so there is no format conversion step and no separate migration project between training and serving. You are not locked to a proprietary model or to one serving stack, which matters when you want to renegotiate cost or move to faster hardware later.
They serve different teams. Bedrock customization is for application developers who want a better model without learning post-training techniques. SageMaker model customization is for ML teams who want to choose the technique, control the data pipeline, evaluate rigorously, and scale the compute. Models customized here can be deployed to Bedrock, so choosing SageMaker does not cut you off from the Bedrock inference experience.
Serverless customization bills per token or per job with nothing to reserve, so there is no idle GPU time between experiments. Training Jobs bill per second based on the compute you select, and HyperPod is persistent cluster capacity. On the inference side the savings come from running a smaller model for the same result on your task. See the SageMaker AI pricing page for current rates.
Start building with model customization
Join the teams shipping custom models on SageMaker AI.
Every request you send to a frontier model for a narrow, repetitive task is money spent on capability you are not using. Your team was hired to build models that know your business, not to negotiate GPU quotas.Customize an open-weight model on your data and a smaller model does the job at a fraction of the cost and latency. You train on data in your own account, and the weights are yours.
Did you find what you were looking for today?
Let us know so we can improve the quality of the content on our pages