Skip to main content

Amazon SageMaker AI

Model customization on Amazon SageMaker AI

Customize models with your data using the broadest set of techniques. From serverless to dedicated clusters, on infrastructure you trust.

models_ss.png

Teams customizing models on SageMaker AI

Missing alt text value
Missing alt text value
Missing alt text value
Missing alt text value
Missing alt text value

The future belongs to those with custom models.

of enterprise GenAI models will be industry-specific and function-specific by 2027

lower cost per AI query at Stratasys, after fine-tuning a smaller model on SageMaker AI

shorter experimentation cycles at Collinear AI, with no training cluster to stand up

Start customizing a popular model in minutes

Loading
Loading
Loading
Loading
Loading

Every technique, your model

Match the model and the technique to your task, and keep the weights when you are done.

techniques_sc.png

Ship custom models in days, not months

Describe the task. The agent builds the.workflow.

  • Describe your use case and the agent handles data prep, technique choice, training, evaluation, and deployment
  • Export editable notebooks and code, ready to reproduce and automate
  • Install open-source skills in the IDE you already use, including Visual Studio and Cursor
  • Work with the coding agent you prefer, including Kiro, Claude Code, and Copilot
agent_guided.gif

Governance from customization to production

The same controls cover training and serving, so security and evaluation are part of the workflow rather than a later step.

  • Train on data in your own S3 buckets, in the Region you choose, with access governed by IAM
  • Your data and the customized weights are not used to train any other model
  • Built-in evaluation for accuracy, bias, and safety, and CloudWatch observability once it ships
governance_sagemaker_ai.png

Proven in production. Chosen by leaders.

Missing alt text value
Collinear AI
With SageMaker AI's serverless model customization, we cut experimentation cycles from weeks to days, and focus on building better training data, not infrastructure.

Soumyadeep Bakshi

Co-founder
Read the case study
Missing alt text value
Robin AI
With serverless model customization in SageMaker AI, we can rapidly experiment with advanced techniques like reinforcement learning with verifiable rewards in just days.

Diana Mincu

Director of Research
Read the case study
Missing alt text value
Amazon Music
We launch parallel tuning jobs across models like Nova and Qwen with different datasets and hyperparameters, without reserving or managing GPU instances.

Fabian Moerchen

Sr. Principal Applied Scientist

Start with serverless. Step up when the workload asks.

Start here with serverless

Recipe-driven, no MLOps required

  • No instance types to select and no quotas to negotiate
  • Pay per token or per job

Step up to Training Jobs

Full control over containers & frameworks

  • Select your own instance type, billed per second
  • AWS-tested recipes, or bring your own framework and libraries

 

Dedicated clusters

Dedicated clusters

  • Persistent clusters with fault recovery, on EKS or Slurm
  • Recipe-driven distributed training for the largest workloads

 

Common questions about model customization

FAQs

Open all

    Use retrieval when the model needs to ground answers in facts that change often, since you update the source without retraining. Customize when the model itself must absorb your domain's language and rules, which is also what lets you run a much smaller model for the same result. They are not exclusive, and plenty of production systems use both. The signal that you need customization is a retrieval system that is well tuned and still not good enough on a business-critical workflow.

    In SageMaker Studio, open the Models browser, select a model, and choose Customize with UI . Pick your technique, point at your dataset, and run serverless training. You watch metrics as it runs and deploy from the same screen. Or describe your use case in natural language and the agent walks you through the same workflow, then hands back editable notebooks and code so nothing is trapped in a UI. There are no instances to select and no quotas to negotiate. SageMaker provisions GPU capacity from the model size and the job, then releases it when the job finishes.

    Your datasets stay in your own Amazon S3 buckets, in the Region you choose, and access is enforced through IAM. Neither your data nor the resulting weights are used to train or improve any AWS or third-party model. With SageMaker Training Jobs you can also run training inside your VPC with no internet egress. Evaluation for accuracy, bias, and safety is part of the workflow rather than added afterward. This is why teams in financial services and healthcare customize here: Region-level data residency, IAM, CloudWatch audit trails, and HIPAA eligibility are inherited from the account you already run in, rather than rebuilt for a separate tuning vendor.

    You own it. Customizing an open-weight model produces weights you control, and they are never shared or used to improve other models. Deploy to Amazon Bedrock for serverless inference or to SageMaker endpoints when you want dedicated capacity and control over throughput and tail latency. Deployment is part of the same workflow, so there is no format conversion step and no separate migration project between training and serving. You are not locked to a proprietary model or to one serving stack, which matters when you want to renegotiate cost or move to faster hardware later.

    They serve different teams. Bedrock customization is for application developers who want a better model without learning post-training techniques. SageMaker model customization is for ML teams who want to choose the technique, control the data pipeline, evaluate rigorously, and scale the compute. Models customized here can be deployed to Bedrock, so choosing SageMaker does not cut you off from the Bedrock inference experience.

    Serverless customization bills per token or per job with nothing to reserve, so there is no idle GPU time between experiments. Training Jobs bill per second based on the compute you select, and HyperPod is persistent cluster capacity. On the inference side the savings come from running a smaller model for the same result on your task. See the SageMaker AI pricing page for current rates.

Join the teams shipping custom models on SageMaker AI.

Every request you send to a frontier model for a narrow, repetitive task is money spent on capability you are not using. Your team was hired to build models that know your business, not to negotiate GPU quotas.Customize an open-weight model on your data and a smaller model does the job at a fraction of the cost and latency. You train on data in your own account, and the weights are yours.

Did you find what you were looking for today?

Let us know so we can improve the quality of the content on our pages