Amazon Bedrock AgentCore Pricing
Tailor AgentCore to your needs—mix and match capabilities, use them independently or together, and pay for what you use as your AI initiatives grow.
Pay only for what you use
Amazon Bedrock AgentCore offers flexible, consumption-based pricing with no upfront commitments or minimum fees. Each feature can be used independently or together, and you pay only for what you use. This modular approach allows you to start small and scale as your agent applications grow.
Get started with AgentCore at no cost - the harness in AgentCore is free of extra charge, you pay only for the underlying resources you use. New AWS customers also receive up to $200 in Free Tier credits. Explore AWS Free Tier benefits and start building today.
Note: There is no charge for using AWS Agent Registry during the preview period.
AgentCore features
Open all- No upfront resource selection required
- Billing is calculated per second, using actual CPU consumption and peak memory consumed up to that second, with a 1-second minimum
- You only pay for actual resource consumption during your session, which spans from microVM boot, initialization, active processing, idle periods, until session termination (microVM shutdown)
- Billing includes system overhead in addition to your application's resource usage
- 128MB minimum memory billing applies for memory
- Storage costs: Container deployment requires ECR storage (billed separately). Direct code deployment bills you for the size of the code artifacts you deployed at S3 Standard rates. The managed session storage (currently in public preview) is an early version of the S3-backed storage; we will come out with changes including pricing to that experience before GA.
- Network data transfer charges apply at standard EC2 rates
- The management fee is a percentage of the EC2 On-Demand price.
- Sessions persist up to 14 days, and you can co-locate multiple agents on a single instance.
- The management fee is calculated on the public EC2 On-Demand price.
- Charged per instance-hour from provisioning (boot) until the instance is stopped or terminated, with an approximately 1-minute minimum.
- EC2 compute runs in your account and is billed to you at your EC2 rates. Your EC2 Savings Plans, Reserved Instances, and On-Demand Capacity Reservations apply to the compute — not to the management fee.
- Amazon EBS storage for persistent volumes is billed at standard EBS rates (including while a session is stopped).
- Network data transfer is billed at standard EC2 rates.
AgentCore Runtime is a secure, serverless runtime purpose-built for deploying and scaling agents and tools. Choose between direct code deployment for rapid iteration or container-based deployment for maximum control. AgentCore Runtime offers two compute types: microVMs (serverless, consumption-based) and Instances (EC2 instance cost plus a management fee).
microVMs: you only pay for the active resources you consume
Unlike traditional compute services that charge for pre-allocated resources (i.e., fixed instance size and cost per second while hosting the agent), with AgentCore Runtime you only pay for active resource consumption. This delivers substantial cost savings for agentic workloads, which typically spend 30-70% of time in I/O wait (waiting for LLM responses, tool / API calls, or database queries). With pre-allocated pricing, you would pay for idle CPU during these wait periods. With the active resource consumption-based pricing in AgentCore Runtime, I/O wait and idle time is free, if no other background process is running.
Billing is based on CPU and memory consumption across your session lifetime, calculated at per-second increments. For CPU resources, you are charged based on actual consumption - if your agent consumes no CPU during I/O wait, there are no CPU charges. For memory resources, you're charged for the peak memory consumed up to that second.
Key details:
Instances: pay EC2 instance cost plus a management fee
For persistent, resource-intensive, or specialized agent workloads, AgentCore Runtime also offers the Instances compute type, which runs your agents on AWS-managed Amazon EC2 instances in your own account. Instances bill on the underlying EC2 instance for the time it runs, plus an AgentCore management fee.
Key details:
- No upfront resource selection required
- Billing is calculated per second, using actual CPU consumption and peak memory consumed up to that second, with a 1-second minimum
- You only pay for actual resource consumption during your session, which spans from microVM boot, initialization, active processing, idle periods, until session termination (microVM shutdown)
- Billing includes system overhead in addition to your application's resource usage
- 128MB minimum memory billing applies for memory
- Storage costs: Browser Profiles require Amazon S3 storage for storing profile artifacts (cookies, local storage) and you will get billed at Amazon S3 Standard rates, starting April 15, 2026.
- Network data transfer charges apply at standard EC2 rates.
AgentCore Browser provides a fast, secure, cloud-based browser runtime to enable agents to interact with websites at scale.
You only pay for the active resources you consume
Unlike traditional compute services that charge for pre-allocated resources (i.e., fixed instance size and cost per second while hosting the agent), with AgentCore Browser you only pay for active resource consumption. This delivers substantial cost savings for agentic workloads, which typically spend 30-70% of time in I/O wait (waiting for LLM responses, tool / API calls, or database queries). With pre-allocated pricing, you would pay for idle CPU during these wait periods. With the active resource consumption-based pricing in AgentCore Browser, I/O wait and idle time is free, if no other background process is running.
Billing is based on CPU and memory consumption across your session lifetime, calculated at per-second increments. For CPU resources, you are charged based on actual consumption - if your agent consumes no CPU during I/O wait, there are no CPU charges. For memory resources, you're charged for the peak memory consumed up to that second.
Key details:
- No upfront resource selection required
- Billing is calculated per second, using actual CPU consumption and peak memory consumed up to that second, with a 1-second minimum
- You only pay for actual resource consumption during your session, which spans from microVM boot, initialization, active processing, idle periods, until session termination (microVM shutdown)
- Billing includes system overhead in addition to your application's resource usage
- 128MB minimum memory billing applies for memory
- Network data transfer charges apply at standard EC2 rates
AgentCore Code Interpreter enables agents execute code securely in sandbox environments, enhancing their accuracy and expanding their ability to solve complex end-to-end tasks.
You only pay for the active resources you consume
Unlike traditional compute services that charge for pre-allocated resources (i.e., fixed instance size and cost per second while hosting the agent), with AgentCore Code Interpreter you only pay for active resource consumption. This delivers substantial cost savings for agentic workloads, which typically spend 30-70% of time in I/O wait (waiting for LLM responses, tool / API calls, or database queries). With pre-allocated pricing, you would pay for idle CPU during these wait periods. With the active resource consumption-based pricing in AgentCore Code Interpreter, I/O wait and idle time is free, if no other background process is running.
Billing is based on CPU and memory consumption across your session lifetime, calculated at per-second increments. For CPU resources, you are charged based on actual consumption - if your agent consumes no CPU during I/O wait, there are no CPU charges. For memory resources, you're charged for the peak memory consumed up to that second.
Key details:
- No minimum fees and no upfront commitments
- Billing is calculated per query submitted to the web search
- You pay only for the queries your agents make, scaling directly with usage
Web Search on AgentCore gives your agents real-time access to current information from the web, enabling them to answer questions, research topics, and ground their responses in up-to-date sources without building and maintaining your own search infrastructure.
Consumption-based pricing
You only pay for what you use. Pricing is simple and usage-based — you are charged based on the number of search queries your agents submit to the web search. There are no upfront commitments, and you scale costs directly with how often your agents need to retrieve information from the web.
Web Search is priced at $7 per 1,000 queries.
Key details:
- No upfront costs or minimum commitments required
- Network data transfer charges apply at standard EC2 rates
- Data egress to customer-owned VPCs is charged a data processing rate of $0.006 per GB in commercial AWS Regions
- Use of Web Search and Bedrock Managed Knowledge Base are charged separately in accordance with the rates defined or linked in the pricing table
- There are no extra charges for customer-defined rate limiting
Amazon Bedrock AgentCore Gateway enables agents to securely access tools, models, and other agents.
Consumption-based pricing
You pay only for the API calls made to Gateway. You're charged based on the number of operations (such as listing available actions, invoking actions, and health checks), search queries, and targets indexed for semantic search functionality.
Key details:
AgentCore policy gives you comprehensive control over actions agents take, helping ensure agents stay within defined boundaries without slowing down.
Consumption-based pricing
You only pay for the authorization requests performed during agent execution. Each time an agent calls a tool through AgentCore Gateway, Policy checks the action against your rules to determine whether it is allowed or denied.
In addition, Policy offers natural language policy authoring, which lets you create policies using simple natural language descriptions. You are charged per one million user input tokens processed when converting natural language into Cedar policy statements. If you use guardrails through policy, you are charged based on the usage of each safeguard (see the Amazon Bedrock pricing page for details).
- No minimum fees and no upfront commitments
- Billing is calculated per successful OAuth token or API key requested to perform a task requiring authorization for a non-AWS resource
- No additional charges incurred when customers use AgentCore Identity through AgentCore Runtime or AgentCore Gateway
AgentCore Identity simplifies agent identity and access management and allows your agents to securely access AWS resources and third-party tools and services on behalf of users or by themselves with pre-authorized user consent.
Consumption-based pricing
Customers who use AgentCore Identity through either AgentCore Runtime or AgentCore Gateway, do not incur any additional charges for their use of AgentCore Identity. For all other scenarios, you pay for only what you use and are charged based on the number of requests from the agent to AgentCore Identity for an OAuth token or an API key.
Key details:
- Short-term memory is priced based on the number of raw events created, giving you predictable costs for in-session context.
- Long-term memory records is priced based on the number of memories processed and stored each month and the number of memory record retrieval calls, so you only pay when your agents store and use processed knowledge.
- To extract long-term memory from raw events, you can choose between built-in memory strategies, which include automatic processing, or more configurable memory strategies that run in your account using your choice of model and prompt.
- No upfront resource selection required
- For short-term memory, billing is calculated per create event request
- For long-term memory storage, billing is calculated per stored memory record and billed hourly
- For long-term memory retrieval, billing is calculated per retrieve memory request
AgentCore Memory makes it easy for developers to build context-aware agents by eliminating complex memory infrastructure management while providing full control over what the agent remembers.
Consumption-based pricing
You only pay for what you use. Our pricing is simple and usage-based, aligning directly with how your agents create value:
Key details:
AgentCore Observability gives developers complete visibility into agent workflows to trace, debug, and monitor agents' performance in production environments.
Consumption-based pricing
You pay as you go for telemetry generated, stored, and queried for your agents. The telemetry data is ingested and stored in your Amazon CloudWatch account. You are charged for data ingestion and storage, queries to retrieve and analyze information, and masking of sensitive/Personally Identifiable Information (PII) data in logs. To review pricing details visit Amazon CloudWatch pricing page.
- No upfront commitments or minimum fees required
- Includes CI/CD integration with configurable quality thresholds
- Production monitoring with sampling rules and dashboard aggregation
- Cost control through percentage-based sampling, conditional sampling, and selective metric monitoring
- Model usage costs are included for built-in evaluators — no separate model charges
- Custom evaluations incur additional model usage charges in your account
- Evaluations rates apply when used through AgentCore optimization batch evaluations
AgentCore Evaluations helps continuously inspect agent quality based on real-world behavior. Teams can perform agentic evaluations using 13 built-in evaluators on common quality dimensions or create custom evaluators for specific business requirements. The results are integrated into AgentCore Observability powered by Amazon CloudWatch for unified monitoring. Evaluations also powers the batch evaluation and A/B testing capabilities in AgentCore optimization.
Consumption-based pricing You pay for what you use. For built-in evaluators, pricing is charged by AgentCore based on input and output tokens processed during evaluation. For custom evaluations using your own LLM infrastructure, you pay per evaluation performed, with separate inference costs based on the model used.
Key details:
- No upfront commitments or minimum fees required
- Insights is free during public preview; pricing will be announced before general availability
- Insights analyzes OpenTelemetry-compatible traces stored in CloudWatch; standard CloudWatch charges apply for underlying telemetry ingestion and storage
- Recommendation generation is included at no charge — you pay only for any Evaluations consumed as part of the workflow, billed at applicable AgentCore Evaluations rates
- Batch evaluations are charged at a 25% discount on standard AgentCore Evaluations rates
- A/B tests are charged based on Gateway, Runtime, and Evaluations consumed during the experiment
AgentCore optimization makes it easy for developers to continuously improve production agents by automating the work of identifying failures, analyzing root causes, generating better configurations, and validating improvements before they reach users. Optimization works through three capabilities: Insights, which delivers rich failure, intent, and trajectory analysis across hundreds of sessions, surfacing patterns no dashboard or one-at-a-time trace review would reveal; Recommendations, which analyze production traces and generate optimized agent configurations (currently system prompts and tool descriptions); and Experiments, which validate those configurations through batch evaluations run offline against a test dataset, or A/B tests run on live production traffic.
Consumption-based pricing Insights is available at no charge during public preview; pricing will be announced before general availability. Recommendations are available at no charge; pay only for any Evaluations consumed as part of the workflow. Experiments are priced at standard AgentCore rates based on the underlying services consumed (Evaluations, Gateway, Runtime).
Key details:
- No upfront commitments or minimum fees required
- You pay only for what you use in terms of Records added in the Registry, and Search, List and Get calls made on Records in the Registry
- For Records, you are only charged for ‘Net Records’ in the Registry present at any given point of time. If you add and then delete a record, it no longer counts towards your Net total number of records
The AWS Agent Registry provides a centralized catalog for organizing, curating, and discovering resources across your organization. With registry, you can publish MCP servers, agents, agent skills, and custom resources into a searchable registry, control access through an approval workflow, and enable both human users and AI agents to discover the right tools and agents using semantic and keyword search.
Consumption based Pricing with Free Tier
You only pay for what you use. Pricing for registry is based on the number of records you have added into your registry, and the number of API calls (Search, List, and Get) you make to discover resources added into your registry. Registry comes with a Free Tier, where every month your first 5,000 records, first 1,000,000 Search API calls, and first 2,000,000 combined Get and List API calls are free of charge. You are charged only for usage above these thresholds.
Key details:
- No upfront commitments or minimum fees required
- Only charged for underlying 3P wallet (Coinbase and Stripe) operations performed through 2 APIs: CreateInstrument, and ProcessPayment,
- Coinbase CDP Wallet: 1 CreateInstrument API invocation = 1 wallet operation fee; 1 ProcessPayment invocation = 1 wallet operation fee
- Stripe Privy Wallet: 1 CreateInstrument API invocation = No charge; 1 ProcessPayment invocation = 1 wallet operation fee
- There are no additional charges for API invocations beyond the standard wallet operation fees charged by your chosen provider.
Amazon Bedrock AgentCore payments is a fully managed service that enables microtransaction payments in AI agents to access paid APIs, MCP servers, and content. AgentCore payments provides a suite of developer-friendly capabilities that enable secure, instant payments to paid services with stablecoin support, open protocols like x402 for cost-effective microtransactions, and configurable guardrails to control agent spending, reducing developer effort from months to days.
Consumption based Pricing
You only pay for what you use. To get started with AgentCore payments, you'll need to select a wallet provider, either Coinbase CDP or Stripe Privy, to fund your agent for making payments. The developer would be charged only for two APIs: CreateInstrument, which generates a unique wallet address for the end user, and ProcessPayment, which signs the transaction to execute the payment; these map directly to underlying Coinbase and Stripe wallet operations (wallet creation and transaction signing, respectively). There are no additional charges for API invocations beyond the standard wallet operation fees charged by your chosen provider. With Coinbase CDP, each CreateInstrument and each ProcessPayment invocation counts as one wallet operation, with pricing per the Coinbase CDP pricing page. With Stripe Privy, CreateInstrument invocations carry no charge, while each ProcessPayment invocation counts as one wallet operation, with pricing per the Privy pricing page.
Key details:
Pricing Table
| Service/Capability | Type | Resource | Price |
| Runtime microVMs | Active Consumption Based | CPU | $0.0895 per vCPU-hour |
| Active Consumption Based | Memory | $0.00945 per GB-hour | |
| Runtime instances | Instance-based | Managed instance | See Runtime instances pricing below. |
| Browser Tool | Active Consumption Based | CPU | $0.0895 per vCPU-hour |
| Active Consumption Based | Memory | $0.00945 per GB-hour | |
| Code Interpreter | Active Consumption Based | CPU | $0.0895 per vCPU-hour |
| Active Consumption Based | Memory | $0.00945 per GB-hour | |
| Web Search on AgentCore |
Consumption Based
|
Queries
|
$7.00 per 1,000 queries
|
| Bedrock Managed Knowledge Base | Consumption Based | Storage | Please see the Amazon
Bedrock pricing page for details. |
| Consumption Based | Queries | Please see the Amazon
Bedrock pricing page for details. |
|
| Gateway | Consumption Based | API Invocations (e.g. ListTools, InvokeTool, Ping) | $0.005 per 1,000 invocations |
| Consumption Based | Search API | $0.025 per 1,000 invocations | |
| Consumption Based | Tool Indexing | $0.02 per 100 tools indexed per month | |
| Identity | Consumption Based | Token or API key requests for non-AWS resources |
$0.010 per 1,000 token or API keys requested by the agent (Note: AgentCore Identity is available at no additional charge to customers when they use it through either AgentCore Runtime or AgentCore Gateway) |
| Memory | Consumption Based | Short-Term Memory | $0.25 per 1,000 new events |
| Consumption Based | Long-Term Memory Storage |
Using built-in memory strategies: $0.75 per 1000 memory records stored per month Using built-in with override or self-managed memory strategies: $0.25 per 1000 memory records stored per month* Note: Billed hourly assuming a 31 day month |
|
| Consumption Based | Long-Term Memory Retrieval | $0.50 per 1000 memory record retrievals | |
| Observability | Consumption Based | Spans, logs, and metrics | Charged as per Amazon CloudWatch pricing |
| Evaluations | Consumption Based | Built-in evaluators - input tokens | $0.0024 per 1000 input tokens |
| Consumption Based | Built-in evaluators - output tokens | $0.012 per 1,000 output tokens | |
| Consumption Based | Custom evaluators | $1.50 per 1,000 evaluations (model usage billed separately) | |
| Consumption Based | Batch evaluations |
$0.0018 per 1,000 input token $0.009 per 1,000 output tokens |
|
| Optimization | Consumption Based | Insights (Preview) | Free during public preview; pricing will be announced before general availability |
| Consumption Based | Recommendations | Recommendation generation is free; pay for any Evaluations consumed as part of workflow | |
| Consumption Based | A/B Tests | Pay for underlying AgentCore resources consumed | |
| Policy | Consumption Based | Authorization Request |
$0.000025 per request Your first 100 temporal policies per policy engine incur no additional authorization charges |
| Consumption Based | Input Tokens Processed | $0.13 per 1,000 tokens | |
| Consumption Based | Guardrail safeguard usage | Pricing for Amazon Bedrock Guardrails is based on the usage of each safeguard. Please see the Amazon Bedrock pricing page for details. | |
| AWS Agent Registry (Preview) | Consumption Based | Registry Records | First 5,000 records free monthly, then $0.400 per 1,000 records |
| Consumption Based | Search API Invocation | First 1,000,000 invocations free monthly, then $0.020 per 1,000 invocations | |
| Consumption Based | List and Get API Invocations | First 2,000,000 combined invocations free monthly, then $0.004 per 1,000 invocations | |
| Payments | Consumption Based | CreateInstrument and ProcessPayment API invocations |
Coinbase CDP Wallet: Wallet operations performed through the AgentCore Payments API are charged per Coinbase's published pricing:
Stripe Privy Wallet: Wallet operations performed through the AgentCore Payments API are charged per Privy's published pricing:
|
| Consumption Based | Rest of the API invocations | No charge |
Pricing Examples
Open allExample: Customer Support Agent Deployment
You plan to deploy a customer support agent that resolves user queries across chat and email. The agent handles order issues, account verification, and policy clarifications. It uses retrieval-augmented generation (RAG) to fetch product policies, and Model Context Protocol (MCP)-compatible tools to query order status and update support tickets. Each agent session involves sophisticated multi-step reasoning with 1 RAG call to a vector store, 2 MCP tool calls (e.g., OrderAPI, TicketAPI), and 2 LLM reasoning steps. You deployed your agent on AgentCore Runtime because you require complete session isolation and the flexibility to scale to thousands of sessions in seconds.
Processing 10M user requests monthly, each session runs for 60 seconds with 70% I/O wait time (waiting for LLM responses and API calls), and no other background process is running during I/O. Each agent session utilizes 1vCPU during active processing. Memory usage starts at 1GB during initialization, increases to 2GB during RAG processing, then peaks at 2.5GB during complex tool calls for the remainder of the session. Your monthly costs break down as follows:
CPU cost per session: 18 seconds (only active processing time) × 1vCPU × ($0.0895/3600) = $0.0004475
Memory cost per session: 10 seconds × 1GB × ($0.00945/3600) + 20 seconds × 2GB × ($0.00945/3600) + 30 seconds × 2.5GB × ($0.00945/3600) = $0.000026 + $0.000053 + $0.000197 = $0.000276
Total cost per session: $0.0007235
Monthly total: 10M sessions × $0.0007235 = $7,235
Storage costs: With container-based deployment, you manage ECR storage separately based on published ECR rates. If you used direct code deployment instead, S3 Standard pricing (starting February 27, 2026) would apply for your code artifacts - for a 100MB agent, this adds up to $0.0023/month in storage costs.
Comparison to pre-allocated pricing: If this workload ran on a service charging for pre-allocated resources, you would need to provision for peak usage (2.5GB memory, 1vCPU) for the entire 60-second duration, resulting in up to 3.3x higher CPU costs and up to 1.4x higher memory costs. In the example above, the CPU cost savings are significant, given that CPU consumption accounted for 62% of the total workload cost. AgentCore Runtime's consumption-based pricing automatically captures the "I/O wait is free" benefit and charges only for actual memory usage as it scales up, delivering substantial cost efficiency compared to traditional compute options.
Example: Multi-agent code-modernization workload
You run a code-modernization pipeline where several collaborating agents — a code writer, a reviewer, and a test runner — share a filesystem and work on the same repository within a single persistent session. Because each job shallow-clones a large repository, holds it in memory, and runs for hours across build-and-test cycles, you choose the Instances compute type. You configure a capacity provider using c7g.2xlarge instances (8 vCPU, 16 GB) in US East (N. Virginia), and you run 1,000 modernization jobs per month, each completing in a single 3-hour session.
Runtime instances bill the underlying EC2 instance for the session duration plus a 12% management fee. Your monthly costs break down as follows:
EC2 cost per session: 3 hours × $0.289/hour (c7g.2xlarge, On-Demand) = $0.867
AgentCore management fee (12% of the On-Demand rate): $0.867 × 0.12 = $0.10404
Total cost per session: $0.97104
Monthly total: 1,000 sessions × $0.97104 = $971.04
Storage and data transfer: Persistent EBS volumes attached to your sessions are billed separately at standard Amazon EBS rates, and network data transfer applies at standard EC2 rates.
Using your EC2 pricing agreements: Because the instances run in your account, EC2 discounts such as Savings Plans, Reserved Instances, and On-Demand Capacity Reservations (ODCRs) apply to the EC2 compute portion. The management fee is always calculated on the On-Demand rate.
GPU instances: G-series families (including Graviton-based gr6) carry a reduced management fee of 7.8%.
Example: Automated Travel Booking System
You plan to create a travel booking agent that automates full trip planning and booking through web interactions. Your implementation requires AgentCore Browser's secure, serverless runtime to dynamically manage headless browsers for searching flights, hotels, simulating clicks, extracting prices, and submitting booking forms. AgentCore Browser tool provides enterprise-grade capabilities including session-isolated sandbox compute and comprehensive observability through Live View and Session Replay.
The agent processes 100K monthly requests. Each browser session runs for 10 minutes with 80% I/O wait time. During active processing it utilizes 2vCPU and 4GB memory continuously, and during I/O it is utilizing 0.4vCPU and 5GB memory. Your monthly costs break down as follows:
CPU cost per session: 120 seconds (adjusting for 80% I/O wait) × 2 vCPU × ($0.0895/3600) = $0.005967
Memory cost per session: 600 seconds × 4GB × ($0.00945/3600) = $0.0063
Total cost per session: $0.012267
Monthly total: 100K sessions × $0.012267 = $1,226.67
Example: Natural Language Data Analysis Automation
You plan to deploy a data analyst agent that supports business and product teams with dataset queries, visualizations, and statistical analysis—all through natural language. Your agent dynamically generates and executes Python code for complex requests like correlation analysis between site traffic and conversion rates. You leverage AgentCore Code Interpreter because it provides isolated sandbox environments compliant with enterprise security policies, pre-built execution runtimes for multiple languages (JavaScript, TypeScript, Python), and large file size support.
The agent processes 10K monthly requests with 3 code executions per request. Each execution runs for 2 minutes with 60% I/O wait time, utilizing 2vCPU during active processing and 4GB memory continuously. Your monthly costs break down as follows:
CPU cost per session: 48 seconds (adjusting for 60% I/O wait) × 2 vCPU × ($0.0895/3600) = $0.002387
Memory cost per session: 120 seconds × 4GB × ($0.00945/3600) = $0.00126
Total cost per session: $0.003647
Monthly total: 30K executions × $0.003647 = $109.40
Example: Grounding a research agent with real-time web search
You are building a market research agent that helps your sales team stay informed about prospects, industry trends, and competitor activity. The agent uses the Web Search through AgentCore Gateway to retrieve current information from the web, ensuring responses are grounded in up-to-date sources rather than stale training data. Your sales team generates 200,000 research queries per month across 500 active users. Each query triggers 2 additional tool calls on average, resulting in 400,000 additional tool calls.
Web Search charges: 200,000 x $7/1K Queries = $1,400.00
InvokeTool API charges: 600,000 x $5/million = $3.00
Monthly total: $1,403.00
Example: Connecting HR Assistant agent to internal tools
You plan to build an HR assistant agent for a mid-sized enterprise, handling internal policy questions, leave balances, benefits enrollment, and payroll inquiries. To serve the user requests, the agent needs to access multiple internal systems (Onboarding, Benefits, Payroll, and Leave Management APIs) as tools. You used AgentCore Gateway to create MCP servers for 200 internal tools that your agent can interact with from anywhere, all without writing any code. To improve tool use accuracy, you leveraged the search capability to index tool metadata and enable dynamic matching of tools during agent invocation based on interaction context.
Each agent interaction requires 1 Search API and 4 InvokeTool API invocations. 50M monthly interactions result in 50M Search and 200M InvokeTool calls. Your monthly costs break down as follows:
SearchToolIndex charges: 200 tools × $0.02 per 100 tools = $0.04
Search API charges: 50M × $25/million = $1,250
InvokeTool API charges: 200M × $5/million = $1,000
Monthly total: $2,250.04
Example:
You plan to develop a procurement automation agent that helps operations teams manage vendor selection, purchase order creation, and invoice approvals. To ensure actions follow defined business rules, you use Policy with AgentCore Gateway tools to automatically verify every action before it executes against your defined policies. Each time the agent attempts to perform an action (for example, sending purchase approval or initiating a payment), Gateway intercepts the tool call to check whether the action is allowed or denied. Let’s assume the agent serves 100K sessions in a month and makes 5 tool calls on average in every session. If you implement one authorization request for each tool call, you make 500K authorization requests per month with a cost break down as follow:
Authorization Requests = 100K sessions x 5 tool calls/session x 1 policy enforced/ tool call = 500K authorization requests
Monthly Total: 500K requests x 0.000025 = $12.50
Before deployment, your team optionally uses natural language policy authoring to simplify onboarding and policy setup. Instead of writing Cedar policies manually, they describe rules in plain language and AgentCore converts them into Cedar policy statements. You are charged a one-time fee based on the number of user input tokens processed during this authoring step. If your team used 10,000 tokens to author several policies, your costs would be:
Policy Authoring = 20K tokens × $0.13 per 1K input tokens = $2.60
Note: Standard CloudWatch rates apply if Observability is enabled.
Example: Secure Customer Support Access Management
You plan to operate a customer support agent that assists technical teams by accessing multiple tools—Slack for support conversations, Zoom to fetch call logs, and GitHub for issue tracking and commit logs. Your implementation uses AgentCore Identity for secure, delegated access for users or support engineers. The system is compatible with existing identity providers ( e.g. Amazon Cognito, Okta, Microsoft Entra ID) and manages all authentication methods from OAuth tokens to API keys, eliminating the need for custom security infrastructure.
Lets assume the agent is being used by 10K monthly active users averaging 5 interactions each, requiring 3 tool accesses per session for each user per month, your monthly costs break down as follows:
Total tokens requested: 10K users × 5 sessions × 3 tools = 150K tokens
Monthly total: 150K requests × $0.010/1,000 = $1.50
Note: AgentCore Identity is included at no additional cost when using AgentCore Runtime or Gateway.
Example: Personalized Coding Assistant Agent Implementation
You plan to develop a coding assistant agent that helps software engineers write, debug, and refactor code across IDEs and terminals. To provide a personalized experience, the agent needs to maintain context during a session and remember user preferences over multiple sessions. Your implementation uses AgentCore Memory for equipping the agent with both short-term memory (immediate conversations and events) and long-term memory (persistent knowledge across sessions).
Each time a user interacts with the agent (e.g., by sending a code snippet or asking a coding question), you send an event to AgentCore Memory for storing it as short-term memory. For long-term memory, you configured built-in extraction strategies to automatically extract and store summarization of debugging sessions and user preferences across sessions. The agent can then retrieve these long-term memories to provide a personalized experience for developers.
With 100,000 monthly short-term memory events, 10,000 stored long-term memory records, and 20,000 monthly memory record retrieval calls, your costs break down as follows:
Short-term memory: 100,000 events × $0.25/1,000 = $25
Long-term memory storage: 10,000 memories × $0.75/1,000 = $7.50
Long-term memory retrieval: 20,000 retrievals × $0.50/1,000 = $10
Monthly total: $42.50
Note: With built-in with override extraction strategies, long-term storage cost would be lower at $0.25 per 1,000 memories stored. However, you may incur additional charges for model usage in your account.
Example: Multi-Agent Financial Advisory Platform
Note: Most development environment observability data volumes are low enough that observability costs are near zero.
You deploy a comprehensive financial advisory platform with multiple specialized agents handling investment research, portfolio analysis, and regulatory compliance checks. Each agent performs complex, multi-step reasoning involving database and web-search queries, API calls to financial data providers, and document analysis. The platform processes thousands of agent calls and generates extensive telemetry data—including traces, metrics, and logs—across all agent interactions. You use AgentCore Observability to monitor performance, debug issues, and ensure compliance with financial regulations through comprehensive audit trails.
In production, your platform has 200,000 agent invocations per month, generating 10 GB of observability data from approximately 10 million spans (assuming 1 KB per span). This includes agent interactions, API calls, and system events. Of those spans, 30% include event logs (input/output for model invocations and tool calls), resulting in approximately 6 GB of additional log data (assuming 2 KB per span event) written to CloudWatch standard logs. Your monthly costs break down as follows:
Monthly Span Ingestion charges: 10 GB × $0.35/GB = $3.50
Monthly Event Logging charges: 6 GB × $0.50/GB = $3.00
Monthly total: $3.50 + $3.00 = $6.50
*Standard CloudWatch rates apply for any metrics and non-telemetry (standard) log data sent to CloudWatch.
- Total interactions evaluated: 15,000 (5,000 development + 10,000 production)
- Built-in evaluators per interaction: 3 (correctness, helpfulness, goal success rate)
- Total evaluation: 15,000 interactions × 3 evaluators = 45,000 evaluations
- Input tokens: 45,000 evaluations × 15,000 tokens = 675M tokens
- Output tokens: 45,000 evaluations × 300 tokens = 13.5M tokens
- Input cost: 675M tokens × $2.40/1M = $1,620
- Output cost: 13.5M tokens × $12.00/1M = $162
- Built-in evaluators subtotal: $1,782
- Total custom evaluations: 15,000 interactions × 1 custom evaluator = 15,000 evaluations
- Custom evaluation cost: 15,000 evaluations × $1.50/1,000 = $22.50
Example: E-commerce Customer Service Agent Quality Monitoring
You plan to deploy a customer service agent that handles order inquiries, returns processing, and product recommendations for an e-commerce platform. To ensure consistent service quality, you use AgentCore Evaluations to monitor agent performance across development and production environments. Your implementation uses 3 built-in trace-level evaluators (correctness, helpfulness, and goal success rate) plus 1 custom evaluator for business-specific quality metrics.
During development, your CI/CD pipeline evaluates 5,000 test interactions monthly. In production, you monitor 2% of live interactions through sampling rules, evaluating 10,000 customer conversations monthly. Each built-in evaluation processes an average of 15,000 input tokens (including conversation history, product catalogs, and order details) and generates 300 output tokens for scoring.
Your monthly costs break down as follows:
Built-in Evaluators:
Custom Evaluations:
Monthly total: $1,804.50
Note: Model usage costs are included for built-in evaluators. Custom evaluations incur additional model usage charges in your account.
- 1 system prompt recommendation: $0
- Total evaluations: 1,000 interactions × 3 evaluators = 3,000 evaluations
- Input tokens: 3,000 × 15,000 = 45M tokens × $0.0018/1,000 = $81.00
- Output tokens: 3,000 × 300 = 900K tokens × $0.009/1,000 = $8.10
- Batch evaluations subtotal: $89.10
- 400,000 sessions × 1.5 tool calls = 600,000 InvokeTool calls × $0.005/1,000 = $3.00
Example: Optimizing a Production Health Coach Agent
You have deployed a health coach agent handling 200,000 active users. After several weeks in production, you notice the agent's response quality has plateaued. You use AgentCore optimization to analyze production traces, generate a system prompt recommendation, and validate the improvement before rolling it out.
You run 1 system prompt recommendation and validate it with a batch evaluation across 1,000 test interactions using 3 built-in evaluators (correctness, helpfulness, and goal success rate). Each evaluation processes an average of 15,000 input tokens and generates 300 output tokens. After confirming the improvement, you run an A/B test on live traffic for 2 weeks, splitting 10% of sessions (approximately 400,000 sessions) between the current and optimized prompt. Each session invokes 1.5 Gateway tool calls on average.
Recommendations:
Batch Evaluations:
A/B Test (Gateway consumption):
Total: $92.10
Note: Runtime and Evaluations consumed during the A/B test are charged at standard AgentCore rates.
- First Month:
- Records: 2,000 MCP servers (full month) + 7,500 pro-rated records (10,000 agents + 5,000 tools added mid-month) = 9,500 effective records. With 5,000 records free, you're charged for 4,500 records at $0.40/1,000 = $1.80
- Search API calls: 300,000 calls (5,000 builders × 2 calls/day × 30 days). Since it is within the 1M free tier, the cost is = $0
- List + Get API calls: 750,000 calls (5,000 builders × 5 calls/day × 30 days). Since it is within the 2M free tier, the cost is = $0
- Monthly total: $1.80
- Second Month:
- Records: 17,000 total records - 5,000 free = 12,000 billable records at $0.40/1,000 = $4.80
- Search API calls: 3,000,000 calls (5,000 builders × 20 calls/day × 30 days) - 1M free = 2M billable calls at $0.020/1,000 = $40.00
- List + Get API calls: 7,500,000 calls (5,000 builders × 50 calls/day × 30 days) - 2M free = 5.5M billable calls at $0.004/1,000 = $22.00
- Monthly total: $66.80
Example: Enterprise AI Tool Marketplace for a Financial Services Firm
You work at a large financial services firm with 5,000+ developers across trading, risk management, compliance, and client services teams. Each team builds and maintains MCP servers, agents, and custom tools, but there's no centralized place to discover existing resources. You leverage AWS Agent Registry as your centralized catalog where teams publish their tools.
You create a single registry where developers across 12 teams register their resources—totaling 17,000 registry records: 2,000 MCP servers (market data feeds, order management, risk calculators, compliance checkers), 10,000 agents (portfolio rebalancing, trade reconciliation, regulatory reporting), and 5,000 custom tools (internal APIs, proprietary models, knowledge base). The 2,000 MCP servers are added at the beginning of the month, while the 10,000 agents and 5,000 custom tools are added mid-month.
On the consumer side, all 5,000 builders query the registry throughout the month. In the first month, builders average 2 Search calls and 5 List + Get calls per day as they familiarize themselves with the system. By the second month, usage increases to 20 Search calls and 50 List + Get calls per day as adoption grows.
- CreateInstrument calls: 200 agents × 1 call = 200 wallet operations at $0.005/operation = $1.00
- ProcessPayment calls: 270,000 calls (200 agents × 15 requests/day × 3 payments/request × 30 days) at $0.005/operation = $1,350.00
- Microtransaction spend (paid to vendors): 270,000 payments × $2.50 avg per request / 3 payments per request = $675,000
- Monthly AgentCore Payments total: $1,351.00
- CreateInstrument calls: 0 calls (wallets already created) = $0
- ProcessPayment calls: 720,000 calls (200 agents × 40 requests/day × 3 payments/request × 30 days) at $0.005/operation = $3,600.00
- Microtransaction spend (paid to vendors): 720,000 payments × $2.50 avg per request / 3 payments per request = $1,800,000
- Monthly AgentCore Payments total: $3,600.00
Example: AI Agent for Financial Services Firm
A financial services firm deploys 200 AI agents through Amazon Bedrock to assist analysts with evaluating financial instruments such as equities, derivatives, fixed income, and structured products. Each analyst makes dozens of data requests daily that require accessing premium, pay-per-use data services from third-party vendors. These requests involve microtransactions — typically ranging from $0.10 to $3.00 per call — across multiple providers. The firm uses AgentCore payments with Coinbase CDP or Stripe Privy Wallet to enable these agents to make secure, instant micropayments to paid services without requiring per-vendor payment integrations.
An equities research analyst asks their AI agent: "Analyze Amazon's (AMZN) Q1 2026 earnings performance, evaluate its current options chain, and compare valuation multiples against mega-cap tech peers." To fulfill this single request, the agent makes three micropayments via ProcessPayment:
Monthly Cost Breakdown
AgentCore payments is provided at no additional charge by AWS. Wallet operations performed through the AgentCore Payments API are charged by the wallet provider (e.g., Coinbase CDP) per their published pricing at $0.005 per wallet operation.
First Month (Pilot — Equities & Fixed Income):
Second Month (Full Rollout — all divisions):