OpenCall achieves 190 ms voice AI latency with Tech 42 on AWS
Learn how Tech 42 helped OpenCall avoid vendor lock-in and achieve about 190 ms latency by self-hosting on AWS.
Benefits
text-to-speech requests per hour at peak
night to conduct a 5,000-patient campaign
Overview
OpenCall helps healthcare systems automate inbound and outbound patient calls and texts, including scheduling, call routing, and urgent outreach. As customer demand increased, OpenCall decided to use Amazon Web Services (AWS) to gain greater control over its voice AI solution while maintaining reliable call experiences. So the company worked alongside AWS Partner Tech 42 to deploy self-hosted speech technology, achieving about 190 ms latency while supporting high-volume calling.
About OpenCall
Based in the United States, voice AI company OpenCall builds intelligent phone agents for healthcare organizations.
Opportunity | Overcoming barriers for OpenCall’s voice AI solution
OpenCall had built its solution with deep dependencies on third-party vendors. Although the solution performed well, the leadership recognized that relying entirely on external providers for the business’s core voice capabilities created operational risk. Any service disruption, pricing change, or quality degradation could directly impact the company’s ability to serve its customers. Additionally, the company had no control over costs as its customer base increased. “We were happy with our previous vendor, and its product worked well for us, but we knew it was going to be a growing expense,” says Arthur Silverstein, cofounder of OpenCall. “Plus, we didn’t want to be tied to one vendor.”
The company stopped working with one vendor after experiencing quality issues and understood how quickly a dependency could become a liability. Beyond vendor lock-in risk, OpenCall faced a technical challenge: latency. Near real-time voice agents are extraordinarily sensitive to response times. To maintain natural, seamless conversational experiences, the company needed the first output from its text-to-speech system delivered in under 300 ms—ideally, closer to 150–200 ms. OpenCall had previously tried to self-host AI models, including on AI-optimized GPUs, but couldn’t achieve the required speed without raising quality concerns.
Solution | Deploying self-hosted voice AI infrastructure on AWS
After evaluating its build-buy-partner options, OpenCall chose to work with AWS Advanced Tier Services Partner Tech 42. Together, they designed a production-ready, containerized text-to-speech solution by using a conversational speech model on AWS. First, the team containerized the model and stored it in Amazon Elastic Container Registry (Amazon ECR), a service for easily storing, sharing, and deploying container software. Then, the team deployed the container on Amazon Elastic Container Service (Amazon ECS), a service for easily building, managing, and running containerized applications.
By using an Application Load Balancer to load balance HTTP and HTTPS traffic, the partner helped OpenCall maintain high availability during peak loads. In parallel, Tech 42 deployed a self-hosted speech-to-text solution, containerizing and orchestrating it through the same Amazon ECS infrastructure. By colocating both the synthesis and transcription services within the same AWS environment, Tech 42 avoided unnecessary network round trips that had previously contributed to latency.
The partner also configured Amazon CloudWatch to observe and optimize workloads. Using Amazon CloudWatch, the solution can collect metrics and trigger alarms to drive scaling policies, automatically scaling in response to concurrent workflows. OpenCall and Tech 42 conducted a multiphase program, including rapid assessment, proof of concept, and production build. Using this structured process, OpenCall could validate performance before committing to full deployment.
Outcome | Achieving 190 ms latency and preventing vendor lock-in
During testing, OpenCall’s text-to-speech and transcription pipeline achieved about 190 ms of latency to first output. This value is 33 percent below the maximum production threshold of 300 ms and within the ideal target range of 150–200 ms. Validated by OpenCall’s cofounder, this result unblocked the path to production deployment. “Using AWS, we can classify intent and chain tool calls, achieving more consistent performance with almost no perceptible lag to the user,” says Silverstein.
The infrastructure handled 200 concurrent calls during testing, with more than 8,000 text-to-speech requests per hour at peak. The team carried out a 5,000-patient outbound campaign in 1 night. Designed to scale well beyond current load, the architecture doesn’t require manual infrastructure intervention. “We expect to increase concurrent calls by two to three times within 4 months, reaching tens of thousands within 1 year,” says Silverstein.
About AWS Partner Partner Tech 42
US company Tech 42 specializes in cloud infrastructure, AI/ML deployments, and containerized-application architecture.
Using AWS, we can classify intent and chain tool calls, achieving more consistent performance with almost no perceptible lag to the user.
Arthur Silverstein
Cofounder, OpenCallAWS Services Used
Did you find what you were looking for today?
Let us know so we can improve the quality of the content on our pages