Reimagining legal analysis and research using AI on AWS with Forlex
Learn how legal-technology company Forlex cut costs and boosted performance by using AWS and NVIDIA Blackwell GPUs.
Benefits
improvement in time to first token
improvement in tokens processed per minute
Overview
Forlex needed to scale its proprietary AI models as its user base and data volumes grew. The company’s existing infrastructure limited its model-training speed, inference performance, and ability to expand into new markets. So Forlex used Amazon Web Services (AWS), which provides a comprehensive portfolio of AI and machine learning services, and the advanced capabilities of NVIDIA Blackwell GPUs. That way, the company reduced training time, increased throughput, and cut operational costs, positioning itself to deliver scalable, high-accuracy legal AI worldwide.
About Forlex
Founded in 2023, Forlex is a legal-technology company that aims to democratize access to justice in Brazil through AI. Its specialized small language models help automate and streamline legal workflows with high accuracy and efficiency.
Opportunity | Using AWS to scale legal AI infrastructure for Forlex
Forlex’s solutions help legal professionals conduct research, analyze documents, and automate routine workflows. As adoption and data volumes increased, the company’s infrastructure could no longer provide the scalability, cost efficiency, or training speed necessary to support growth. “Our proprietary models require continual training on high-quality datasets,” says Daniel Bichuetti, co-CEO and chief technology officer at Forlex. “But we couldn’t maintain service quality with the existing infrastructure.”
That infrastructure also restricted future expansion, especially as new markets and evolving legal frameworks added to training and compute needs. “Our limited compute resources made us slow to update models and respond to legal changes,” says Bichuetti.
In addition, expanding into new markets would require the company to meet new data sovereignty obligations and help customers scale properly in their respective jurisdictions. To continue growing while maintaining a lean team and facilitating compliance, Forlex migrated its AI workloads to AWS.
Solution | Migrating to AWS and NVIDIA Blackwell GPUs
To achieve high GPU performance for AI training and inference, the company used Amazon Elastic Compute Cloud (Amazon EC2) P6-B200 Instances, which are accelerated by NVIDIA Blackwell GPUs. Besides its high performance and compatibility with popular AI frameworks, the Blackwell architecture supports 4-bit floating-point (FP4) precision. This helps Forlex process calculations efficiently and speed up innovation with its small language models while maintaining model accuracy.
To improve its ability to satisfy security and compliance requirements, Forlex encrypts data by using AWS Key Management Service (AWS KMS) to create and control keys. The company also implemented Amazon Elastic Kubernetes Service (Amazon EKS), a service for building, running, and scaling production-ready Kubernetes applications. This way, the team can flexibly optimize server configuration for performance and dynamic scaling. “Using Amazon EC2 P6-B200 Instances and tools such as Amazon EKS, we dramatically reduced our operational overhead,” says Bichuetti. “For small teams in new businesses, this is like a dream.”
In addition, Forlex joined the NVIDIA Inception program, which offers guidance on NVIDIA products and access to educational resources. The company adopted several NVIDIA tools to orchestrate multimodal large language models (LLMs), accelerate model inference, and train models at scale. For example, Forlex powers natural language processing with NVIDIA CUDA and optimizes text generation and summarization with TensorRT LLM. “We’re using AWS and NVIDIA technology to stand out from competitors,” says Bichuetti.
Outcome | Increasing efficiency and performance for users worldwide
Using the compute power and tooling of AWS and NVIDIA, Forlex’s 13 specialized models can automate document analysis, legal research, and general workflows in 30–60 seconds. “We’ve increased the tokens that we process per minute by 170 percent and reduced the time to first token by about 67 percent,” says Bichuetti.
The company cut operational costs by about 40 percent and expects a full return on investment within 3–6 months. Additionally, by off-loading infrastructure management, the team can be more responsive to its user base and focus its efforts on product improvement.
With the global reach of AWS and the performance of NVIDIA Blackwell GPUs, Forlex has turned AI scalability into a competitive advantage. “We closed a deal in Brazil because we could scale quickly to meet demand, and we’ve recently started deals with the United Kingdom,” says Bichuetti. “Using GPU-accelerated infrastructure, we’re prepared to serve a growing user base worldwide.”
Using Amazon EC2 P6-B200 Instances and tools such as Amazon EKS, we dramatically reduced our operational overhead. For small teams in new businesses, this is like a dream.
Daniel Bichuetti
Co-CEO and Chief Technology Officer, ForlexAWS Services Used
More Customer Stories
Did you find what you were looking for today?
Let us know so we can improve the quality of the content on our pages