Skip to main content

Reimagining legal analysis and research using AI on AWS with Forlex

Learn how legal-technology company Forlex cut costs and boosted performance by using AWS and NVIDIA Blackwell GPUs.

Benefits

reduction in operational costs

improvement in time to first token

improvement in tokens processed per minute

Overview

Forlex needed to scale its proprietary AI models as its user base and data volumes grew. The company’s existing infrastructure limited its model-training speed, inference performance, and ability to expand into new markets. So Forlex used Amazon Web Services (AWS), which provides a comprehensive portfolio of AI and machine learning services, and the advanced capabilities of NVIDIA Blackwell GPUs. That way, the company reduced training time, increased throughput, and cut operational costs, positioning itself to deliver scalable, high-accuracy legal AI worldwide.

About Forlex

Founded in 2023, Forlex is a legal-technology company that aims to democratize access to justice in Brazil through AI. Its specialized small language models help automate and streamline legal workflows with high accuracy and efficiency.

Opportunity | Using AWS to scale legal AI infrastructure for Forlex

Forlex’s solutions help legal professionals conduct research, analyze documents, and automate routine workflows. As adoption and data volumes increased, the company’s infrastructure could no longer provide the scalability, cost efficiency, or training speed necessary to support growth. “Our proprietary models require continual training on high-quality datasets,” says Daniel Bichuetti, co-CEO and chief technology officer at Forlex. “But we couldn’t maintain service quality with the existing infrastructure.”

That infrastructure also restricted future expansion, especially as new markets and evolving legal frameworks added to training and compute needs. “Our limited compute resources made us slow to update models and respond to legal changes,” says Bichuetti.

In addition, expanding into new markets would require the company to meet new data sovereignty obligations and help customers scale properly in their respective jurisdictions. To continue growing while maintaining a lean team and facilitating compliance, Forlex migrated its AI workloads to AWS.

Solution | Migrating to AWS and NVIDIA Blackwell GPUs

To achieve high GPU performance for AI training and inference, the company used Amazon Elastic Compute Cloud (Amazon EC2) P6-B200 Instances, which are accelerated by NVIDIA Blackwell GPUs. Besides its high performance and compatibility with popular AI frameworks, the Blackwell architecture supports 4-bit floating-point (FP4) precision. This helps Forlex process calculations efficiently and speed up innovation with its small language models while maintaining model accuracy.

To improve its ability to satisfy security and compliance requirements, Forlex encrypts data by using AWS Key Management Service (AWS KMS) to create and control keys. The company also implemented Amazon Elastic Kubernetes Service (Amazon EKS), a service for building, running, and scaling production-ready Kubernetes applications. This way, the team can flexibly optimize server configuration for performance and dynamic scaling. “Using Amazon EC2 P6-B200 Instances and tools such as Amazon EKS, we dramatically reduced our operational overhead,” says Bichuetti. “For small teams in new businesses, this is like a dream.”

In addition, Forlex joined the NVIDIA Inception program, which offers guidance on NVIDIA products and access to educational resources. The company adopted several NVIDIA tools to orchestrate multimodal large language models (LLMs), accelerate model inference, and train models at scale. For example, Forlex powers natural language processing with NVIDIA CUDA and optimizes text generation and summarization with TensorRT LLM. “We’re using AWS and NVIDIA technology to stand out from competitors,” says Bichuetti.

Outcome | Increasing efficiency and performance for users worldwide

Using the compute power and tooling of AWS and NVIDIA, Forlex’s 13 specialized models can automate document analysis, legal research, and general workflows in 30–60 seconds. “We’ve increased the tokens that we process per minute by 170 percent and reduced the time to first token by about 67 percent,” says Bichuetti.

The company cut operational costs by about 40 percent and expects a full return on investment within 3–6 months. Additionally, by off-loading infrastructure management, the team can be more responsive to its user base and focus its efforts on product improvement.

With the global reach of AWS and the performance of NVIDIA Blackwell GPUs, Forlex has turned AI scalability into a competitive advantage. “We closed a deal in Brazil because we could scale quickly to meet demand, and we’ve recently started deals with the United Kingdom,” says Bichuetti. “Using GPU-accelerated infrastructure, we’re prepared to serve a growing user base worldwide.”

Missing alt text value
Using Amazon EC2 P6-B200 Instances and tools such as Amazon EKS, we dramatically reduced our operational overhead. For small teams in new businesses, this is like a dream.

Daniel Bichuetti

Co-CEO and Chief Technology Officer, Forlex

Did you find what you were looking for today?

Let us know so we can improve the quality of the content on our pages