Skip to main content

What is data fabric?

Data fabric is an architectural design pattern that uses metadata and automation to discover, integrate, process, and share data across systems, departments, and hybrid, multi-cloud infrastructure. The data fabric layers contain a series of data services that operate across existing and legacy databases, data warehouses, object stores, data lakes, and other data storage systems. The data fabric uses active metadata with machine learning (ML) techniques to automate data discovery, maintain data governance policies, and provide recommendations for data.

Why is data fabric important?

Data fabric in data governance

Data fabric helps organizations to surface data that resides in disparate or siloed data storage. As data systems grow, data engineers face challenges in consolidating, governing, and accessing data across disparate data sources. This results in slower time-to-insights, compliance challenges, and data consistency. By implementing a data fabric, engineers can improve observability, governance, and data transformation without massive data movements. Instead of manually adding more data pipelines, data teams implement an integrated environment that enables end-to-end data lifecycle management, a unified view, and self-service access.

What are the key components of data fabric?

Data fabric enhances data management tasks by combining different data integration styles with machine learning, semantics graph, and visualization technologies. While some data fabric implementations are simple, they can be expanded to cover these layers.

ML-augmented data catalog

A data fabric solution uses machine learning to connect with data lakes, data warehouses, and other databases. Here, data engineers create an inventory that consolidates disparate data assets and provides context for business queries. This data catalog includes tags, historical tracking, and access management.

Active metadata layer

Active metadata describes the continuously updated access patterns, transformation, usage, and other information associated with specific datasets. With active metadata, data teams track data lineage, identify patterns, improve data governance, and perform accurate analysis. Data fabric uses an active metadata layer to turn this information into useful recommendations. For example, it identifies schema drift, maps possible inaccuracies, and suggests actions for data teams.

Knowledge graph

A knowledge graph is a semantic model that represents the relationships among datasets and the underlying business entities. You can think of knowledge graphs as a web that connects every related object and concept together. By using machine learning, data fabric systems can further enrich a knowledge graph, enabling data analysts to discover hidden patterns across disparate systems.

Data integration layer

Data fabric provides multiple data integration styles that data engineers can use to automatically access, manage, and analyze disparate data sources. This includes ETL, ELT, data virtualization, data streaming, and data replication. Based on the active metadata, the data fabric system recommends an integration approach to achieve the desired business outcomes.

Federated governance and policy enforcement

Data fabric enables organizations to enforce policies centrally while supporting granular implementation. Instead of applying domain-specific governance, policies can propagate downstream to specific databases. For example, a data fabric solution can scan multiple databases for sensitive data and mask them.

Data quality management

Data fabric uses machine learning to continuously monitor data quality signals in the active metadata layer. Using these data signals, data teams detect and resolve data quality issues before they impact downstream analytics.

How does data fabric work?

Data fabric leverages automation, AI/ML, data processing, and analytics tools to implement a flexible data architecture. Unless necessary, data is seldom moved between sources in most applications post-ingestion, but instead combined and transformed in virtualized overlays. Here is how data fabric makes data more accessible to non-technical users, business analysts, and product managers.

Metadata ingestion

The data fabric extracts metadata, such as schema, lineage, access logs, and quality signals from ingested data. Depending on use cases, the architecture allows connection to structured and unstructured data originating from enterprise software, mobile apps, Internet of Things (IoT), and social media. You can connect data from various sources, whether hosted in the cloud, on-premises, or both, for data discovery and ingestion.

Metadata analysis

Once ingested, the metadata are categorized accordingly in the data catalog. At this stage, the metadata is considered passive, which doesn’t reflect real-time changes that data engineers use to augment analytics, governance, or quality. To enrich passive metadata to transform it to active metadata, the data fabric system applies machine learning and graph analytics. Active metadata is useful for supporting real-time business intelligence, compliance, and quality checks.

Relationship mapping

Data fabric uses a knowledge graph to map relationships between entities represented in datasets, which helps provide business context. For example, it models relationships among customer feedback, sales transactions, and inventory reports, which are stored in separate databases.

Recommendations

By creating a contextually rich metadata layer, the data fabric can identify hidden patterns and surface recommendations for data users. Data teams can be informed about the best integration paths between datasets, potential data quality issues, and the policies they must address. For example, data fabric systems can automatically create a data pipeline with the appropriate schema, structure, and rate limits.

Governance

Data fabric allows consistent enforcement of policies across the architecture layers of data sources. For example, you can implement dynamic data masking to automatically discover credit card numbers in financial databases and mask them.

Continuous adaptation

Rather than a one-off implementation, data fabric is a continuous process. As data sources, schemas, and policies change, it uses ML-powered analytics to continuously update the underlying metadata layer.

Data fabric architecture

What are the benefits of data fabric?

A data fabric solution creates a metadata-rich layer that provides a foundational context for data discoverability, governance, and management. It helps organizations overcome data management challenges in several ways.

Reduces manual data integration

Data fabric solves challenges that organizations face with massive data movement. It accelerates analytics workflows that transform operational data into useful insights. Additionally, data fabric solutions automate data integration to replace inefficient, error-prone manual data pipelines.

Finds data for analytics and AI/ML

Analytics and AI/ML workloads require ingesting, cleaning, and processing distributed data assets. With data fabric, data scientists can train or feed models with high-quality data using layers of automated data services.

Enhances conformity

Data fabric provides unified data across complex environments without reconstructing existing infrastructure. By applying data fabric solutions, you can establish lineage tracking, policy enforcement, and data consistency.

Provides ready-to-use metadata for all use cases

Data fabric breaks down data silos by replacing traditional data integration methods with centralized virtualization. This allows business teams to enhance analytical and operational performance across different use cases.

Continuous improvement

Another benefit data fabric offers is helping to make sure that data is always fresh, accurate, and relevant. With a feedback loop that continuously updates its active metadata, it makes decisions that are based on the underlying datasets.

What are the use cases of data fabric?

Organizations can apply data fabric architecture in these areas.

Enterprise AI/ML training data

AI/ML development requires training and fine-tuning the underlying AI model with clean, high-quality, and relevant datasets. This is often a challenge as data is distributed across multiple data stores. ML teams use data fabric solutions to discover, govern, and maintain data quality across pipelines, ensuring accurate model inference.

Regulatory compliance

Some industries are compelled by law to track data lineage, data privacy, and generate accurate data reports. By using data fabric solutions, compliance and audit teams can access, monitor, and manage data flow throughout their entire lifecycle. Additionally, you can enforce compliance policies that target specific data stores.

Customer analytics

Business teams derive customer insights from diverse sources, such as CRM, transaction logs, social media, and reviews. Data fabric connects information originating from distributed data sources without moving large volumes of data. By using an ML-powered dashboard, you can identify patterns, surface customer sentiment, and predict shopping behavior across data stores.

What is the difference between data fabric and data mesh?

Data mesh is a decentralized data architecture that allows individual teams to take ownership of managing, governing, and securing domain-specific data. The goal of a data mesh is data-as-a-product that can be accessed by self-service.

Meanwhile, data fabric is an architectural design pattern that allows data engineers to access, share, analyze, and govern data across disparate sources. Both data fabric and data mesh complement each other by enabling individual teams to assume delegated data ownership while maintaining centralized oversight.

What are the challenges of data fabric?

Data fabric is still an emerging data architecture. Despite its benefits, data teams might face these challenges in real-world implementation.

Metadata quality and completeness

Incomplete, outdated, or inaccurate metadata directly impacts the data fabric’s performance. If a data source fails to propagate schema changes, the metadata layer will operate with an outdated context. Some data ecosystems, especially those not built for modern data workflows, might not provide metadata that the data fabric requires.

Tooling complexity

Data fabric requires assembling, integrating, and operationalizing multiple tools, such as data catalog, virtualization, semantic knowledge graphs, governance, and observability. Often, such efforts require setting up and training multi-disciplinary teams across business functions.

Organizational data and ML maturity

Data fabric builds upon basic data cataloging, governance, and data quality practices. If these fundamentals are missing, the data fabric ecosystem cannot deliver the benefits described. Moreover, many organizations are still at the early stage of adopting knowledge graphs and ML-driven automation. Such technical debts can prevent effective data fabric implementation.

How can AWS support your data fabric requirements?

AWS offers a set of services to create your data fabric architecture, with metadata management, data integration, data governance, and ML availability across distributed hybrid and multi-cloud environments. Choose to build with:

  • Amazon DataZone is a data management service that makes it faster and easier for customers to catalog, discover, share, and govern data stored across AWS, on premises, and third-party sources. Amazon DataZone makes it easier for engineers, data scientists, product managers, analysts, and business users to access data throughout an organization so that they can discover, use, and collaborate to derive data-driven insights.
  • AWS Glue provides all the capabilities needed for data integration, so you can gain insights and put your data to work quickly. AWS Glue provides a fully managed, serverless toolkit to design and automate modern data pipelines—with built-in ETL, schema discovery, and cross-service integration.
  • AWS Lake Formation helps you centrally govern, secure, and globally share data for analytics and machine learning. With Lake Formation, you can centralize data security and governance using the AWS Glue Data Catalog, letting you manage metadata and data permissions in one place with familiar database-style features. It also delivers fine-grained data access control, so you can help users have access to the right data down to the row and column level.

Get started with data fabric on AWS by creating a free account today.

Browse all cloud computing concepts

Browse all cloud computing concepts content here:

Loading
Loading
Loading
Loading
Loading

Did you find what you were looking for today?

Let us know so we can improve the quality of the content on our pages