A data fabric is an architecture for connecting data that lives in many places, such as cloud applications, databases, data warehouses, data lakes and on-premises systems, so people and applications can find it, access it and use it under consistent rules. Its core is shared metadata: a continuously updated record of what data exists, where it lives, what it means and who may use it. Integration, virtualization and governance tools use that metadata to deliver data where it is needed, sometimes without copying it into one central store.
At a glance
- An architecture, not a single product, built from catalogs, integration tools and governance controls.
- Uses metadata about data (location, meaning, lineage, sensitivity) to connect sources across cloud and on-premises systems.
- Can query data in place, copy it, or both, depending on the tools and the workload.
- Often adds automation, such as suggesting data sources, classifying data or flagging quality issues, depending on the vendor.
- Largely a technology answer to scattered data; it does not by itself assign ownership or fix quality.
What problem it solves
Most mid-market organizations run dozens of SaaS applications plus a few databases, a warehouse or lake, and files in shared drives. Each new report, integration or AI project tends to start the same way: someone finds the right sources, works out what fields mean, gets access and builds another pipeline. Over time that produces a tangle of point-to-point connections, duplicate copies and inconsistent definitions.
A data fabric tries to replace that one-off work with a shared layer. Once a source is connected and described, it can be found and reused by other teams, and the same access and privacy rules can be applied across connected sources, to the extent the tools support each one. It also gives security and compliance teams a better view of where sensitive data sits.
How it works
Metadata and catalog. Connectors scan sources and record technical details (tables, fields, formats) and business details (definitions, owners, quality scores). Lineage shows where data came from and what depends on it. Data classification tags sensitive data such as personal or financial records.
Integration. Data is delivered in whatever way suits the use: batch pipelines such as extract, transform, load (ETL), streaming for near-real-time feeds, APIs, or data virtualization, which queries sources in place and presents them as if they were one system.
Governance and policy. Access, masking and retention rules are defined once and applied across the connected sources, to the extent the tools support each source. This is where a fabric overlaps most with data governance programs.
Automation. Many products use machine learning to suggest matches between datasets, recommend sources or detect unusual changes. Treat these as aids that still need human review.
Consumption. Analysts, analytics and business intelligence (ABI) tools, applications and AI models draw on the connected data through the catalog, a query layer or delivered datasets, which may land in a data lake or data lakehouse.
When it matters for buyers
- When data is spread across many systems and clouds. The more sources and teams you have, the more a shared metadata layer can save.
- When compliance requires knowing where data is. Privacy and data residency questions are easier to answer with a catalog and lineage.
- When AI projects stall on data access. A fabric can shorten the path from “which data do we have?” to a usable dataset.
- When integration costs keep rising. Reusable connections can replace some one-off pipelines, though they still need maintenance.
Our analytics and business intelligence overview covers the platforms involved, and our multi-cloud overview covers managing data and workloads across providers.
Questions to ask vendors
- Which of our sources (applications, databases, clouds, on-premises systems) can you connect to today, and which need custom work?
- Do you query data in place, copy it, or both, and how does that affect performance and cost?
- How is metadata collected and kept current, and how much manual curation should we expect?
- How are access and masking policies enforced in each connected source?
- Which automated features are included, and how are their suggestions reviewed?
- How is the platform priced (by connector, data volume, users or compute)?
- Can we export our catalog and lineage if we change tools?
How it differs from a data mesh
A data fabric and a data mesh are often presented as rivals, but they answer different questions. A data fabric is mostly about technology: a shared, metadata-driven layer, often run by a central team, that connects data across systems. A data mesh is mostly about organization: business domains such as sales or operations own their data and publish it as products for others to use, with a self-service platform and common standards. A mesh still needs connecting technology, and a fabric still needs owners, so many organizations use fabric tools to support a mesh-style operating model.
