What Is a Data Lakehouse?

Also called: Lakehouse, Lakehouse architecture

Related problems: Paying to store and load the same data in both a data lake and a data warehouse; Reports and data science teams working from different copies that disagree; Data lake files that analysts can't query reliably or quickly; Warehouse costs rising as raw and semi-structured data volumes grow

A data lakehouse is a way of building a data platform that combines the low-cost, flexible storage of a data lake with the structure and query performance traditionally associated with a data warehouse. In a common pattern, data sits in object storage, often in open file formats, and a table layer on top adds features such as schemas, transactions, versioning and fast SQL queries. File and table formats, which query engines can use the data, and how storage and compute are scaled and billed all vary by platform. The goal is one copy of data that can serve business reporting, data science and AI, instead of loading the same data into two separate systems.

At a glance

  • Typically stores data in object storage, often in open file formats, with a table layer that adds warehouse-style features.
  • Aims to serve reporting, data science and machine learning from one copy of data.
  • Often built on open table formats, so more than one query engine can read the same tables, depending on the platform.
  • Sold as a managed cloud platform by several vendors, or assembled from open-source components.
  • Still needs governance, cataloging and cost controls; the architecture alone does not make data trustworthy.

What problem it solves

Many organizations ended up with two data platforms. A data lake held large volumes of raw, semi-structured and unstructured data cheaply, which suited data science. A data warehouse held cleaned, modeled data for dashboards and finance reporting. Moving data between them meant extra pipelines, duplicate storage, delays and, often, two versions of the truth.

Data lakes on their own had well-known weaknesses for reporting: files could be partially written or overwritten, schemas drifted, and queries over millions of files were slow. Warehouses, on the other hand, could become expensive as raw data volumes grew, and some were awkward for machine learning tools that want direct access to files. A lakehouse tries to close that gap by giving the lake the reliability features analysts need.

How it works

Storage. Data is commonly kept in cloud or on-premises object storage (see file systems and object storage), usually in columnar file formats designed for analytics. Many platforms separate storage from compute, which can let each be scaled, and sometimes billed, on its own; how far that goes, and how it is priced, varies by platform.

Table layer. An open table format or a vendor’s own table layer adds metadata that turns folders of files into tables. Common capabilities include atomic transactions (a write fully succeeds or fails), schema enforcement and evolution, time travel to earlier versions, and efficient updates and deletes, which matter for corrections and privacy requests.

Query engines. SQL engines, data science notebooks and machine learning (ML) frameworks can read the same tables, to the extent the platform and table format support each engine. Many platforms add caching, indexing and file compaction to bring query speed closer to a dedicated warehouse, with results that vary by workload.

Data flow. Data typically arrives through ELT or extract, transform, load (ETL) pipelines and is refined in stages, often described as raw, cleaned and business-ready layers. Analytics and business intelligence (ABI) tools connect to the curated layer.

Governance. A catalog records what tables exist, who owns them and who can access them. Access controls, lineage and auditing are provided by the platform or by add-on tools, and their depth differs between products.

When it matters for buyers

  • When you run, or are about to build, both a lake and a warehouse. A lakehouse may let you consolidate, but only if it meets your reporting performance and concurrency needs.
  • When AI and data science need the same data as finance. One governed copy reduces disagreements between teams.
  • When warehouse bills grow with data volume. Moving raw and historical data to object storage can lower storage cost, though compute spend still needs monitoring.
  • When lock-in is a concern. Open table formats can make it easier to switch query engines, so ask how open a given platform really is.

Our analytics and business intelligence overview covers how organizations choose and run data platforms.

Questions to ask vendors

  • Which table formats do you support, and can other engines read and write our tables without your platform?
  • How is pricing structured for storage, compute and data transfer, and what would our current workloads cost?
  • What query performance and concurrency can we expect for our largest dashboards, and can we test with our own data?
  • How do you handle updates, deletes and privacy requests at scale?
  • What catalog, access control, lineage and audit features are included, and which cost extra?
  • Where is our data stored, and can we keep it in our own cloud account?
  • What does it take to export our data and metadata if we leave?

How it differs from a data lake and a data warehouse

A data lake stores raw files cheaply and applies structure when data is read, which makes it flexible but harder to query reliably. A data warehouse, whether run in-house or bought as data warehouse as a service (DWaaS), stores modeled data in its own managed format and is tuned for fast, consistent reporting. A lakehouse keeps data in lake storage but adds a table layer so it behaves more like a warehouse. In practice, the lines blur: many warehouses now read lake files, and many lakehouse platforms add warehouse-style performance features. Compare specific capabilities and costs rather than labels.

A lakehouse is also different from a data fabric, which is an integration and metadata layer that connects data across many systems; a lakehouse is one place where data is stored and queried, and it can be one of the sources a fabric connects. Whatever the design, data governance decides whether the data is trusted.

Frequently Asked Questions

Is a data lakehouse a product or an architecture?
An architecture. Several cloud data platforms are sold as lakehouses, and you can also assemble one, commonly from object storage, a table format and one or more query engines. Because the label is used loosely, compare what each platform actually supports rather than whether it uses the word.
Does a lakehouse replace our data warehouse?
Sometimes. Some organizations consolidate onto a lakehouse; others keep a warehouse for core financial and operational reporting and use the lakehouse for raw data, data science and AI. Test your heaviest reports and concurrency needs before retiring a warehouse.
What is an open table format?
A specification for organizing data files in object storage as tables, with a metadata layer that tracks versions, schema and changes. Open formats let more than one query engine read the same tables, which can reduce lock-in, although vendors differ in how fully they support each format.
Is a lakehouse cheaper than a warehouse?
It can be, mainly because storage in object stores is usually inexpensive and one copy of data can serve more uses. Total cost still depends on compute usage, query patterns, data egress and the staff time needed to run it, so model your own workloads.
Does a lakehouse fix data quality problems?
No. Table features such as schema enforcement help, but quality, ownership and definitions still come from data governance. A lakehouse without catalogs and owners can become as hard to trust as a neglected data lake.

You Don’t Need Another Sales Call. You Need an Answer.

30 minutes. No pitch. Just an honest conversation about where you are, what you need, and whether working together makes sense.

We use your details to set up and prepare for the call, and send the newsletter only if you ask for it. Privacy policy.