A data lakehouse is a way of building a data platform that combines the low-cost, flexible storage of a data lake with the structure and query performance traditionally associated with a data warehouse. In a common pattern, data sits in object storage, often in open file formats, and a table layer on top adds features such as schemas, transactions, versioning and fast SQL queries. File and table formats, which query engines can use the data, and how storage and compute are scaled and billed all vary by platform. The goal is one copy of data that can serve business reporting, data science and AI, instead of loading the same data into two separate systems.
At a glance
- Typically stores data in object storage, often in open file formats, with a table layer that adds warehouse-style features.
- Aims to serve reporting, data science and machine learning from one copy of data.
- Often built on open table formats, so more than one query engine can read the same tables, depending on the platform.
- Sold as a managed cloud platform by several vendors, or assembled from open-source components.
- Still needs governance, cataloging and cost controls; the architecture alone does not make data trustworthy.
What problem it solves
Many organizations ended up with two data platforms. A data lake held large volumes of raw, semi-structured and unstructured data cheaply, which suited data science. A data warehouse held cleaned, modeled data for dashboards and finance reporting. Moving data between them meant extra pipelines, duplicate storage, delays and, often, two versions of the truth.
Data lakes on their own had well-known weaknesses for reporting: files could be partially written or overwritten, schemas drifted, and queries over millions of files were slow. Warehouses, on the other hand, could become expensive as raw data volumes grew, and some were awkward for machine learning tools that want direct access to files. A lakehouse tries to close that gap by giving the lake the reliability features analysts need.
How it works
Storage. Data is commonly kept in cloud or on-premises object storage (see file systems and object storage), usually in columnar file formats designed for analytics. Many platforms separate storage from compute, which can let each be scaled, and sometimes billed, on its own; how far that goes, and how it is priced, varies by platform.
Table layer. An open table format or a vendor’s own table layer adds metadata that turns folders of files into tables. Common capabilities include atomic transactions (a write fully succeeds or fails), schema enforcement and evolution, time travel to earlier versions, and efficient updates and deletes, which matter for corrections and privacy requests.
Query engines. SQL engines, data science notebooks and machine learning (ML) frameworks can read the same tables, to the extent the platform and table format support each engine. Many platforms add caching, indexing and file compaction to bring query speed closer to a dedicated warehouse, with results that vary by workload.
Data flow. Data typically arrives through ELT or extract, transform, load (ETL) pipelines and is refined in stages, often described as raw, cleaned and business-ready layers. Analytics and business intelligence (ABI) tools connect to the curated layer.
Governance. A catalog records what tables exist, who owns them and who can access them. Access controls, lineage and auditing are provided by the platform or by add-on tools, and their depth differs between products.
When it matters for buyers
- When you run, or are about to build, both a lake and a warehouse. A lakehouse may let you consolidate, but only if it meets your reporting performance and concurrency needs.
- When AI and data science need the same data as finance. One governed copy reduces disagreements between teams.
- When warehouse bills grow with data volume. Moving raw and historical data to object storage can lower storage cost, though compute spend still needs monitoring.
- When lock-in is a concern. Open table formats can make it easier to switch query engines, so ask how open a given platform really is.
Our analytics and business intelligence overview covers how organizations choose and run data platforms.
Questions to ask vendors
- Which table formats do you support, and can other engines read and write our tables without your platform?
- How is pricing structured for storage, compute and data transfer, and what would our current workloads cost?
- What query performance and concurrency can we expect for our largest dashboards, and can we test with our own data?
- How do you handle updates, deletes and privacy requests at scale?
- What catalog, access control, lineage and audit features are included, and which cost extra?
- Where is our data stored, and can we keep it in our own cloud account?
- What does it take to export our data and metadata if we leave?
How it differs from a data lake and a data warehouse
A data lake stores raw files cheaply and applies structure when data is read, which makes it flexible but harder to query reliably. A data warehouse, whether run in-house or bought as data warehouse as a service (DWaaS), stores modeled data in its own managed format and is tuned for fast, consistent reporting. A lakehouse keeps data in lake storage but adds a table layer so it behaves more like a warehouse. In practice, the lines blur: many warehouses now read lake files, and many lakehouse platforms add warehouse-style performance features. Compare specific capabilities and costs rather than labels.
A lakehouse is also different from a data fabric, which is an integration and metadata layer that connects data across many systems; a lakehouse is one place where data is stored and queried, and it can be one of the sources a fabric connects. Whatever the design, data governance decides whether the data is trusted.
