A data lake is a central repository that stores large volumes of data in its raw, original form, whether structured tables, semi-structured files such as logs and JSON, or unstructured content such as documents, images and audio. Data is usually kept on low-cost, scalable storage, most often cloud object storage, and is organized and interpreted only when someone reads it for a specific purpose. Data lakes are commonly used to feed analytics, data science and machine learning, often alongside a data warehouse rather than in place of one.
At a glance
- Stores data as it arrives, in many formats, without requiring it to fit a fixed structure first.
- Usually built on cloud object storage, which is far cheaper per terabyte than database or warehouse storage.
- Structure is applied when data is read (“schema on read”), which keeps loading simple but pushes work onto whoever queries it.
- Commonly used for logs, event data, IoT readings, files and the raw inputs to analytics and AI.
- Without catalogs, ownership and access rules, a data lake can turn into a “data swamp” nobody trusts.
What problem it solves
Organizations generate data in many systems: applications, websites, sensors, call recordings, security tools, spreadsheets. Much of it does not fit neatly into the rows and columns of a traditional data warehouse, and loading it all into a warehouse would be slow and expensive. As a result, it either gets thrown away or sits in silos where nobody can combine it.
A data lake gives that data one inexpensive place to land. Teams can keep raw data for years, combine sources that were never designed to work together and go back to the original records when a new question comes up. Data scientists and machine learning projects in particular need access to large, detailed, unfiltered datasets, which a data lake is designed to hold.
How it works
Ingestion. Data flows in from source systems in batches or as continuous streams. Many teams load data largely as-is rather than reshaping it on the way in, though pipelines such as extract, transform, load (ETL) or its variants are still used to move and prepare it.
Storage. Files are kept in object storage or a distributed file system, often in open, analytics-friendly file formats. Data is commonly organized into zones, for example raw, cleaned and curated, so users can tell how processed each dataset is.
Catalog and governance. A data catalog records what each dataset contains, where it came from, who owns it and who may use it. Data classification and access policies control who can see sensitive fields. This layer is what separates a usable lake from a swamp.
Processing and query. Separate compute engines read the data: SQL query engines for analysts, distributed processing frameworks for large transformations, and notebooks and ML tools for data scientists. Because storage and compute are separate, each can scale on its own, and some platforms charge for compute only while queries run.
Serving. Curated data is often loaded into a data warehouse or business intelligence tool for reporting, or read directly by “lakehouse” platforms that add table features on top of lake storage.
See our analytics and business intelligence page for how organizations put this data to use.
When it matters for buyers
- Starting an analytics or AI program. A lake is often the first shared foundation for data that doesn’t fit existing systems.
- When warehouse costs climb. Moving raw, rarely queried data out of a warehouse and into a lake can reduce storage spend, though query costs need watching.
- Consolidating after growth or acquisition. A lake can bring data from many systems together without first agreeing on a single data model.
- Compliance and retention. Large stores of raw data raise questions about personal data, retention periods and data governance that must be answered up front.
- Choosing a platform. Cloud providers, data platform vendors and open-source stacks all offer data lake capabilities with different pricing and lock-in trade-offs.
Questions to ask vendors
- How is storage priced, and how are queries, compute and data movement priced separately?
- Which file and table formats do you use, and can other tools read our data if we leave?
- What catalog, lineage and data quality tools are included?
- How are access controls applied at the level of datasets, columns and rows?
- How do you handle personal or regulated data, including deletion requests and retention rules?
- What are the egress or export costs to move our data out?
- Can the same data serve reporting and data science, or will we need a separate warehouse?
How it differs from a data warehouse
These are traditional tendencies, not fixed rules. A data warehouse, whether run in-house or bought as data warehouse as a service (DWaaS), is optimized for governed, modeled data, which makes it fast and consistent for business reporting but has typically made it less suited to raw or unstructured data and more expensive per terabyte. A data lake emphasizes scalable object or file storage and flexible formats, usually applying structure when data is read, which tends to make it cheaper and more flexible but harder for business users to query directly. ELT pipelines, which load raw data into warehouses before transforming it, and lakehouse designs, which add warehouse-style tables to lake storage, blur where schema and transformation happen. Many organizations run both, using the lake as the landing zone and the warehouse for curated reporting.
