What Is a Data Lake?

Related problems: Data from many systems scattered in places nobody can query together; Need to keep large amounts of raw data cheaply for analytics or AI; Our data warehouse is too expensive for logs, files and event data; Analysts can't find or trust the data that has been collected

A data lake is a central repository that stores large volumes of data in its raw, original form, whether structured tables, semi-structured files such as logs and JSON, or unstructured content such as documents, images and audio. Data is usually kept on low-cost, scalable storage, most often cloud object storage, and is organized and interpreted only when someone reads it for a specific purpose. Data lakes are commonly used to feed analytics, data science and machine learning, often alongside a data warehouse rather than in place of one.

At a glance

  • Stores data as it arrives, in many formats, without requiring it to fit a fixed structure first.
  • Usually built on cloud object storage, which is far cheaper per terabyte than database or warehouse storage.
  • Structure is applied when data is read (“schema on read”), which keeps loading simple but pushes work onto whoever queries it.
  • Commonly used for logs, event data, IoT readings, files and the raw inputs to analytics and AI.
  • Without catalogs, ownership and access rules, a data lake can turn into a “data swamp” nobody trusts.

What problem it solves

Organizations generate data in many systems: applications, websites, sensors, call recordings, security tools, spreadsheets. Much of it does not fit neatly into the rows and columns of a traditional data warehouse, and loading it all into a warehouse would be slow and expensive. As a result, it either gets thrown away or sits in silos where nobody can combine it.

A data lake gives that data one inexpensive place to land. Teams can keep raw data for years, combine sources that were never designed to work together and go back to the original records when a new question comes up. Data scientists and machine learning projects in particular need access to large, detailed, unfiltered datasets, which a data lake is designed to hold.

How it works

Ingestion. Data flows in from source systems in batches or as continuous streams. Many teams load data largely as-is rather than reshaping it on the way in, though pipelines such as extract, transform, load (ETL) or its variants are still used to move and prepare it.

Storage. Files are kept in object storage or a distributed file system, often in open, analytics-friendly file formats. Data is commonly organized into zones, for example raw, cleaned and curated, so users can tell how processed each dataset is.

Catalog and governance. A data catalog records what each dataset contains, where it came from, who owns it and who may use it. Data classification and access policies control who can see sensitive fields. This layer is what separates a usable lake from a swamp.

Processing and query. Separate compute engines read the data: SQL query engines for analysts, distributed processing frameworks for large transformations, and notebooks and ML tools for data scientists. Because storage and compute are separate, each can scale on its own, and some platforms charge for compute only while queries run.

Serving. Curated data is often loaded into a data warehouse or business intelligence tool for reporting, or read directly by “lakehouse” platforms that add table features on top of lake storage.

See our analytics and business intelligence page for how organizations put this data to use.

When it matters for buyers

  • Starting an analytics or AI program. A lake is often the first shared foundation for data that doesn’t fit existing systems.
  • When warehouse costs climb. Moving raw, rarely queried data out of a warehouse and into a lake can reduce storage spend, though query costs need watching.
  • Consolidating after growth or acquisition. A lake can bring data from many systems together without first agreeing on a single data model.
  • Compliance and retention. Large stores of raw data raise questions about personal data, retention periods and data governance that must be answered up front.
  • Choosing a platform. Cloud providers, data platform vendors and open-source stacks all offer data lake capabilities with different pricing and lock-in trade-offs.

Questions to ask vendors

  • How is storage priced, and how are queries, compute and data movement priced separately?
  • Which file and table formats do you use, and can other tools read our data if we leave?
  • What catalog, lineage and data quality tools are included?
  • How are access controls applied at the level of datasets, columns and rows?
  • How do you handle personal or regulated data, including deletion requests and retention rules?
  • What are the egress or export costs to move our data out?
  • Can the same data serve reporting and data science, or will we need a separate warehouse?

How it differs from a data warehouse

These are traditional tendencies, not fixed rules. A data warehouse, whether run in-house or bought as data warehouse as a service (DWaaS), is optimized for governed, modeled data, which makes it fast and consistent for business reporting but has typically made it less suited to raw or unstructured data and more expensive per terabyte. A data lake emphasizes scalable object or file storage and flexible formats, usually applying structure when data is read, which tends to make it cheaper and more flexible but harder for business users to query directly. ELT pipelines, which load raw data into warehouses before transforming it, and lakehouse designs, which add warehouse-style tables to lake storage, blur where schema and transformation happen. Many organizations run both, using the lake as the landing zone and the warehouse for curated reporting.

Frequently Asked Questions

What is the difference between a data lake and a data warehouse?
Traditionally, a data warehouse is optimized for governed, modeled data organized for reporting, while a data lake emphasizes scalable object or file storage and flexible formats, with structure applied when data is read. Warehouses have tended to suit consistent business reporting; lakes have tended to be cheaper per terabyte and more flexible for data science and large volumes. ELT and lakehouse designs blur where schema and transformation happen, so compare specific platforms rather than labels.
What is a data swamp?
An informal name for a data lake that has filled with data nobody can find, understand or trust, because it was loaded without catalogs, ownership or quality rules. Governance, not technology, is usually what prevents it.
What is a data lakehouse?
A design that adds warehouse-style features, such as tables, transactions and faster queries, on top of data lake storage, so one copy of the data can serve both reporting and data science. Many current platforms are sold this way, but capabilities vary, so test with your own workloads.
Do we need a data lake for AI?
Not necessarily, but many AI and machine learning projects need large amounts of raw or semi-structured data, and a data lake is a common place to collect it. The more important questions are whether the data is labeled, governed and permitted to be used for that purpose.
Is a data lake secure by default?
Not necessarily. A data lake concentrates a lot of data in one place, often including sensitive data. Access controls, encryption, classification and monitoring must be designed in, and the defaults differ by platform.

You Don’t Need Another Sales Call. You Need an Answer.

30 minutes. No pitch. Just an honest conversation about where you are, what you need, and whether working together makes sense.

We use your details to set up and prepare for the call, and send the newsletter only if you ask for it. Privacy policy.