A data lake is a store for raw data of every kind, kept in the form it arrived in. AWS defines it as "a centralized repository that allows you to store all your structured and unstructured data at any scale". Wikipedia puts it more simply: "a system or repository of data stored in its natural, raw format". Unlike a data warehouse, a lake doesn't require a structure up front. In AWS's words, "The structure of the data or schema is not defined when data is captured."
The risk
Storing everything is easy. Finding and trusting it later is not. AWS names the main challenge plainly: "raw data is stored with no oversight of the contents". Without ways to catalogue and secure it, AWS warns, data "cannot be found, or trusted resulting in a 'data swamp'". Wikipedia's entry notes that "Poorly managed data lakes have been facetiously called data swamps."
In a timber business
Much timber data arrives as documents and machine files rather than system records: contracts, tickets, settlements, harvester files and scanner output. Keeping them as received is worth doing. The value comes later. The facts inside are read into records that can be joined with everything else, each linked back to its page.
Whether that needs a lake is another question. AWS, which sells data lake services, cites a 451 Research survey, undated on its page, in which "more than half of surveyed enterprises have a data lake implemented today, with another 22% citing plans". Those are large enterprises with very large volumes of mixed data. A mid-sized timber business holds far less. In our view, a well-organised document store and an extraction step can usually do the same job.
What it isn't
A lake isn't a warehouse, which holds organised, defined data for analysis. Many businesses use both. The lake holds raw files, and the warehouse holds the records built from them.
What to check
Ask how files are catalogued, so you can find every contract for a property. Ask how the facts inside them become records, and how each record links back to its page. And ask who can see what, since contracts and settlements are sensitive.
How Quarri works explains the platform as a layer over existing systems, not a migration.
Sources
- AWS, "What is a data lake?": aws.amazon.com
- Wikipedia, "Data lake": en.wikipedia.org
Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.