Data lineage security: why encryption is a property of the flow, not the bucket
Here is a pattern that turns up in almost every data platform review. The source database is encrypted with a customer-managed key. The data lake it feeds is encrypted. The warehouse is encrypted. Every individual resource passes its encryption check. And yet, somewhere in the middle, a nightly job writes an intermediate extract to a staging location that has default settings, broad read access and no lifecycle rule. The data is sitting there in a form nobody intended.
No per-resource check will flag this as a data protection problem in context, because each resource is judged alone. The weakness is in the path. That is what data lineage security is about.
What data lineage security means
Data lineage is the record of where data comes from, what transforms it, and where it ends up. Data teams have used lineage for years to debug pipelines and answer "where did this number come from?". Data lineage security applies the same idea to protection: for every route sensitive data takes, are the controls consistent from the first hop to the last?
The questions change from "is this bucket encrypted?" to:
- Where does data from this sensitive source end up?
- Is every hop along that route encrypted, at rest and in transit?
- Does access widen as data moves downstream?
- Does any hop cross a region or account boundary that your commitments do not allow?
Why pipelines create exposure
Pipelines are built for correctness and speed, not for consistent controls. The common causes:
- Intermediate stores. Staging buckets, temporary tables and export folders are created for one job and configured with whatever defaults applied.
- Different owners per hop. The application team owns the source, the data engineering team owns the transforms, the analytics team owns the destination. Each secures its own piece.
- Access widens downstream. Production databases are usually tightly controlled. Analytics destinations are designed to be read by many people. Sensitive columns that were never meant to leave the source travel with the rest.
- Copies outlive their purpose. A one-off migration or backfill leaves a full copy behind, and nothing deletes it.
How teams handle it today
Most teams handle this partly, through a mix of:
- Data catalogue tools maintained by the data team, which describe lineage for analytics but rarely carry security context.
- Per-resource posture checks, which confirm each store is encrypted and private but do not connect them.
- Architecture reviews when a pipeline is first built, which are accurate on the day and drift afterwards.
The gap is the join. The data team knows the flow. The security team knows the controls. Nobody has both on one page, so the weak hop sits between two teams' responsibilities.
A checklist for securing data flows
Map the flows that carry sensitive data
- Start from the stores you have classified as sensitive, not from every pipeline.
- For each, list the downstream destinations: replicas, exports, warehouse tables, lake partitions, caches, and SaaS destinations.
- Include scheduled jobs, event-driven functions and manual export processes. Manual exports are the ones most often forgotten.
Check controls hop by hop
For each hop on a sensitive route:
- Encryption at rest is on, and the key policy is no weaker than at the source.
- Transport between hops is encrypted.
- Read access is no broader than the data's sensitivity justifies.
- Public access is blocked.
- Logging is enabled, so you can see who read the data at that hop.
- A retention rule exists, especially for staging and intermediate stores.
Look for widening and crossing
- Flag any hop where the set of identities that can read the data grows sharply compared with the source.
- Flag any hop that moves data into another account, another cloud, or another region.
- Flag any hop that writes to a store with no owner.
Grade the route, not just the resources
Give each route a single risk grade based on its weakest hop. A route with four strong hops and one weak one is a weak route. Ranking by route rather than by store also cuts duplicate work: one finding about one flow, instead of separate tickets for each store it touches.
Fix at the point of creation
Most weak hops are created by pipeline code or infrastructure templates. Fix the job definition or the template that creates the staging location, not just the instance of it you found today, or it will be recreated on the next run.
Re-check after every pipeline change
New jobs, new destinations and new consumers change the route. Treat a new downstream destination for a sensitive source as something that should trigger a review.
A short worked example
A customer table lives in an encrypted managed database. A nightly job exports it to an object storage folder, then loads it into a warehouse.
- Source: encrypted, private, access limited to the application role.
- Export folder: encrypted with a default key, readable by a broad data engineering role, no lifecycle rule.
- Warehouse: encrypted, readable by the analytics group.
Every resource passes a basic encryption check. The route still has a problem: the export folder keeps every night's full copy forever, readable by more identities than the source. The fix is a lifecycle rule and a narrower role on the export location, made in the template that creates it.
How Onam approaches it
Data lineage is part of Onam's Data Security engine, its DSPM implementation. Onam links the relationships your cloud's configuration already records — replication, backup, ETL writes, streaming, export and import — into chains that start at the original source and follow each path up to eight hops. Every hop that crosses a region or an account is flagged, because those are the moments a copy leaves the controls on the original.
Each store in a chain carries its own DSPM checks — encryption, public access, logging, lifecycle — and its metadata-based classification, so the route and the state of each stop on it can be read together. Lineage sees what configuration describes: a copy made by a script with its own credentials leaves no relationship to follow. How chains are built is set out in the lineage documentation.
To see the chains in your own environment, request a demo or start a 14-day trial.