DSPM implementation checklist: a practical rollout for cloud data
Most teams do not decide to buy data security posture management in the abstract. Something forces the question. An auditor asks for a map of where customer data lives. A data export turns up in a bucket nobody remembers creating. A new analytics platform arrives and nobody can say which identities can read it. DSPM is the discipline of answering those questions continuously instead of once, in a panic, the week before an audit.
This post is the checklist we would use to roll it out. It is deliberately tool-neutral until the last section.
The problem DSPM exists to solve
Cloud data does not stay where it was put. A single production database becomes read replicas, warehouse copies, nightly exports to object storage, a staging copy for testing, and a cache that outlived the feature it served. Each copy is created by a different team for a good reason, and each one inherits whatever access and encryption settings were convenient on the day.
Three things go wrong at once:
- Nobody has a complete inventory. The data stores the security team knows about are a subset of the ones that exist.
- Sensitivity is unknown. A bucket called
exports-2024might hold marketing images or a full customer table. - Access is decided elsewhere. Who can read a store is the product of identity policies, resource policies and network paths, none of which live next to the data.
DSPM brings those three together: what data you have, how sensitive it is, and who or what can reach it.
How teams usually handle it today
Before a DSPM programme, data security is usually a mix of:
- A spreadsheet of "crown jewel" databases maintained by hand and out of date within a quarter.
- Per-service checks from the cloud provider: is this bucket public, is this volume encrypted.
- Periodic manual reviews of who has access to production data, often driven by a compliance deadline.
None of these is wrong. The gap is that each answers one question about one resource. Exposure is usually a combination: a store that is encrypted, private, and still readable by an over-privileged role that is itself reachable from the internet.
The rollout checklist
Work through these in order. Each step makes the next one cheaper.
1. Define what "sensitive" means for you
Before scanning anything, agree on the categories you care about. Most organisations start with:
- Personal data (names, emails, national identifiers, dates of birth)
- Health data
- Payment card data
- Secrets and credentials stored as data (keys in config exports, tokens in logs)
Write down which regulations and contracts apply to each. That list decides your priorities later.
2. Inventory every data store, across every account
Cover object storage, managed relational databases, NoSQL stores, warehouses, data lakes, block and file volumes, and any SaaS data platforms you run. Include non-production accounts: test copies of production data are one of the most common places sensitive data ends up with weak controls.
Checklist for this step:
- Every cloud account and subscription is connected, not just production.
- Every region is covered, including ones you do not think you use.
- Each store has an owner, even if the first owner is "unknown, needs triage".
3. Classify, and record your confidence
Classification can be done by reading contents or by reading metadata: names, tags, schemas, column names and configuration. Each has trade-offs. Content inspection is more precise but means a tool reads your data. Metadata classification is less intrusive and faster, but it can be wrong when names are vague.
Whatever approach you use, keep a confidence level on each label. A table with columns named ssn and dob is high confidence. A bucket called data-backup is not, and should be flagged for a human to confirm.
4. Map who can reach each sensitive store
This is the step most programmes skip, and it is the one that matters most. For each store labelled sensitive, answer:
- Which human users, roles and service accounts can read it?
- Which can write or delete?
- Is it reachable from the internet, directly or through a public endpoint?
- Are there cross-account or external principals in that list?
Access is rarely granted on the store itself. It comes through role policies, group memberships and trust relationships, so this step needs identity analysis, not just a resource check.
5. Check the basic controls on every sensitive store
For each sensitive store, verify:
- Encryption at rest is enabled, with a key you control where policy requires it.
- Encryption in transit is enforced.
- Public access is blocked unless there is a documented reason.
- Access logging is on, so you can answer "who read this?" after the fact.
- Retention and lifecycle rules match your data retention policy.
- The store is in a region your residency commitments allow.
6. Prioritise by exposure, not by count
You will find more issues than you can fix in a sprint. Rank them by combining sensitivity, reachability and identity:
- Sensitive data reachable from the internet.
- Sensitive data readable by identities that are themselves over-privileged or unused.
- Sensitive data without encryption or logging.
- Everything else.
7. Make it continuous
A one-off DSPM assessment is out of date within weeks. Set a cadence for re-classification, alert on new stores and on permission changes to sensitive ones, and track the trend, not just the current count.
Common mistakes
- Scanning production only. Copies in development and test accounts are where weak controls live.
- Treating classification as the finish line. A labelled store with unknown access is still an unknown risk.
- Fixing per bucket instead of per root cause. If one template creates every export bucket, fix the template.
- No owner. A finding without an owner does not get fixed.
How Onam approaches it
Onam's Data Security engine is its DSPM implementation. It inventories data stores across your connected clouds through read-only posture roles — object storage, managed databases and warehouses, streams and Kubernetes secrets — plus self-hosted databases, including Snowflake, that you onboard.
Classification is metadata-based: Onam labels each store PII, PHI, PCI, financial or confidential from its name, description, tags and schema, without reading the contents. Because the labels come from names and tags, the reason is readable in the store's own metadata, and a missing label is fixed with a tag.
Each store then carries the ways it can be reached: the grants that make it public, the other accounts its bucket policy lets in, the principals seen accessing it in the last 30 days, and the attack paths on the shared security graph that end at it. Alongside that you get encryption, logging, lifecycle and residency checks, a 0–100 governance score per store, and data lineage across replication, backup, ETL, streaming and export hops. Everything is re-evaluated on every scan. The detail is in the DSPM documentation.
If you want to see what this looks like on your own accounts, a 14-day trial is the quickest way to find out.