Key takeaways
- Data discovery for GDPR is the systematic identification of where personal data sits in your systems, how it got there, and who can reach it. It is not a GDPR obligation in itself. It is the precondition for at least five of them.
- Without a current inventory, Art. 30 (records of processing), Art. 15 (access), Art. 17 (erasure), Art. 33 (the 72-hour breach notification), and Art. 32 (security) cannot be met in any defensible way.
- The blind spot is almost never the database. It is the systems adopted without the compliance function: the marketing tool on a trial, the AI assistant on a team card, the spreadsheet on a shared drive.
- An inventory taken once a year is between zero and twelve months out of date when you need it, six on average. Art. 33 gives you 72 hours to state which categories of data were affected.
- The Art. 30(5) exemption from the record-keeping duty still applies only below 250 employees, and falls away anyway where processing carries a risk to rights and freedoms, is not occasional, or involves special categories. The proposed extension to 750 employees under Omnibus IV was not adopted as of August 2026.
What is data discovery, and why does the GDPR require it?
Data discovery is the process of locating personal data across every system a company uses, classifying it, and tracing how it moves. The output is not a list of systems. It is an answer to three questions: what personal data do we process, where does it sit, and who has access to it.
The GDPR never uses the term. It assumes the result. Art. 5(2) makes you accountable, which means demonstrating that you comply with the principles, and you cannot demonstrate anything about processing you do not know about. Our complete guide to the GDPR places that accountability duty alongside the regulation's other obligations.
In practice this makes data discovery an inventory problem rather than a legal one. Like any inventory, it reflects the day it was taken.
Which GDPR duties are unachievable without it
Five obligations fail directly on an incomplete inventory. Not because the law demands an inventory, but because the required answer does not exist without one.
The link between Art. 30 and Art. 15 is the one most teams underrate. Since the CJEU's judgment of 12 January 2023 (C-154/21), an access request entitles the person to the actual identity of the recipients, not merely categories of recipient. That list comes from the record of processing, and the record comes from the inventory. How to run an access request end to end, from the deadline calculation to redacting third parties, is covered in our guide to handling a DSAR under the GDPR.
Where personal data actually sits
Discovery rarely fails on the systems you think of. CRM, HR, and finance appear in every record. What gets missed are the stores that accumulated on the side.
The fastest-growing category is the last one. A transcription service that sits in customer calls processes personal data on both sides of the conversation and needs a contract under Art. 28. It is usually adopted by whoever has to write up the meeting, not by the compliance function. How to surface systems nobody approved is covered on the Tool and Data Discovery page.
An illustrative figure for scale, worked through for a 60-person B2B SaaS company: IT approval lists 14 systems. The first complete discovery finds 23. The nine extra are not a policy breach, they are the ordinary residue of three years of growth in which teams adopted tools that each looked harmless. Of those nine, six need a processor agreement and three process data outside the EU.
The data discovery process in five steps
Step 3 is where purely technical tooling stops. A scanner finds a field containing a phone number. Whether that field belongs to the processing activity "order fulfillment" on the contractual basis of Art. 6(1)(b), or to direct marketing under legitimate interests at Art. 6(1)(f) with a documented balancing test, is not something it can decide. That assignment is a professional judgment and stays one.
Three mistakes that make an inventory worthless
First, treating it as a one-off. When a supervisory authority asks for the record, or a breach has to be notified, what counts is the position today, not the position at the last survey. The difference between an inventory and a maintained record is not thoroughness, it is frequency.
Second, looking only at structured data. Databases can be scanned; mailboxes and shared drives cannot be scanned to the same standard. Yet that is exactly where the effort concentrates in an access request, because that is where third-party data has to be redacted by hand.
Third, mistaking discovery for assessment. A tool finds systems. It does not decide which processing activities follow, which legal basis carries them, what retention period applies, or whether a data protection impact assessment under Art. 35 is required. An exported tool report is a system list, not a record of processing.
A fourth point belongs here even though it is not strictly a mistake. Discovery regularly surfaces processing with no legal basis at all. That is uncomfortable, and it is the entire purpose of the exercise. Processing you do not know about is processing you can neither justify nor stop.
How Kertos handles data discovery
Kertos pairs the technical detection with the professional assessment, because separating the two is precisely where inventories stall. Detection surfaces the systems actually in use, including the ones nobody approved, and maps the categories of personal data each one holds; the Tool and Data Discovery page shows how that works.
The record of processing, the technical and organizational measures, and the processor agreements are then built from that same data set rather than maintained as three separate files, with changes derived from connected systems so the position does not age between annual surveys. The RoPA and GDPR documentation page covers how those three come together, and the wider privacy management system is what keeps them on one basis.
Certified experts handle the assignment work, and where the role cannot be filled internally, as an external data protection officer.
To find out how many systems a discovery run would surface in your environment, book a call.
Frequently asked questions
What is data discovery in data protection?
Data discovery is the systematic identification of which systems in a company process personal data, which categories are involved, and who has access. It is the basis for the record of processing activities required by Art. 30 GDPR, and therefore for demonstrating accountability under Art. 5(2).
Is data discovery required by the GDPR?
The GDPR does not name it as a requirement. It does require outcomes that cannot exist without it: a complete record under Art. 30, a copy of the data plus a list of actual recipients under Art. 15, an effective erasure under Art. 17, and a statement of the categories of data affected within 72 hours under Art. 33.
How often should you repeat a discovery run?
Continuously rather than periodically. An annual survey leaves you with a position that is six months out of date on average when it matters. It is more useful to derive changes from connected systems, and to trigger discovery additionally on every new tool, every new supplier, and any significant organizational change.
What is the difference between data discovery and shadow IT?
Shadow IT refers to systems in use without the approval of the IT or compliance function. Data discovery is the broader process of locating personal data regardless of whether the system was approved. Shadow IT is therefore one output of data discovery, and usually its least comfortable one.
Is an automated tool enough for GDPR documentation?
No. A tool finds systems and data fields. Assigning them to a processing activity, selecting the legal basis under Art. 6(1), setting the retention period, and deciding whether a data protection impact assessment under Art. 35 is required are professional judgments. Automation shortens the discovery; it does not replace the assessment.





