Data Protection

Data Discovery for GDPR: Finding Personal Data Across Your Systems

Why the record of processing, the access request, and the erasure all hang on the same question: where is the data?

Author
Andy Mura
Date
4.2.2026
Updated on
18.8.2026
Data Discovery for GDPR: Finding Personal Data Across Your Systems

Key takeaways

  • Data discovery for GDPR is the systematic identification of where personal data sits in your systems, how it got there, and who can reach it. It is not a GDPR obligation in itself. It is the precondition for at least five of them.
  • Without a current inventory, Art. 30 (records of processing), Art. 15 (access), Art. 17 (erasure), Art. 33 (the 72-hour breach notification), and Art. 32 (security) cannot be met in any defensible way.
  • The blind spot is almost never the database. It is the systems adopted without the compliance function: the marketing tool on a trial, the AI assistant on a team card, the spreadsheet on a shared drive.
  • An inventory taken once a year is between zero and twelve months out of date when you need it, six on average. Art. 33 gives you 72 hours to state which categories of data were affected.
  • The Art. 30(5) exemption from the record-keeping duty still applies only below 250 employees, and falls away anyway where processing carries a risk to rights and freedoms, is not occasional, or involves special categories. The proposed extension to 750 employees under Omnibus IV was not adopted as of August 2026.

What is data discovery, and why does the GDPR require it?

Data discovery is the process of locating personal data across every system a company uses, classifying it, and tracing how it moves. The output is not a list of systems. It is an answer to three questions: what personal data do we process, where does it sit, and who has access to it.

The GDPR never uses the term. It assumes the result. Art. 5(2) makes you accountable, which means demonstrating that you comply with the principles, and you cannot demonstrate anything about processing you do not know about. Our complete guide to the GDPR places that accountability duty alongside the regulation's other obligations.

In practice this makes data discovery an inventory problem rather than a legal one. Like any inventory, it reflects the day it was taken.

Which GDPR duties are unachievable without it

Five obligations fail directly on an incomplete inventory. Not because the law demands an inventory, but because the required answer does not exist without one.

Duty Article What is missing without an inventory
Records of processing activities Art. 30 Processing nobody reported never reaches the record. Supervisory authorities ask for the record first.
Right of access Art. 15 The copy of the data and the list of actual recipients, both inside one month.
Erasure Art. 17 The locations. An erasure that reaches three systems out of five is not an erasure.
Breach notification Art. 33 Which categories of data and how many people were affected, within 72 hours of becoming aware.
Security of processing Art. 32 The risk basis. Security appropriate to the risk assumes you know what you are protecting.

The link between Art. 30 and Art. 15 is the one most teams underrate. Since the CJEU's judgment of 12 January 2023 (C-154/21), an access request entitles the person to the actual identity of the recipients, not merely categories of recipient. That list comes from the record of processing, and the record comes from the inventory. How to run an access request end to end, from the deadline calculation to redacting third parties, is covered in our guide to handling a DSAR under the GDPR.

Where personal data actually sits

Discovery rarely fails on the systems you think of. CRM, HR, and finance appear in every record. What gets missed are the stores that accumulated on the side.

Location Typical examples Why it gets missed
Structured business systems CRM, ERP, HR system, ticketing Rarely missed as systems, but the free-text fields inside them are
Marketing and sales tools HubSpot, Mailchimp, mailing lists, webinar platforms Bought by the function, often on a trial, and never decommissioned
Collaboration and storage Shared drives, Notion, Confluence, chat history Unstructured, searchable only with effort, and growing unobserved
Mailboxes Personal and shared mailboxes, archives Almost always contain third-party data, and they are the most expensive item in any access request
Analytics BI tools, data warehouse, product analytics Assumed to be anonymous, and frequently are not once an individual can be singled out
Backups and copies Backups, test environments holding production data, exports on endpoints Fall through erasure because nobody owns them
AI tools Assistants, transcription services, code assistants New, adopted quickly, and frequently without a processor agreement

The fastest-growing category is the last one. A transcription service that sits in customer calls processes personal data on both sides of the conversation and needs a contract under Art. 28. It is usually adopted by whoever has to write up the meeting, not by the compliance function. How to surface systems nobody approved is covered on the Tool and Data Discovery page.

An illustrative figure for scale, worked through for a 60-person B2B SaaS company: IT approval lists 14 systems. The first complete discovery finds 23. The nine extra are not a policy breach, they are the ordinary residue of three years of growth in which teams adopted tools that each looked harmless. Of those nine, six need a processor agreement and three process data outside the EU.

The data discovery process in five steps

Step What to do Output
1. Find the systems Technical detection through single sign-on, network traffic, and expense records, plus interviews with the functions The complete system list, not the approved one
2. Classify the data Per system, establish which categories of personal data are present and whether special categories under Art. 9 are among them Data category mapped to location
3. Build processing activities Group systems into business activities with a purpose and a legal basis under Art. 6(1) Draft record of processing under Art. 30
4. Establish recipients and transfers Identify processors, sub-processors, and third-country exposure per activity Recipient list for Art. 15, review list for Chapter V
5. Keep it current Derive changes from connected systems instead of surveying once a year A record that is accurate on the day it is asked for

Step 3 is where purely technical tooling stops. A scanner finds a field containing a phone number. Whether that field belongs to the processing activity "order fulfillment" on the contractual basis of Art. 6(1)(b), or to direct marketing under legitimate interests at Art. 6(1)(f) with a documented balancing test, is not something it can decide. That assignment is a professional judgment and stays one.

Three mistakes that make an inventory worthless

First, treating it as a one-off. When a supervisory authority asks for the record, or a breach has to be notified, what counts is the position today, not the position at the last survey. The difference between an inventory and a maintained record is not thoroughness, it is frequency.

Second, looking only at structured data. Databases can be scanned; mailboxes and shared drives cannot be scanned to the same standard. Yet that is exactly where the effort concentrates in an access request, because that is where third-party data has to be redacted by hand.

Third, mistaking discovery for assessment. A tool finds systems. It does not decide which processing activities follow, which legal basis carries them, what retention period applies, or whether a data protection impact assessment under Art. 35 is required. An exported tool report is a system list, not a record of processing.

A fourth point belongs here even though it is not strictly a mistake. Discovery regularly surfaces processing with no legal basis at all. That is uncomfortable, and it is the entire purpose of the exercise. Processing you do not know about is processing you can neither justify nor stop.

How Kertos handles data discovery

Kertos pairs the technical detection with the professional assessment, because separating the two is precisely where inventories stall. Detection surfaces the systems actually in use, including the ones nobody approved, and maps the categories of personal data each one holds; the Tool and Data Discovery page shows how that works.

The record of processing, the technical and organizational measures, and the processor agreements are then built from that same data set rather than maintained as three separate files, with changes derived from connected systems so the position does not age between annual surveys. The RoPA and GDPR documentation page covers how those three come together, and the wider privacy management system is what keeps them on one basis.

Certified experts handle the assignment work, and where the role cannot be filled internally, as an external data protection officer.

To find out how many systems a discovery run would surface in your environment, book a call.

Frequently asked questions

What is data discovery in data protection?

Data discovery is the systematic identification of which systems in a company process personal data, which categories are involved, and who has access. It is the basis for the record of processing activities required by Art. 30 GDPR, and therefore for demonstrating accountability under Art. 5(2).

Is data discovery required by the GDPR?

The GDPR does not name it as a requirement. It does require outcomes that cannot exist without it: a complete record under Art. 30, a copy of the data plus a list of actual recipients under Art. 15, an effective erasure under Art. 17, and a statement of the categories of data affected within 72 hours under Art. 33.

How often should you repeat a discovery run?

Continuously rather than periodically. An annual survey leaves you with a position that is six months out of date on average when it matters. It is more useful to derive changes from connected systems, and to trigger discovery additionally on every new tool, every new supplier, and any significant organizational change.

What is the difference between data discovery and shadow IT?

Shadow IT refers to systems in use without the approval of the IT or compliance function. Data discovery is the broader process of locating personal data regardless of whether the system was approved. Shadow IT is therefore one output of data discovery, and usually its least comfortable one.

Is an automated tool enough for GDPR documentation?

No. A tool finds systems and data fields. Assigning them to a processing activity, selecting the legal basis under Art. 6(1), setting the retention period, and deciding whether a data protection impact assessment under Art. 35 is required are professional judgments. Automation shortens the discovery; it does not replace the assessment.

The Founder's Guide about NIS2: Prepare your company Now before

Protect your startup: Discover how NIS2 can impact your business and what you need to consider now. Read the free white paper now!

Ready, your compliance to put on autopilot?
Andy Mura

Andy Mura

Head of Marketing

Andy Mura is Head of Marketing at Kertos, where he leads growth strategy for the company's compliance automation platform. A marketer and growth strategist by trade, he has spent years working in highly regulated industries such as payments, which is where his interest in compliance, data privacy, and information security first took root. That foundation has since been sharpened by extensive field research and by ongoing conversations with the CISOs and IT security leaders Kertos serves as customers. He writes about the practical realities of building and running security and compliance programs, drawing on what practitioners tell him works and what does not.

About Kertos

Kertos is the modern backbone of the data protection and compliance activities of scaling companies. We enable our customers to implement integrated data protection and information security processes in accordance with GDPR, ISO 27001, TISAX®, SOC2 and many other standards quickly and cheaply through automation.

Ready to simplify GDPR compliance?

CTA Image

📅 Schedule Your 5min Compliance Check

Please enter your business email to continue. We require a company email address to ensure we can best serve your organization.

📞 5min Compliance Check