“PII scan” sounds simple.
Search for email.
Search for phone.
Search for name.
Done.
Except real production databases are not that clean.
A field called metadata can contain names, phone numbers and addresses. A field called customer_ref can become personal data when it is linkable to an individual. A JSON document can carry an entire profile. A table created during a migration can preserve years of old records.
A useful PII discovery process combines schema intelligence and content inspection.
Step 1: Inventory schemas and tables
Start broad.
Collect:
- Database
- Schema
- Table
- Column
- Type
- Row count
- Owner
- Environment
This gives you the map before you inspect values.
Step 2: Use column-name heuristics
High-signal names include:
email
phone
mobile
address
dob
passport
pan
aadhaar
upi
account_number
But heuristics are not proof.
A column called contact might be an email address.
A column called ref might be an internal non-personal identifier.
Use heuristics to prioritise scanning.
Step 3: Inspect representative values
Sample carefully.
Look for:
- Email patterns
- Phone patterns
- PAN patterns
- Aadhaar-like values with appropriate validation
- Names and addresses
- Free-text personal data
For Indian identifiers, format validation should be more sophisticated than simple regex where a standard validation method exists.
Step 4: Scan JSON and unstructured fields
This is where basic scripts fail.
Example:
{
"profile": {
"full_name": "Asha Rao",
"contact": "9876543210"
}
}
The column can still be named payload.
Use parsers and entity-recognition techniques for ambiguous fields.
Step 5: Treat identifiers as contextual
PII is not only obvious government identifiers.
A random-looking ID can become personal data if it can be associated with an identifiable person.
So a good scanner should understand relationships.
user_id
↓
users table
↓
email
The user ID may be useful for tracing personal data even when the identifier itself is not obviously identifying.
Step 6: Scan production safely
Never run an uncontrolled full-table export just to inspect for PII.
Use:
- Metadata-first discovery
- Sampling
- Read-only credentials
- Rate limits
- Redaction
- Minimal retention of scan results
The scanner itself should become a privacy control, not another place where sensitive data accumulates.
Step 7: Find the hidden copies
Once you identify a production field, look for downstream copies:
Postgres
↓
warehouse
↓
analytics
↓
CRM
↓
logs
The production database is often the beginning of the data map, not the end.
What a good finding looks like
Bad finding:
1,238 PII records found.
Useful finding:
System: production.users
Column: email
Category: Contact
Purpose: account communication
Owner: Platform
Recipients: email provider
Retention: defined policy
Deletion: account workflow
Last scanned: 2026-09-04
Now someone can act.
Common mistakes
Only scanning column names.
Full database exports.
No context.
No re-scan after remediation.
Ignoring staging and backups.
Where Privra fits
Privra uses infrastructure discovery to identify personal data where it actually lives, then connects findings to purpose, ownership, processors, retention, and remediation.
The output is not “we found some emails.”
It is:
Here is where personal data exists, why it is there, and what should happen next.