Finding personal data across files

I wrote a Python script to screen more than 1,000 files for possible personal-data exposure. The useful part was giving the review a smaller set of matches to investigate.

This project comes from my résumé. The public examples use made-up records; employer and client material stays private.

Approach
Python and regular expressions
Scope
TXT logs, PDF reports and XLSX files
Result
Reduced manual screening effort

Why I built it

Personal data can turn up in an ordinary log or report export. When the file population is large, opening each file and searching it manually takes time. It’s also easy to apply slightly different checks from one file to the next.

I built a Python scanner that used regular expressions to identify possible sensitive-data exposure across more than 1,000 unstructured files. The formats included TXT, PDF and XLSX. The project gave the review a repeatable screening step and reduced the manual effort involved.

My part in the work

My contribution was the scanner and its use in the data-privacy review. The original source files, client context and workpapers remain private. The output shown in this portfolio is a reconstruction using invented filenames and masked values.

What a match actually tells us

A match says that some text fits a rule. It doesn’t yet tell us whose data it is, whether the data belongs in that file or who can access it. I’d review the surrounding context before treating it as an observation.

For example, an identifier-like number might be a report reference. A contact field might be legitimate for the file’s purpose. If the presence is confirmed, the next questions concern necessity, access and the applicable handling or retention requirement.

The file population matters too

For a review like this, I’d keep track of the files received, the files successfully read and anything that couldn’t be processed. A clean-looking output is difficult to interpret if it silently leaves part of the population out.

I’d pay particular attention to PDF extraction. A scanned page can be an image rather than readable text, so ordinary text extraction may miss its contents. The pypdf documentation explains that limitation; it is a useful coverage question, not a claim that this project used that library or included OCR.

How I’d review the results

  1. Check that the output can be traced back to the file and the rule that produced the match.
  2. Read enough context to separate a candidate match from a false positive or an expected use of data.
  3. Record the reviewer decision and the basis for it, with sensitive values masked where the workpaper allows.
  4. Take confirmed observations into the control review, rather than treating the raw match count as the finding.

The public sample shows three different dispositions: a match needing context, a fictional false positive and confirmed data presence that still needs a handling assessment.

Read the sample output

What changed

The scanner reduced the time spent checking files manually. I prefer to explain that contribution plainly: it supported the screening procedure and helped direct the reviewer’s attention. The public example doesn’t include an accuracy study or a measured time-saving benchmark.

The part that interests me is how a small script can make an audit step more consistent while leaving the conclusion with the person reviewing the evidence.