PDM

HIPAA Safe Harbor coverage

HIPAA’s Safe Harbor method requires removing 18 categories of identifiers to de-identify health data. Here’s PDM’s coverage of each one.

PDM detects and masks the Safe Harbor identifier categories in structured columns and in free text — the de-identification step Safe Harbor calls for. Coverage is measured separately for each, and where a category has only been measured on one of the two, the row below says so. Files with clinical structure automatically receive clinical identifier coverage, with a Business Associate Agreement attestation step before processing.

1.Names

Measured

Structured columns 100% (640 values). Free text: proper case 96.3% (285 / 296); lowercase 57.4–58.2% (162–164 / 282).

2.Geographic units smaller than a state

Measured

Structured address columns 100% (357 values). Free text: neighbourhoods and informal locations 13.6% (23 / 169). Standard-form street addresses in free text are not measured — the test data plants them in structured columns only.

3.Dates — birth, admission, discharge, service

Measured

Structured date columns 100% (645 values). Free text 99.7% (319 / 320).

4.Telephone numbers

Measured

Structured columns 100% (345 values). Free-text coverage is not measured — the test data plants telephone numbers in structured columns only.

5.Fax numbers

Measured

Free text 100% (175 / 175). Tested against one format: (999) 999-9999.

6.Email addresses

Measured

Structured columns 100% (159 values). Free text 100% (91 / 91), across 37 distinct formats.

7.Social Security numbers

Measured

Structured columns 100% (271 values). Free text 100% (103 / 103). Tested against one format: 999-99-9999.

8.Medical record numbers (MRN)

Measured

Structured columns 100% (215 values). Free text: with a nearby label 100% (83 / 83), tested as seven digits; bare digits with no label 48.8% (21 / 43).

9.Health-plan beneficiary numbers

Measured

Structured columns 100% (105 values). Free-text coverage is not measured separately.

10.Account numbers

Measured

Structured columns 100% (1,013 values across account, routing, IBAN and SWIFT). Free text 84.4% (626 / 742) for those four types combined.

11.Certificate / license numbers

No reliable method

No standard format distinguishes these from part numbers, lot codes and other organization-specific references. Not detected: 0 of 313 planted values, together with row 13.

12.Vehicle identifiers (VIN)

Measured

Free text 100% (201 / 201), across 200 distinct formats.

13.Device identifiers / serial numbers

No reliable method

No standard format distinguishes these from part numbers, lot codes and other organization-specific references. Not detected: 0 of 313 planted values, together with row 11.

14.Web URLs

Measured

Free text 100% (221 / 221), across 25 distinct formats.

15.IP addresses

Measured

Free text 99.5% (206 / 207). The single miss is an address whose leading octet was masked and whose remainder was not.

16.Biometric identifiers

Out of scope

Not present in a text export.

17.Full-face photographs

Out of scope

Image modality, not text.

18.Any other unique identifying number or characteristic

Out of scope

A judgment standard for the covered entity's Expert Determination, not a format a tool can check.

An identifier counts as detected only when the entire value is masked. Partial masks count as misses.

Structured columns are confirmed and masked at the column level: once a column is identified as containing a given identifier type, every value in it is masked. A structured figure therefore reflects column confirmation rather than per-value detection.

Where a figure reads 100%, the number tested is shown beside it, and the formats tested are named where a category was tested against only one.

Full methodology and benchmark results available on request.

See also: detection accuracy · BAA terms.

Beyond HIPAA: GDPR and CCPA

The same masking that supports Safe Harbor de-identification supports GDPR pseudonymisation and CCPA deidentification practices.

Masking supports GDPR pseudonymisation and data-minimisation practices — the safeguards Articles 25 and 32 call for when sharing personal data.

Masking supports CCPA deidentification and data-minimization practices — reducing what your shared files carry under California and similar state privacy laws.

More on the sealed perimeter and controls on /trust.