HIPAA Safe Harbor coverage
HIPAA’s Safe Harbor method requires removing 18 categories of identifiers to de-identify health data. Here’s PDM’s coverage of each one.
PDM detects and masks the Safe Harbor identifier categories in structured columns and in free text — the de-identification step Safe Harbor calls for. Coverage is measured separately for each, and where a category has only been measured on one of the two, the row below says so. Files with clinical structure automatically receive clinical identifier coverage, with a Business Associate Agreement attestation step before processing.
| # | Identifier category | Coverage | Detail |
|---|---|---|---|
| 1 | Names | Measured | Structured columns 100% (640 values). Free text: proper case 96.3% (285 / 296); lowercase 57.4–58.2% (162–164 / 282). |
| 2 | Geographic units smaller than a state | Measured | Structured address columns 100% (357 values). Free text: neighbourhoods and informal locations 13.6% (23 / 169). Standard-form street addresses in free text are not measured — the test data plants them in structured columns only. |
| 3 | Dates — birth, admission, discharge, service | Measured | Structured date columns 100% (645 values). Free text 99.7% (319 / 320). |
| 4 | Telephone numbers | Measured | Structured columns 100% (345 values). Free-text coverage is not measured — the test data plants telephone numbers in structured columns only. |
| 5 | Fax numbers | Measured | Free text 100% (175 / 175). Tested against one format: (999) 999-9999. |
| 6 | Email addresses | Measured | Structured columns 100% (159 values). Free text 100% (91 / 91), across 37 distinct formats. |
| 7 | Social Security numbers | Measured | Structured columns 100% (271 values). Free text 100% (103 / 103). Tested against one format: 999-99-9999. |
| 8 | Medical record numbers (MRN) | Measured | Structured columns 100% (215 values). Free text: with a nearby label 100% (83 / 83), tested as seven digits; bare digits with no label 48.8% (21 / 43). |
| 9 | Health-plan beneficiary numbers | Measured | Structured columns 100% (105 values). Free-text coverage is not measured separately. |
| 10 | Account numbers | Measured | Structured columns 100% (1,013 values across account, routing, IBAN and SWIFT). Free text 84.4% (626 / 742) for those four types combined. |
| 11 | Certificate / license numbers | No reliable method | No standard format distinguishes these from part numbers, lot codes and other organization-specific references. Not detected: 0 of 313 planted values, together with row 13. |
| 12 | Vehicle identifiers (VIN) | Measured | Free text 100% (201 / 201), across 200 distinct formats. |
| 13 | Device identifiers / serial numbers | No reliable method | No standard format distinguishes these from part numbers, lot codes and other organization-specific references. Not detected: 0 of 313 planted values, together with row 11. |
| 14 | Web URLs | Measured | Free text 100% (221 / 221), across 25 distinct formats. |
| 15 | IP addresses | Measured | Free text 99.5% (206 / 207). The single miss is an address whose leading octet was masked and whose remainder was not. |
| 16 | Biometric identifiers | Out of scope | Not present in a text export. |
| 17 | Full-face photographs | Out of scope | Image modality, not text. |
| 18 | Any other unique identifying number or characteristic | Out of scope | A judgment standard for the covered entity's Expert Determination, not a format a tool can check. |
1.Names
MeasuredStructured columns 100% (640 values). Free text: proper case 96.3% (285 / 296); lowercase 57.4–58.2% (162–164 / 282).
2.Geographic units smaller than a state
MeasuredStructured address columns 100% (357 values). Free text: neighbourhoods and informal locations 13.6% (23 / 169). Standard-form street addresses in free text are not measured — the test data plants them in structured columns only.
3.Dates — birth, admission, discharge, service
MeasuredStructured date columns 100% (645 values). Free text 99.7% (319 / 320).
4.Telephone numbers
MeasuredStructured columns 100% (345 values). Free-text coverage is not measured — the test data plants telephone numbers in structured columns only.
5.Fax numbers
MeasuredFree text 100% (175 / 175). Tested against one format: (999) 999-9999.
6.Email addresses
MeasuredStructured columns 100% (159 values). Free text 100% (91 / 91), across 37 distinct formats.
7.Social Security numbers
MeasuredStructured columns 100% (271 values). Free text 100% (103 / 103). Tested against one format: 999-99-9999.
8.Medical record numbers (MRN)
MeasuredStructured columns 100% (215 values). Free text: with a nearby label 100% (83 / 83), tested as seven digits; bare digits with no label 48.8% (21 / 43).
9.Health-plan beneficiary numbers
MeasuredStructured columns 100% (105 values). Free-text coverage is not measured separately.
10.Account numbers
MeasuredStructured columns 100% (1,013 values across account, routing, IBAN and SWIFT). Free text 84.4% (626 / 742) for those four types combined.
11.Certificate / license numbers
No reliable methodNo standard format distinguishes these from part numbers, lot codes and other organization-specific references. Not detected: 0 of 313 planted values, together with row 13.
12.Vehicle identifiers (VIN)
MeasuredFree text 100% (201 / 201), across 200 distinct formats.
13.Device identifiers / serial numbers
No reliable methodNo standard format distinguishes these from part numbers, lot codes and other organization-specific references. Not detected: 0 of 313 planted values, together with row 11.
14.Web URLs
MeasuredFree text 100% (221 / 221), across 25 distinct formats.
15.IP addresses
MeasuredFree text 99.5% (206 / 207). The single miss is an address whose leading octet was masked and whose remainder was not.
16.Biometric identifiers
Out of scopeNot present in a text export.
17.Full-face photographs
Out of scopeImage modality, not text.
18.Any other unique identifying number or characteristic
Out of scopeA judgment standard for the covered entity's Expert Determination, not a format a tool can check.
An identifier counts as detected only when the entire value is masked. Partial masks count as misses.
Structured columns are confirmed and masked at the column level: once a column is identified as containing a given identifier type, every value in it is masked. A structured figure therefore reflects column confirmation rather than per-value detection.
Where a figure reads 100%, the number tested is shown beside it, and the formats tested are named where a category was tested against only one.
Full methodology and benchmark results available on request.
See also: detection accuracy · BAA terms.
Beyond HIPAA: GDPR and CCPA
The same masking that supports Safe Harbor de-identification supports GDPR pseudonymisation and CCPA deidentification practices.
Masking supports GDPR pseudonymisation and data-minimisation practices — the safeguards Articles 25 and 32 call for when sharing personal data.
Masking supports CCPA deidentification and data-minimization practices — reducing what your shared files carry under California and similar state privacy laws.
More on the sealed perimeter and controls on /trust.