Extracting PII Entities
openmed.extract_pii finds PHI/PII spans and returns them without changing the
text. Use it when you need to see the identifiers — to audit, route to a
custom redactor, or decide a policy — rather than produce redacted output. It runs
on-device.
When to use
- You want the spans and labels of identifiers, with the original text intact.
- You need a preview of what
deidentifywould act on before committing. - You are feeding detected spans into a downstream redactor (your own,
Presidio, or
deidentify). - You want to normalize model labels to a stable canonical taxonomy.
If you instead want redacted/masked output directly, use
deidentifying-clinical-text (openmed.deidentify). If you need reversible
masking, see reidentifying-text.
extract_pii vs deidentify
extract_pii |
deidentify |
|
|---|---|---|
| Changes the text? | No | Yes (mask/remove/replace/hash/shift) |
| Returns | PredictionResult (spans) |
DeidentificationResult (redacted text) |
| Default threshold | 0.5 |
0.7 (safety-biased) |
| Use for | detection, audit, routing | producing safe output |
Install
pip install "openmed[hf]"
Quick start
import openmed
note = "Patient John Doe (MRN 00481726), DOB 1970-01-15, phone 617-555-0142."
result = openmed.extract_pii(note, confidence_threshold=0.5)
for ent in result.entities:
print(f"{ent.label:10} {ent.text!r:18} {ent.confidence:.2f} [{ent.start}:{ent.end}]")
extract_pii(...) returns a PredictionResult. Its .entities are PIIEntity
objects (synthetic example fields shown):
ent.text # the identifier surface string, e.g. "617-555-0142"
ent.label # detected label, e.g. "PHONE"
ent.confidence # model score in [0, 1] (NOTE: .confidence, not .score)
ent.start / ent.end # character offsets into the original note
ent.canonical_label # label mapped to OpenMed's canonical taxonomy (if set)
ent.entity_type # same as label
The text is unchanged — result.text is your original input.
Signature & key parameters
openmed.extract_pii(
text,
model_name="OpenMed/OpenMed-PII-SuperClinical-Small-44M-v1", # default EN model
confidence_threshold=0.5, # raise for precision, lower for recall
use_smart_merging=True, # merge fragmented spans into whole units
lang="en", # en es pt fr de it nl hi te ar tr ja
loader=None, # reuse a ModelLoader across calls
)
use_smart_merging=True(default) reassembles fragmented predictions into complete units (a full phone number, a full date) — keep it on.langselects the language-appropriate default model and regex patterns. Pass the right language; do not run the English model on non-English text. Useopenmed.get_default_pii_model(lang)to confirm coverage.
Normalize labels to the canonical taxonomy
Different models may emit slightly different label spellings. Normalize them to OpenMed's canonical set so downstream logic is stable:
import openmed
from openmed import CANONICAL_LABELS, normalize_label
result = openmed.extract_pii("Email jane.roe@example.org; SSN 123-45-6789.")
for ent in result.entities:
canon = ent.canonical_label or normalize_label(ent.label)
assert canon in CANONICAL_LABELS or canon == "OTHER"
print(ent.text, "->", canon)
CANONICAL_LABELS is a frozenset of UPPER_SNAKE_CASE labels (e.g. PERSON,
DATE, PHONE, EMAIL, SSN, ID_NUM, LOCATION). normalize_label(label)
accepts messy inputs ("FIRSTNAME", "first_name", "B-EMAIL") and maps unknown
labels to OTHER rather than raising.
Feed spans to a downstream redactor
extract_pii gives you offsets; you decide the action. A simple offset-based
redactor (replace highest-offset first so positions stay valid):
import openmed
note = "Patient John Doe, MRN 00481726, seen 2024-03-02."
result = openmed.extract_pii(note, confidence_threshold=0.6)
redacted = note
for ent in sorted(result.entities, key=lambda e: e.start, reverse=True):
redacted = redacted[:ent.start] + f"[{ent.label}]" + redacted[ent.end:]
print(redacted) # Patient [PERSON], MRN [ID_NUM], seen [DATE].
For production redaction, masking strategies, and policy profiles, hand the work
to openmed.deidentify instead of hand-rolling — it adds a safety sweep and
date-shifting (see deidentifying-clinical-text).
Hand-off to / from OpenMed
- To
deidentifying-clinical-text: once you have reviewed the spans, callopenmed.deidentify(note, method="mask", policy="hipaa_safe_harbor")to produce safe output — it re-detects with a higher default threshold for safety. - To
reidentifying-text: if you need reversibility, useopenmed.deidentify(..., keep_mapping=True)and store the mapping securely. - To Presidio / custom anonymizers:
extract_piispans (label,start,end) translate cleanly into other recognizers' result formats; OpenMed also ships anAnonymizer(openmed.Anonymizer) for richer surrogate generation.
Edge cases & gotchas
- Attribute is
.confidence, not.score.PIIEntityextendsEntityPrediction. - Detection is not redaction.
extract_piinever changes text — if a caller expected redacted output, they wantdeidentify. - Threshold trade-off:
0.5favors recall (good for finding PHI to review). For removing PHI, preferdeidentify's safety-biased0.7default. - Language matters: wrong
langsilently lowers recall. Verify withget_default_pii_model(lang). - No raw PHI in logs/audit. Record offsets, labels, and hashes — never the identifier text. Use synthetic data in examples and tests.
- Local-first. No cloud calls in PHI workflows; models run on-device after a one-time download.
Standards & references
- HIPAA Safe Harbor 18 identifiers: 45 CFR §164.514(b)(2) — https://www.hhs.gov/hipaa/for-professionals/privacy/special-topics/de-identification/
- OpenMed canonical labels:
openmed.CANONICAL_LABELS/openmed.normalize_label. - OpenMed PII models: https://huggingface.co/OpenMed