No signup · no account · works offline

Also on Product Hunt, where questions are answered in the thread.

Safe Harbor De-identifier

Paste clinical text. It finds the mechanical HIPAA Safe Harbor identifiers and either masks them, or replaces them with realistic fakes that keep the note readable and keep the intervals between dates intact. Then you read it back as a stranger would, because no rule set can know that a rare diagnosis in a small town identifies someone.

Nothing leaves your device. This page runs entirely in your browser. There is no server call, no upload, no analytics on this page, and no storage. Every name list and medical term list it uses is embedded in this one file, and the page carries a Content-Security-Policy of connect-src 'none', which makes an outbound request impossible rather than merely absent. Turn on airplane mode and it still works. Open DevTools, watch the Network tab, and check.

Masking is one route and a business associate agreement is the other. Which one your note needs, and why taking the name out is not enough, is written up with its sources in If I take the name out, can I paste the note into ChatGPT?

1

Your text

An example is loaded so you can see what it does. Replace it with your own. Browser extensions can read this box, so turn writing and AI extensions off before you paste real text.

The treating facility is left in on purpose. A hospital name is not one of the eighteen, which cover the patient and the people around them, not the people treating them. But a named hospital plus a rare presentation can still identify someone. That call is yours.

2

Masked text, read it back

Updates as you type. Read it back before you use it.

What it found

Nothing pasted yet

Paste text on the left, or tap Example.

This does not make any AI tool HIPAA-compliant, and it does not certify de-identification. It is a first pass that catches the mechanical identifiers. Safe Harbor also requires that you have no actual knowledge the remaining information could identify the person, and a rare presentation plus a small catchment area can identify someone with every field above removed. Read the output back as a stranger would. If someone who knows this patient could recognize them, strip more or don't use the case at all.

And none of this grants permission. Your organization's policy and your vendor's contract decide whether you may use a tool at all.

Prove it: run the accuracy self-test in your own browser

A tool that claims accuracy without a number is marketing. This page carries a small corpus of synthetic notes, invented patients, invented numbers, each one annotated with the correct answer. The button below runs the same code that just processed your text against that corpus and reports what it actually caught. Nothing is sent anywhere; the corpus is in this file and you can read it in View Source.

Recall is the share of real identifier characters the tool masked. A miss is the expensive failure, so recall is what matters. Precision is the share of masked characters that were really identifiers; low precision means the tool is eating your clinical content. Certain only is what the Redact and Surrogate modes do on their own. With review is what you get if you open Review mode and accept every amber candidate too.

What is in the name lists, where they came from, and what they cost you

Embedded lists

  • Surnames 25,000, the most frequent surnames from the U.S. Census Bureau Frequently Occurring Surnames from the Census 2000 file app_c.csv, taken in count order down to a cutoff of 933 bearers. That covers 85.9% of the people in that file. Published by a federal agency; usa.gov states that a government work is one "created by a U.S. government officer or employee as part of their official duties".
  • Given names 5,494, every entry in the U.S. Census Bureau 1990 dist.male.first and dist.female.first files. The Bureau's own methodology note limits those files to "the minimum number of entries that contain 90 percent of the population in that data file", and warns that "The fact that a name doesn't appear in these three files does not mean that it is non existent, only that it is reasonably rare."
  • Medical collision list 754, words that are in the name lists and appear in the U.S. National Library of Medicine's MeSH 2025 tree file, so they are demoted to amber instead of being masked on sight. This is what stops Parkinson, Bell, Graves, Cushing and Rose from being deleted out of your assessment. Courtesy of the U.S. National Library of Medicine.
  • Ordinary-English collision list 4,985, words that are in the name lists and in the Webster's Second International word list shipped with BSD and macOS at /usr/share/dict/web2, whose README states the 1934 copyright has lapsed. Used only to hold fire at the start of a sentence, where Brown sputum and Mr Brown look identical to a machine.
  • Eponym followers 814, every word that follows a person-name in a MeSH heading (syndrome, disease, palsy, cells). A name-shaped word followed by one of these is treated as a diagnosis, not a patient, and is left alone entirely.
  • Clinical nouns 1,793, words that appear after the first word of a MeSH heading at least three times and are not themselves names. A name-shaped word followed by one of these (Foley catheter) is demoted to amber rather than masked outright. A short list of bedside objects MeSH does not name, sign, catheter, manoeuvre, is written by hand and marked as such in the source.
  • Place names 900, the most widely repeated U.S. place names in the Census Bureau 2024 Gazetteer place file, used only to invent fake cities in surrogate mode.
  • A short list of ward, department and note-heading words is written by hand for this tool and is not taken from any dataset. It is marked as such in the source.

What the lists cannot do

A gazetteer is a popularity contest. A surname held by fewer than 933 people in the United States is not in it, and no U.S. list carries the world's names. If your patient's name is rare, or not American, the list will not save you, the title, relationship and two-word-name rules are what catch it, and Review mode is where you catch the rest. That is the honest limit of this design, and it is why the output still has to be read by a human.

How surrogate mode works, and the one thing that makes it dangerous

Redaction advertises its own failures. Every [NAME] in a note tells a reader exactly where to look, so the one name the tool missed sits in the open, surrounded by markers, and it glows. Surrogates do the opposite: a missed name sits among plausible invented names and nothing marks it as special. That is the entire argument for this mode.

What stays consistent

  • The same person keeps the same fake name everywhere in the document, and a surname mentioned on its own later still maps to the surname you already saw.
  • Every date moves by one single hidden offset, so admission-to-discharge, symptom-onset-to- presentation and every other interval survives intact. That is what a researcher actually needs.
  • Record numbers, phone numbers and postcodes keep their shape, same digit count, same punctuation, so anything downstream that validates a format still works.

The dangerous part

The output looks like a real record and is not one. Never file it, never quote a surrogate value back to a colleague as fact, and never let it reach a chart. The warning banner is on by default for that reason.

The linkage key box exists for cohorts: type the same key on two different notes and both get the same date offset and the same fake names, so records about one patient still line up. That key is now a re-identification key. Anyone holding the key and the output can undo the date shift. Treat it like a password, keep it away from the de-identified data, and if you do not need cross-document linkage, leave the box empty, a fresh random offset is then generated per page load and is discarded when you close the tab.

The pocket card

The 18 identifiers and the five-step protocol on one page you can print and pin above the workstation. Free, no email, no purchase.

This is one page of the system

The Attending is a workflow system for the administrative half of medicine, 8 EHR-ready phrase templates, 100 clinical prompts each marked GREEN / AMBER / RED for patient data, seven complete workflows, and the full Safe Harbor protocol. One payment. No subscription.

The 18 identifiers

Detected automaticallyYou must check by hand
Names, after a title or a relationship word, as a two-word pair matching the embedded name lists, or as a single listed name that is not also a clinical or ordinary English word, dates, ages over 89, phone, fax, email, SSN, MRN, account and health-plan numbers, NHS and hospital numbers, license, NPI and certificate numbers, device serials, URLs, IP addresses, street address, ZIP, UK postcode, county, and city where it is written next to a state, right after a street address, or after "lives in" A city written with no state and no residence verb, biometric identifiers, full-face photographs, employer or facility names, relationship descriptors, a rare surname that is not in any U.S. name list, and any other unique characteristic or code, including a rare diagnosis in a small population

What this tool is, and what it is not

This tool applies the mechanical part of the HIPAA Safe Harbor method described at 45 CFR 164.514(b)(2)(i). It does not certify de-identification, it does not make you or any tool HIPAA compliant, and it does not create a business associate relationship. Nothing you paste is transmitted anywhere.

Detection is incomplete, by nature and by design. Category (R) of the Safe Harbor list, "any other unique identifying number, characteristic, or code", cannot be detected by any rule set. The covered entity, not this tool, makes the de-identification determination, including the actual knowledge test at 45 CFR 164.514(b)(2)(ii). Read every output before it leaves your building.

Surrogate mode does not produce Safe Harbor output. Safe Harbor requires that all elements of dates other than year be removed. Surrogate mode shifts dates by a hidden offset and keeps the month and the day so that intervals survive. Use Redact mode, which replaces dates with [DATE], when you are working toward Safe Harbor. Redact mode removes date elements; it does not by itself make a document Safe Harbor de-identified.

The linkage key is a re-identification key. Anyone holding it can undo the date shift. Keep it apart from the de-identified text.

No warranty. This tool is provided as is, without warranty of any kind, express or implied, including merchantability, fitness for a particular purpose and non-infringement. We do not warrant that it detects every identifier, or that its output satisfies any legal or regulatory standard. To the maximum extent permitted by law, total liability for any claim arising out of this tool is limited to the amount you paid for it, and we are not liable for indirect, incidental, consequential or special damages, including regulatory penalties, breach notification costs, or loss of data. Nothing here excludes liability that cannot lawfully be excluded.

Your organization's policy and your vendor contracts decide whether you may use any tool at all. Nothing here grants that permission.