DocketX / Anonymize case files before AI
Every lawyer who has pasted a client document into a chatbot has had the same second thought. The fix is not to stop using the models; it is to make sure they never receive a name. A scrambler replaces every person, address, docket and account in a matter with a stable placeholder on hardware you control, sends only the placeholders out, and restores them in the answer. Here is how it works, what it can honestly promise, and what it cannot.
The pipeline, honestly divided
Only the passages the task needs leave the file. A two-hundred-page complaint is not sent to answer a question about one cause of action.
Social security numbers, phone numbers, emails, account numbers, dates of birth and docket patterns are replaced deterministically before any model runs. These are never missed, because no model is involved.
A model running on your own hardware reads the text and replaces every remaining person, organisation and address with a typed, numbered placeholder: CLIENT_1, JUDGE_2, ADDRESS_3. It returns structured JSON and a mapping table. Nothing has left the building.
The same local model reads the scrubbed text with one question: is any real entity still here, and is any placeholder used for two different people? A model is far better at reviewing a redaction than at producing one perfectly.
The output must parse as JSON and pass a second regex sweep. Anything that fails is rejected and re-run. Nothing falls through to the raw text.
Only now does the matter go to the frontier model, as placeholders. The answer comes back and the mapping table restores every name on your side. The provider never received one.
The part vendors leave out
Replace the client's name with CLIENT_1 and the document still says she is the only cardiologist in a town of two thousand. Re-identification from residual facts is a property of the facts, not the scrubber. Anyone who promises anonymity from placeholder substitution is selling risk.
Over-scrubbing removes context the analysis needs. The trade-off between anonymity and usefulness is real and has to be tuned per matter, not hidden.
A single 'remove all personal information' prompt lands around seventy to eighty percent. The multi-pass system above is what moves that number, and the number has to be measured on real documents, not asserted.
Where this is going
DocketX already runs a private in-house model for reranking and citation work. The scrambler is a second job for that model, on the same machines, never a third-party service in the middle.
One mapping table per matter, persisted, so the same person is the same placeholder across every document of a case and every session. This is where most home-built scramblers fail.
We will publish a leak rate per thousand entities, by type, measured against documents a human already redacted. Until that number exists, we will not quote one.
We do not claim the scrambler is available today; it is designed and in build. We do not claim placeholder substitution is anonymization. We do not claim any leak rate we have not measured on real documents. If you want early access when the measurement exists, tell us.
Yes, and you should. The reliable way is a local model on hardware you control that replaces every person, address and identifier with a stable placeholder before the text leaves, then restores them in the answer. A single prompt to the same public model defeats the purpose, because the document has already been sent.
No. Pseudonymization swaps identifiers for placeholders; anonymization means the person cannot be re-identified. A matter can be fully pseudonymized and still identify someone from the facts that remain. Treat scrambling as exposure reduction, not a guarantee.
A well-prompted 27B-class model running locally can, if it is used as a pipeline: regex pre-filter, structured first pass, verification second pass, strict parsing, and a persistent per-matter mapping table. As a single prompt it cannot.
It is architected and in build on the private model DocketX already runs. It will ship with a measured leak rate or not at all.