We use cookies and similar tools to understand how our site is used and to improve your experience. This may include information about pages visited, browser and device settings and approximate location. We use this data for site analytics and performance. By continuing to use this site, you agree to the use of cookies.

Is Text De-identification Enough for AI Model Training?

Removing names and other obvious identifiers is important. But it does not show that the remaining re-identification risk is acceptable for a specific AI use. That requires a separate assessment of what may have been missed, what can still be linked to an individual and the controls around the data.

Text can be one of an organization's richest data sources. Emails, transcripts, clinical notes, case files, support tickets, survey responses and other narrative records contain context and detail that structured databases often miss.

They can also be difficult to use responsibly for AI model training. Personal information can appear anywhere in a document, and details that seem harmless on their own can identify someone when combined.

So, is running the text through a de-identification tool enough?

No. A de-identification tool can find and redact, replace, mask or re-synthesize identifiers. But transformation alone does not establish that residual re-identification risk is low enough for the intended use. That requires a separate risk assessment.

EviData automates that assessment, so it can be completed in hours rather than months. This allows multiple rapid iterations to reduce the risk, scale the assessment and integrate it within a data pipeline.

Why can privacy risk remain after text is de-identified?

Text behaves differently from a structured database. In a table, you know which field contains a date of birth or postal code. In a document, identifying information may appear anywhere and in many forms.

The assessment must consider two broad types of information:

  • Direct identifiers, such as names, email addresses, identification numbers and full residential addresses.
  • Indirect identifiers, such as dates, locations and other details that may identify someone when combined.

Removing direct identifiers does not resolve the risk created by indirect identifiers. Tools can also miss information. If a name appears three times and only two occurrences are detected, the remaining occurrence is still exposed. That leakage must be measured and included in the risk assessment.

What should a text risk assessment measure?

Risk evaluation for text can build on established approaches for structured data, but it must also account for how well the tool finds personal information in the documents.

A defensible assessment should examine three things:

1. Information leakage. Precision and recall are common measures of tool performance. For privacy, recall is especially important because it shows how many identifiers present in the text were actually detected. Missed information remains exposed.

2. The vulnerability of the remaining text. The assessment must consider whether details left, that information leakage, in the documents could be linked to a person using other available information.

3. The context in which the text will be used. The recipient, intended use and security, privacy and contractual controls all affect the likelihood of re-identification.

Recall is important, but it is not an overall risk score. Tool performance must be considered alongside the vulnerability of the data and the controls around its use to estimate residual re-identification risk against an accepted threshold appropriate to that use.

Does every document need to be reviewed manually?

A traditional approach is to evaluate a representative sample. Human reviewers tag the direct and indirect identifiers to create a ground truth, and the de-identification tool is run on the same sample so its performance can be compared.

If the tool performs well and residual risk is below an accepted threshold appropriate to the document type and intended use, it can be applied to the larger collection and similar future documents. The assessment should be repeated when something material changes, such as the tool itself, document type, recipient, intended use or controls.

But manual evaluation can be slow and expensive. EviData automates the evaluation, residual-risk assessment, reassessment and evidence generation rather than requiring a new manual expert exercise for every dataset.

Transformation changes the text.
Assurance shows whether the result is ready for its intended use.

What does independent assessment add?

A de-identification tool asks where identifiers are and how they should be changed. Independent assessment asks the next question: did those changes reduce re-identification risk enough for this specific use?

For AI training-data providers, that evidence can help earn confidence from data owners upstream and AI buyers downstream. It can support privacy review, procurement, data licensing and governance discussions, and travel with the data rather than remaining an unsupported vendor claim.

For text de-identification vendors, independent assessment can sit downstream as a complementary assurance layer. It can show whether transformed output is ready, needs further risk reduction and reassessment or should not proceed.

When is de-identified text ready to use?

Text is ready when residual re-identification risk is below an accepted threshold appropriate to the intended use and the basis for that decision has been documented.

This is not a permanent or universal label. The answer depends on the documents, the de-identification approach, the recipient, the safeguards and the intended use.

A controlled, non-public use can combine data transformations with security, privacy and contractual protections. Public release is more difficult because those additional controls are limited or may be absent.

De-identification is essential, but it is not the final step. Before text trains an AI model, there should be evidence that residual re-identification risk is low enough for the intended use and a documented basis for moving the data forward.

WATCH THE WEBINAR

AI training-data providers: BOOK A DEMO

Text de-identification vendors: DISCUSS A PARTNERSHIP

About the author: Dr. Khaled El Emam is founder and CEO of Woodway Assurance, professor at the University of Ottawa and senior scientist at the CHEO Research Institute.