Identify
Spot the personal information and the part genuinely needed. Detection works on known fields and formats; an identifier spelled out in a paragraph calls for a read-through.
For many tasks, personal data is not needed at all. Look into protecting it before sending anything to a model.
A made-up illustration of a substitution. The real mechanism depends on the processing and its configuration.
Camille Martin is requesting a return for order ORD-0042. Their device will not start.
[PERSON_1] is requesting a return for order [ORDER_1]. Their device will not start.
Spot the personal information and the part genuinely needed. Detection works on known fields and formats; an identifier spelled out in a paragraph calls for a read-through.
Look into detecting and replacing identifiers, within the options available. The method is decided task by task: an internal summary and published content do not call for the same level.
Review what remains and the scope for re-identification. A check on a real sample is worth more than a promise about the configuration: that is where the edge cases show up.
Reducing the data used to what the task strictly requires.
Replacing identifiers while potentially keeping re-identification possible.
Preventing re-identification effectively and lastingly: a requirement to assess rigorously.
A check is still needed. Indirect information, an unusual context or a poor-quality document can leave scope for identification.
The formats, the categories of information, the quality of the extraction and the expected result all have to be checked. The demonstration is where the terms that apply to your case are framed.
We can start from a made-up example and define the protection requirements.