Anonymising data before it goes to the AI – is that enough?
Partly. Redacting input before it is sent noticeably lowers the risk and is better than nothing. Legally, though, it almost always remains pseudonymisation rather than anonymisation – because the mapping exists somewhere, otherwise the answer would not come back readable. That means the data stays personal data, the GDPR still applies, and the processing still takes place at an external provider. Anyone who wants to avoid that does not need better masking; they need the data not to leave the building at all.
What masking actually achieves
The method is quickly explained: before a text goes to a language model, a detection step looks for personal details – names, addresses, dates of birth, IBANs, case numbers – and replaces them with placeholders. The model works on the masked text, and the original values are put back into the answer. To the user it looks as though the AI had read the real file.
The gain is real and should not be talked down here. One fewer name in the prompt is one fewer name in the provider's logs. Against the most common everyday mistake – someone pastes a complete letter to a customer into a chat window – automatic masking helps more reliably than any training session.
But the decisive question is not whether it helps. It is: Is it enough to meet the obligations that made you think about this in the first place? And there the answer gets less tidy.
Where the limit runs
A detection step that finds personal references in free text works with probabilities. It is imperfect not because it is badly built, but because the task itself has no clear-cut solution: whether a string identifies someone depends on context, not on its form.
That is why providers of such methods quote detection rates rather than guarantees. A rate of 95 per cent sounds high – but it means that one in twenty occurrences is not recognised as one. At a thousand pages of correspondence a month, that stops being a theoretical figure.
That last row is the important one, and it leads straight to the legal core.
Pseudonymisation is not anonymisation
The GDPR draws a sharp line between the two. Data is anonymous only if nobody can re-establish the link to a person with reasonable effort – under recital 26, such data falls outside the regulation entirely. Data is pseudonymised if it can be re-attributed using additional information (Art. 4(5) GDPR). Pseudonymous data remains personal data, and every obligation continues to apply unchanged.
A method that sets placeholders and resolves them again in the answer is by definition the second. The mapping has to exist, otherwise the product does not work. So it remains processing of personal data – with a data processing agreement, a record of processing activities, a legal basis and, where applicable, a data protection impact assessment.
For professionals bound by confidentiality under section 203 of the German Criminal Code – law firms, medical practices, advisory professions – that shifts the question once more. There, disclosure to a third party is already the critical step, regardless of how well the name was redacted. Masking can soften that step, but it does not undo it.
When masking is the right choice anyway
None of this makes the method useless. There are cases where it fits exactly:
- When a particular external model is needed. Some tasks currently run best on one of the large provider models. Anyone who depends on that is considerably better off with masking than without.
- As a net against human error. Even alongside an internal solution, the temptation remains to quickly drop something into a private chat window. Enforced masking catches that.
- For material without personal data. For technical documentation, marketing copy or public tender documents, the whole discussion is moot.
The mistake is not to mask. It is to treat masking as a location decision. It changes what goes out – not whether something goes out, and to where.
The alternative: do not send it in the first place
The second route reverses the order. Instead of masking data before it leaves the building, nothing leaves the building. The model runs where the documents already are – on-premise on your own infrastructure or in a German data centre.
Several questions then dissolve rather than needing an answer. There is no detection rate, because nothing has to be detected. There is no third-country transfer and no CLOUD Act question, because no third country is involved. And answer quality does not suffer from masking: the model sees the complete file, because it is allowed to.
That is exactly how KOSMO is built. The assistant works on your own sources, names the reference for every answer, and no model is trained on your content. What that looks like day to day is under How it works; the legal assessment for cloud services under ChatGPT, Claude & Co. with customer data.
Note: This page sorts through a technical and legal pattern and is no substitute for legal advice. Whether a particular use is permissible depends on your sector, the data you process and how things are set up in the individual case – if in doubt, to be clarified with your own data protection officer or a specialist law firm.
Frequently asked questions
Is anonymisation enough to use ChatGPT in a GDPR-compliant way?
It helps, but on its own it usually is not. Methods that resolve placeholders again in the answer are legally pseudonymisation – the data stays personal data, and processing still takes place at the external provider. The data processing agreement, the legal basis and the third-country transfer question all remain.
What is the difference between anonymisation and pseudonymisation?
Anonymous data can no longer be attributed to a person by anyone with reasonable effort; under recital 26 it falls outside the GDPR. Pseudonymous data can be re-attributed using additional information (Art. 4(5) GDPR) and remains fully protected. Anyone who translates placeholders back necessarily holds that additional information.
How reliably do such methods detect personal data?
Well, but not completely. Structured details such as IBANs or dates of birth can be found safely by pattern. Harder are in-house identifiers without a fixed format and, above all, combinations: a rare diagnosis together with a town and a date identifies a person even without a name.
Does answer quality suffer from masking?
That is a trade-off. The more thoroughly you mask, the less context the model has left. In legal or medical texts, where the relationships between the people involved are the actual information, that can become noticeable.
Does this also apply to professionals bound by confidentiality?
There the situation is stricter. What is critical is already the disclosure to a third party, regardless of the quality of the redaction. Masking reduces the risk but does not remove the step. Anyone processing client, patient or advisory data should have this point checked explicitly.
Does KOSMO need anonymisation?
No. KOSMO runs on-premise or in a German data centre – the data does not leave your infrastructure, so there is no recipient to mask anything from. No model is trained on your content either.
Can you combine the two?
Yes, and in larger organisations that often makes sense. The internal assistant covers everyday work with real data; for the few cases where a particular external model is needed, masking remains the tool of choice.
An assistant where the question does not arise
KOSMO works on your own sources – on-premise or in a German data centre. Nothing is handed outside, nothing is used for training.







