OpenAI Introduces Invisible Watermarks for Generated Text in the EU

In response to the EU AI Act, OpenAI will begin implementing invisible digital watermarks into text generated by its models. The aim is to enhance transparency and enable the identification of content created by AI systems.

OpenAI Introduces Invisible Watermarks for Generated Text in the EU

OpenAI has announced it will start adding invisible digital watermarks to texts generated by models like ChatGPT and Codex. This measure applies to users in the European Union and is a response to the new EU AI Act, whose transparency rules came into effect on August 2nd. The goal is to make AI-generated content recognizable by other systems.

The rollout of these watermarks will occur in the coming weeks for eligible users within the EU. Interestingly, developers using the OpenAI API anywhere in the world can activate this feature for selected models now, although it is off by default. The company does not currently plan to implement watermarks globally as a default setting.

The watermarking mechanism is not a traditional graphical symbol. It operates on the principle of subtly influencing the model's word choice, creating a pattern that the human eye does not detect, but a detection system can recognize. A key feature is that this "imprinted" pattern is an inseparable part of the text and remains preserved even when copied and pasted. OpenAI emphasizes that the watermark does not identify a specific user, and its activation has had no measurable impact on model performance.

OpenAI has also published a detailed technical report on this technology, named "textGrain," developed in collaboration with researchers from the University of Pennsylvania and Yale University. The method uses a secret key to arrange predictions of the next word. By accumulating hundreds of these subtle "hints," a detector can identify AI-generated content, solely based on the text itself and the given key.

Dominik Medal

Got an idea for a website or app?

Let's talk it through — non-binding, with no pressure, and a concrete next step.

However, OpenAI's tests have shown that the watermark can be removed by editing. For example, replacing 10% of words with synonyms reduced detection success from approximately 92% to 66%. Detection is also more difficult with short text segments, mathematical answers, and translated content. Precisely because of these limitations, initial access to the detector will be granted only to approved researchers and expert organizations who will help evaluate its reliability and responsible use.

The company also warns that the absence of a watermark does not automatically imply human authorship. The text could have been too short, significantly edited, or originated from another company's AI system. Watermarks may thus indicate that an OpenAI system was involved in generating or processing part of the text, but they say nothing about the degree of human judgment, editing, or creativity put into the work.

This move comes several months after Anthropic announced the global implementation of watermarks for text generated by its Claude model, which elicited mixed reactions from some users.

Dominik Medal

Let's talk about your project

Tell me what you need and together we will work out the best way forward.

Stop scrolling, call me

+420 735 505 585