5 min Security

OpenAI introduces AI watermark for texts in the EU

Applies to ChatGPT and Codex, but turns out to be a sham

OpenAI introduces AI watermark for texts in the EU

In the coming weeks, OpenAI will add an invisible watermark to text generated by ChatGPT and Codex in the European Union. With this technology, called textGrain, the company aims to comply with the AI Act. At the same time, OpenAI itself admits that the watermark is easy to circumvent.

The watermark is not visible when reading or copying the text. textGrain subtly adjusts the model’s word choice during the generation process. This creates a statistical pattern that a detector can later identify. The rollout applies to users of all subscription plans in the EU, but only to eligible output. It will not become a global standard for now.

Outside of Europe, there are fewer changes. As of yesterday, API customers worldwide can choose to enable watermarks for selected models, though the feature is disabled by default. In addition, OpenAI is working with cloud partners to make watermarks available for models purchased through their services.

European pressure and transparency under the AI Act

The immediate catalyst for this move is European regulation surrounding artificial intelligence. Under Article 50 of the AI Act and the accompanying Code of Practice, providers of generative models must make AI-generated content machine-readable and traceable. In doing so, Europe is prioritizing a mandatory label on AI content to combat misleading deepfakes and the uncontrolled spread of synthetic media. While the law mandates detectability, it deliberately avoids imposing a fixed technical standard, allowing technology providers to develop their own methods.

OpenAI is therefore not the first to take this step. For example, Anthropic recently announced that Claude will include an invisible watermark in all generated text. However, Anthropic chose to implement this marking immediately worldwide and for all new models, rather than following OpenAI’s regional approach. Google has been pursuing a similar path for some time with SynthID and has also experimented with visual markings. For example, Google previously made a visible AI watermark optional, while the underlying invisible technology continues to run. As analyses of market-wide adoption show, AI watermarks are now invisible but present virtually everywhere, with tech giants primarily embedding generative watermarks in the model’s probability distribution (logits).

Editing weakens the signal

It is striking how candid OpenAI is about the practical limitations of textGrain. In tests with 400-token passages, detection rates dropped from about 92 to 66 percent when just 10 percent of the words were replaced with synonyms. When 25 percent of the words were altered, only 17 percent of the signal remained. Text length also plays a decisive role. With a false-positive target of 1 percent, the detector found the watermark in about 80 percent of the 200-token psychology texts, compared to about 95 percent for 400-token texts. For math texts, where the model simply has less freedom in word choice, the scores were significantly lower.

“The absence of a detected watermark does not prove that a human wrote the text,” OpenAI warns. Text may be too short, heavily edited, or translated. Conversely, a detected watermark says nothing about the specific user, the account, or the prompt entered. Furthermore, it remains unclear how much a human actually contributed; someone who simply has their own text proofread or rewritten by the model will still receive the watermark.

And what about output quality? According to OpenAI, textGrain has no noticeable impact on it. The benchmark scores of the frontier model GPT-6 Astra remain comparable with and without a watermark, for example on GPQA Diamond (94.44 versus 93.94 percent).

Watermarks for the development assistant Codex as well

The watermark requirement applies to Codex as well as ChatGPT, aligning with OpenAI’s broader strategy. While the tool once served purely as an underlying model engine for code generation, OpenAI has expanded Codex into a more comprehensive development assistant. The software now functions as an autonomous agent that executes terminal commands, reviews files, and integrates with enterprise tools such as code management and ticketing systems. Because Codex not only delivers raw code but also generates documentation, pull request summaries, and to-do lists, that textual output is also subject to the regulator’s detection requirements in Europe.

Detector not public for now

Due to the real risk of missed watermarks and false positives, OpenAI is not making the text detector publicly available for the time being. Only approved researchers and expert organizations can request access to assist with further evaluation. Eventually, the company does plan to release textGrain as open source. In doing so, it is following the lead of Google DeepMind, which made SynthID Text available as open source in 2024. According to OpenAI, textGrain performs at least as well as that technology in its own comparisons.

Incidentally, this issue is not entirely new for OpenAI. It was previously reported that the company had a working method for text watermarking on the back burner, but did not activate it in ChatGPT at the time for fear of a competitive disadvantage and user resistance. For multimedia, the company has been using the C2PA standard and SynthID signals for some time. Unlike the text checker, however, the verification tools for images and audio via the web and the Content Provenance API remain accessible to everyone.