On October 5, OpenAI published details of textGrain, an entropy-calibrated watermarking technique designed to embed subtle, statistical signals into language model token selections. Starting immediately, API developers globally can opt into watermarking on select models, though the feature remains switched off by default. Over the coming weeks, OpenAI plans to watermark eligible ChatGPT and Codex text outputs in the European Union across all plans. The company framed the deployment as an initiative responding to transparency expectations under the EU AI Act, according to its announcement.

Access to detection tools is tightly restricted at launch. While OpenAI maintains separate, public verification tools for generated images and audio, textGrain detection is not publicly accessible. Instead, OpenAI is opening an application process to grant detector access on a case-by-case basis to vetted researchers and specialized organizations. OpenAI stated that it plans to open-source the underlying technology in the future, with no release date given in the announcement.

Mechanism and Vendor Evaluations

The textGrain system functions during text generation by biasing the model's token sampling based on a keyed statistical distribution. Detection relies on evaluating whether a text sample contains this keyed bias, without identifying the user, account, or original prompt. This describes what the watermark identifies, rather than a guarantee about service logging.

OpenAI's announcement reports evaluations using a target false-positive rate of 1 percent. In company tests using psychology-related passages, OpenAI reported an approximate 80 percent detection rate for 200-token samples and roughly 95 percent for 400-token samples. Detection rates were materially lower for mathematics content, where lower entropy and constrained word choices restrict watermarking opportunities.

OpenAI also evaluated the watermark's resilience against synonym replacement using 400-token English ELI5 samples:

Modification LevelReported Detection Rate
Unmodified text92%
10% synonym replacement66%
25% synonym replacement17%

These figures reflect controlled vendor experiments rather than independent, third-party audits or real-world reliability guarantees. OpenAI noted that idealized mathematical calibrations do not assure a uniform error rate across varied deployment keys and arbitrary text domains. Addressing output quality, OpenAI stated that the watermark caused no meaningful degradation across its Astra benchmark suite, according to its own evaluation.

Technical Boundaries and Detection Limits

OpenAI emphasized several material constraints governing textGrain:

  • The watermark does not identify the user, account, or prompt.
  • Detection does not determine factual truthfulness, human editorial contribution, legal ownership, or liability.
  • The absence of a detectable watermark does not prove human authorship, as unwatermarked models, low-entropy text, and heavily edited outputs can fail to register a signal.
  • The rollout covers eligible EU ChatGPT and Codex outputs, with an optional API feature on select models; it is not a universal watermark for AI-generated text.
Figure2 from the textGrain paper showing token-block recovery and a statistical watermark test
The paper’s Figure2 sketches keyed detection against a null distribution. Detection thresholds do not establish who authored or edited a passage. Image: OpenAI and coauthors · textGrain technical report