Overview
textGrain is a statistical text watermarking scheme OpenAI announced on October 5, 2026, laid out in a post titled Our approach to EU text provenance rules. It adds an invisible statistical signal through the model word choices: at every step of generation there is a cluster of near-equally likely candidate tokens, and a secret key combined with the preceding words quietly decides which of those candidates gets favored. No single word looks odd, but across a long enough passage the pattern becomes measurable. A detector holding the same key re-derives the preferred options at each position, counts how often the text followed them, and compares that rate to what chance would produce. A large gap means the text very likely came from the watermarked model.
Regulation drove this, not product strategy. Article 50 of the EU AI Act requires systems that generate synthetic text to mark their output in a machine-readable way, and reporting around the announcement puts the start of those transparency rules at August 2, 2026. That is why the rollout begins in the EU and why OpenAI framed the post as its approach to EU text provenance rules.
Key Features
- A statistical signal inside word choice: No pixels touched, no metadata written, no hidden Unicode characters. At each generation step a set of near-equal candidates exists, and a key decides which subset gets favored, spreading the bias across the entire response in a way a reader cannot see and a tool can count
- Detection re-derives the preference: A detector with the key recomputes the favored candidates at every position, measures how often the text followed them, and compares that rate against chance. The wider the gap, the stronger the evidence the text came from a watermarked model
- Length and entropy set the ceiling: Statistical confidence grows with token count, so a two-line email carries almost no signal while a long report carries plenty. Where the next token is nearly forced, such as boilerplate, code syntax, or quoted material, there is no room to bias and therefore no signal to embed
- A layered rollout: API customers worldwide can opt in for select models starting on announcement day, off by default. Over the coming weeks the invisible watermark is added to eligible ChatGPT and Codex text for EU users across all plans. It is not a global default at launch
- Detector access starts gated: Access is initially limited to approved researchers and expert organizations on a case-by-case basis, because the technology is weak on edited or fragmented content and carries both false negative and false positive risk
- Limits stated up front: A watermark can only indicate whether text was generated or processed by an OpenAI system. It cannot reveal user identity, measure how much a human contributed, determine ownership or copyright, or verify accuracy. OpenAI states outright that failing to detect a watermark does not prove a person wrote the text
Use Cases
- Product teams serving AI-generated text to EU users who need to meet the Article 50 machine-readable marking requirement
- Developers generating EU-facing text through the OpenAI API who must decide whether to enable watermarking for select models
- Schools, newsrooms, and hiring workflows considering watermark checks, who need the limits before relying on a verdict
- Research groups working on provenance and compliance who want detector access to publish independent results
Pros
- The mark travels with the content itself, with no external registry and no dependency on metadata surviving later
- No hidden characters and no visible tags, so the reading experience is untouched
- OpenAI reports no meaningful output-quality degradation
- The scope and the limits are spelled out concretely, including an explicit list of what it cannot do
- Off by default on the API, leaving the choice with the developer
- Connects to the C2PA metadata OpenAI already uses on images, extending provenance into text, the hardest modality
Pricing
textGrain carries no separate charge. API usage keeps its existing token-based billing, and turning the watermark on does not change pricing. ChatGPT and Codex output includes it for EU users at no extra cost on their existing plans. The detector is not distributed publicly; approved researchers and expert organizations apply case by case, and there is no paid tier for general users.
Summary
textGrain reads more accurately as a compliance instrument than as a detector. It answers the Article 50 requirement for machine-readable marking and offers a usable provenance signal along the way.
Its strengths and weaknesses both follow from the mechanism. Longer text carries more signal: OpenAI's published figures put detection at roughly 80% around 200 tokens and 95% at 400 tokens. Editing erodes it fast, with about 10% of words swapped dropping detection from 92% to 66%, and roughly a quarter swapped leaving about 17%. Translation replaces the word choices entirely and likely destroys the signal, while code and structured output have too little entropy to carry much in the first place. Those numbers come from OpenAI and have been relayed through secondary coverage; no third party with detector access has published an independent reproduction yet.
One point is easy to miss. The watermark indicates whether text was generated or processed by an OpenAI system, not whether it was generated by AI. Text a model rewrote or polished also carries the signal, so a positive result does not mean the whole passage was machine-written, and a negative result is even weaker evidence of human authorship. OpenAI says both things explicitly.
If you serve AI-generated text in the EU, or you are weighing whether to build watermark checks into your own workflow, this is worth understanding now. Just do not ask it to carry conclusions it was not built to carry.
Version History
- textGrain launches (2026-10-05): OpenAI published its approach to text provenance under EU AI Act Article 50 and introduced textGrain, a statistical text watermark that adds an invisible signal through the model word choices. API customers worldwide can opt in for select models from day one, off by default. Eligible ChatGPT and Codex text for EU users across all plans gets the invisible watermark over the coming weeks, and it is not a global default at launch. Detector access is limited to approved researchers and expert organizations. OpenAI also states it cannot identify users or settle ownership, and that a missing watermark does not prove human authorship