What the watermark can prove
OpenAI’s textGrain inserts an invisible statistical signal through the model’s word choices. The company says it plans to open-source the technology and that its internal tests matched or outperformed other approaches, including Google’s SynthID.
Detection depends on both length and subject. With the detector set to a 1% false-positive rate, OpenAI says it identifies watermarks in about 95% of 400-token psychology excerpts. At 200 tokens, that falls to about 80%. Mathematical writing is harder to detect because the model has less freedom in its choice of words.
OpenAI suggests longer passages may strengthen the signal, but provides no data to support that claim. Its reported results also show how quickly editing can weaken detection: replacing 10% of the words in a 400-token excerpt with synonyms drops accuracy from about 92% to 66%; replacing a quarter brings it down to 17%.
Detection depends heavily on text length and subject matter. At 400 tokens, roughly 300 words, the detector identifies about 94 percent of watermarks in psychology passages but only about 60 percent in math passages. | Image: OpenAI
Source: the-decoder.com
That fragility is more than a technical caveat. My read is that a watermark easy to remove could push people who want to conceal their use of ChatGPT toward open-weight models. A mark that cannot reliably survive routine editing may be useful for some checks, but it is a weak basis for treating an unmarked passage as human-written.
Even limited editing sharply reduces detection rates. Replacing 25 percent of the words pushes detection below 20 percent, even in 400-token passages. | Image: OpenAI
Source: the-decoder.com
OpenAI says textGrain does not reduce answer quality. In tests of its advanced Astra model across eight benchmarks, including GPQA Diamond, BrowseComp and DeepSWE, the company found no significant difference between marked and unmarked outputs. Those results do not address writing quality, however; critics have raised a similar question about Claude’s watermark.
A mark is not a verdict
OpenAI’s own cautions define the limits of what the signal means. A detected watermark does not establish authorship or responsibility, reveal the user’s identity, show how much human writing or editing went into a passage, or confirm that its claims are accurate. And a missing mark does not prove a person wrote the text: it may be too short, edited, translated or generated by a model the detector does not support.
The detector will initially be available only to selected researchers and relevant organizations, which can apply through a dedicated form. OpenAI says access will be granted individually under the EU Code of Practice. The tool will report only whether it found an OpenAI watermark; it will not identify users or expose their prompts or conversations.
That restraint reflects a real risk: the detector can mistake unmarked text for marked text, or miss a genuine watermark. OpenAI says it will widen access when it considers the results responsible to interpret, but has not said when that will be. Existing image and audio verification tools remain open to everyone at openai.com/verify and through the Content Provenance API.
What I’d want to know is how often the detector gets those calls wrong outside the company’s tests. Until independent users can examine its performance, OpenAI is asking institutions to treat a signal as evidence while keeping the means to check that evidence out of broad reach.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X