OpenAI to watermark ChatGPT and Codex text in Europe, here's what it means for users
ChatGPT-maker OpenAI has announced that it will add invisible watermarks to text generated by ChatGPT and Codex for users in the European Union, as the
ChatGPT-maker OpenAI has announced that it will add invisible watermarks to text generated by ChatGPT and Codex for users in the European Union, as the company moves to comply with new requirements under the EU AI Act. The company also said that the rollout will being over the coming weeks and will apply to eligible ChatGPT and Codex text outputs across all plans in the EU. The move is designed to make AI-generated text identifiable in a machine-readable way while promoting greater transparency around AI-created content.What is changingOpenAI said that over the coming weeks it will add the watermark to eligible text output from ChatGPT and Codex. The change covers all plans in the EU and nowhere else. The company said it is not making text watermarking a global default at launch and described the regional approach as a way to learn from real-world use and feedback.The company is also making watermarking available to developers. Starting immediately, API customers worldwide can opt in to watermarked text for select models. It stays off by default in the API. OpenAI said it is working with cloud partners to offer watermarking for its models through their services in the coming weeks.How the watermark worksOpenAI's technology, called textGrain, adds a statistical signal to the model's word choices. Readers cannot see it. A detector looks for that pattern to judge whether a passage carries an OpenAI watermark. OpenAI has published a technical report on the method and said it plans to release the technology as open source.The company said textGrain matched or beat the other approaches it tested, including SynthID for text. It added that strong results under ideal conditions do not guarantee reliable detection in everyday use.Limits of detectionOpenAI's own evaluations show where the technology struggles:* Length matters. At a 1% false positive rate, the detector found watermarks in about 80% of 200-token passages and about 95% of 400-token passages in content such as psychology. Detection was substantially lower for content such as mathematics, where there is less flexibility in word choice.* Editing weakens the signal. In tests on 400-token passages, swapping 10% of words for synonyms cut detection from about 92% to 66%. Swapping 25% cut it to 17%.What a watermark does and does not showOpenAI listed several things a detection result cannot tell anyone:* It does not measure how much a human contributed to the text.* It does not establish who owns the text or who is responsible for it.* It does not identify the user. It does not link the text to a person, account, prompt or conversation.* It does not verify whether the text is accurate.* The absence of a watermark does not prove a human wrote the text. Short, edited or translated text may evade detection. Text may also come from an unsupported model, predate watermarking or come from another company's tools.What it means for usersFor people using ChatGPT or Codex in the EU, the visible experience should not change, since the marker cannot be seen. OpenAI said it found no meaningful quality difference in tests of its latest frontier model, Astra, with and without watermarking. On eight benchmarks, the scores were close, with the watermarked version slightly ahead on some and slightly behind on others.Users should understand that a watermark is not a verdict on authorship. Because it can be weakened by editing and does not identify users, it is not a reliable way to prove who wrote a passage. That matters for anyone who might otherwise treat a detection result as proof, such as teachers, employers and publishers.OpenAI said the detector does not reveal users' prompts or conversations.Who gets the detectorOpenAI is not releasing the text detector to the public. Researchers and expert organizations can apply for access, which will be granted case by case under the EU's Code of Practice on AI-generated content. The company said the risk of missed watermarks and false positives is the reason it is holding back a public release.Its tools for images and audio remain publicly available. These include a web verification tool at openai.com/verify and a Content Provenance API.You use AI every day. Now get your AI Quotient. Take the AIQ test.
Topics in this story
Gathered from external sources. Rights to this text belong to whoever originally published it.