Blog Article
Can You Prove Which Company Content Is AI-Written? Why AI Text Watermarks Are Not Enough
Could your company say which of its content was written with AI? OpenAI is watermarking ChatGPT text in the EU, but the mark is fragile. Why records matter and what to do.
Imagine a customer, an auditor or a regulator asks one simple question: which of your content was written with AI? If your honest answer is “we don’t know”, you have a problem that is easy to prevent today and hard to fix later. This post explains why, using news from this week, in plain terms for business leaders who have not heard of text provenance before.
The Risk of Not Knowing
AI tools now write emails, proposals, reports, web pages and support answers across most companies, often without anyone tracking it. That is fine until someone needs an answer. This part is our view, and your legal team should assess your own obligations:
- You may not be able to answer a basic question. Which documents were drafted or edited with AI, by which tool, and who checked them? If nobody recorded it when the work was done, there is nothing to look up later.
- You cannot reconstruct it afterwards. As the next section shows, detection tools are weak, so “we will check later” is not a plan.
- Rules are arriving. The EU AI Act now has transparency rules for AI-generated content, and AI providers are starting to mark their output because of them. Customers and partners in the EU are likely to ask more questions about how you use AI.
Why Detection Will Not Save You
The obvious fix is to scan your content for AI and see what turns up. OpenAI’s own numbers show why that does not work well. These are the vendor’s tests, not independent ones:
- Shorter or more constrained text is harder. At a target false positive rate of 1%, OpenAI’s detector found the watermark in about 80% of 200-token passages and about 95% of 400-token passages, for content such as psychology. OpenAI says detection rates were substantially lower for content such as mathematics, where there is less flexibility in word choice.
- Editing can weaken the watermark. In OpenAI’s evaluation of 400-token English passages, replacing 10% of words with synonyms reduced detection from about 92% to 66%. Replacing 25% of words reduced it to 17%.
- Rewriting and translation make it worse. OpenAI’s help center says substantial rewriting, paraphrasing or translation can make detection less reliable, and that even unchanged wording does not guarantee a detectable mark.
- No mark proves nothing. OpenAI says a missing watermark does not prove a person wrote the text, and a detected one does not show how much human judgment went into it.
- The detector is not public. OpenAI is opening applications for access, which will initially be limited to approved researchers and expert organizations. OpenAI says strong performance under ideal conditions “does not guarantee reliable detection in everyday use”.
The practical point: a record you did not keep when the work was done cannot be recovered by a tool afterwards.
What Did OpenAI Announce?
Facts in this section come from OpenAI’s announcement, which OpenAI published on October 5, and its help center article, plus reporting by TechCrunch and Search Engine Journal.
OpenAI says the EU AI Act requires generative AI providers to make generated text identifiable in a machine-readable way, and calls text watermarking and detection “early technologies with significant limitations”. Its phased approach has three parts: API customers worldwide can opt in to text watermarking for select models starting October 5 (off by default); over the coming weeks, an invisible watermark will be added to eligible ChatGPT and Codex text output in the EU; and OpenAI is opening applications for access to its text detector. The method, which OpenAI calls textGrain, adds an invisible statistical signal to the model’s word choices that a detector can look for. Because the signal lives in the words, it stays with the text when it is copied and pasted, according to TechCrunch. OpenAI plans to release the technology as open source, and its technical report will be updated in the coming weeks.
OpenAI says text watermarking is not a global default at launch. OpenAI’s help center says API customers choose which models get the watermark and that coverage will extend to legacy models over the coming weeks. Search Engine Journal notes that the announcement does not say whether EU ChatGPT users can turn the watermark off. This applies to text only: OpenAI says its verification tools for images and audio, including openai.com/verify and the Content Provenance API, stay publicly available. OpenAI also says it is working with cloud partners to embed watermarks in eligible outputs, and that availability may vary by partner.
TechCrunch reports that the EU AI Act’s transparency rules took effect on August 2 and require AI providers to mark AI-generated content so other systems can identify it. Search Engine Journal adds that systems placed on the market before that date have until December 2 to meet the marking and detection obligation, and that Anthropic marks Claude text worldwide. Because the ChatGPT watermark is EU only, the same prompt can produce marked text in Berlin and unmarked text in Toronto.
What Should Your Business Change?
This section is our advice, not OpenAI’s. The aim is to prove what your own systems did, instead of trying to detect it afterwards:
- Keep provenance records. For important documents, log which AI tool and model produced or edited the text, when, and who reviewed it.
- Update content workflows. Add a review and sign-off step for AI-assisted text, and record it where your team already works.
- Update your AI use policy. State which tools staff may use, for what content, and who is accountable for the output.
- Ask your AI vendors. How do they mark output? Is it on by default? Does it differ by region? What can you do with the mark, and what can you not?
- Plan for EU and non-EU scope. Staff in the EU may produce marked text while colleagues elsewhere do not, even on the same plan. Do not assume the mark is present or absent.
- Do not use detection as proof. Avoid clauses or rules that treat a watermark result as proof of authorship.
- Keep visible disclosures. OpenAI says watermarks do not replace visible labels or other notices that may be required, and that you should have your legal team assess your own obligations.
How Incresco Helps
Incresco does AI digital transformation. The gap between a new rule and what the technology can prove is where we work with companies. Our AI transformation services cover the changes above:
- AI strategy and consulting: a readiness assessment of where AI text is created in your company, an AI use policy, and a roadmap that accounts for rules such as the EU AI Act.
- Custom AI development: AI agents, copilots and workflow automation with logging built in, so provenance records are created as the work happens.
- AI integration: connecting those records to your existing systems through workflow automation, API integration and legacy system modernisation.
If you want to be able to answer “which of our content is AI-written?” with a record instead of a guess, talk to our team.