AI Watermark Check

What the EU AI Act requires of AI-generated content

Article 50(2) is the reason Claude's output carries a watermark and the reason every major provider is now talking about marking synthetic content. It is also more narrowly drafted than most summaries of it suggest.

The obligation itself

Article 50 of the AI Act covers transparency obligations for certain AI systems. Paragraph 2 is the one that matters for text: providers of AI systems that generate synthetic content must ensure that output is marked in a machine-readable format and is detectable as artificially generated or manipulated.

Four words in that sentence carry the weight.

Providers. The obligation sits with whoever puts the system on the market, not with the person typing the prompt. If you use a model to draft an email, you are not the party that has to mark it.

Machine-readable. Not a visible disclaimer. A signal that software can check, which pushes implementations toward watermarking and cryptographic provenance rather than a line of text at the bottom of an answer.

Detectable. Marking is only useful if something can read the mark. That word is why detection tooling arrives alongside the watermark rather than years afterwards.

Effective, interoperable, robust and reliable, the qualifiers the Article attaches to technical solutions, tempered by what is technically feasible. That last clause is doing real work: it acknowledges that some modalities are far harder to mark than others, and that a solution good enough for images may not exist for a two-sentence text reply.

The code of practice

Regulation states an outcome. A code of practice is how providers demonstrate that they have reached it. The transparency code attached to Article 50 is the instrument through which the major providers coordinated their implementations, and signing it is a public commitment rather than a legal reclassification.

This is why the vendor landscape looks the way it does. Google already had SynthID deployed across text, images, audio and video before the deadline. Anthropic confirmed on 11 August 2026 that Claude models released after the 2 August cutoff watermark their text and mark generated files with C2PA, with older models to follow. OpenAI signed the code and applies C2PA to images, without a deployed text watermark at the time of writing. Same obligation, different positions on the implementation curve.

What the Act does not require

Three misreadings are common enough to be worth stating explicitly.

It does not require you to disclose that you used AI to write something. There is a separate transparency obligation for deepfakes and for AI-generated text published to inform the public on matters of public interest, which lands on deployers rather than providers. General personal or business use is not covered.

It does not make watermark removal illegal. The Article obliges providers to mark. It does not create an offence for a user who edits marked text. Whether a given removal breaches some other law, contract, fraud, academic regulations, is a separate question with a separate answer.

It does not guarantee detection works. Marking and detecting are different capabilities. A watermark that survives is only useful to someone who can check it, and for statistical schemes that means holding the provider's key. The Article does not require providers to give that key away, and none of them do.

Why this produced statistical watermarks specifically

"Machine-readable" and "robust" together rule out most of the alternatives for text.

A visible disclaimer is not machine-readable in any useful sense. Metadata is machine-readable but not robust: it is stripped the moment text is copied out of the file that carries it, and text is copied constantly. Hidden Unicode characters are both machine-readable and trivially removable, a single regular expression defeats them, which is exactly why no serious provider uses them as a mark.

That leaves the class of schemes that embed the signal in the generated content itself. For text, that means biasing token selection during sampling: the mark is carried by which words were chosen, it survives copying because the words survive copying, and it degrades only when the words change. Every deployed text watermark works this way, and the difference between statistical and character marks is the single most useful thing to understand about the whole field.

What it means in practice

Primary sources: Regulation (EU) 2024/1689, Article 50(2); the transparency code of practice associated with it; and provider announcements from Anthropic and Google DeepMind. We revisit this page when any of those change.