AI Tools

Claude's watermark can tell us AI was involved. It cannot tell us who did the thinking

Brendan Tack Brendan Tack · · 8 min read
Claude's watermark can tell us AI was involved. It cannot tell us who did the thinking

Claude's watermark can tell us AI was involved. It cannot tell us who did the thinking

Anthropic's new watermark is not a visible badge stamped across a page. It is not a hidden username, a tracking code or a row of invisible characters waiting to expose you.

It is a statistical pattern in the words Claude chooses.

That distinction matters because the watermark can answer one narrow question: was Claude likely involved in producing this text? It cannot tell us whether Claude fixed two clumsy sentences, translated a human-written article, reorganised a rough draft or wrote the whole thing from a three-line prompt.

My view is that watermarking is broadly a good move. The internet needs better provenance. But if employers, schools, publishers and clients use it as a binary test of who "really wrote" something, it will create more confusion than clarity.

Provenance is useful. A verdict on authorship is something else entirely.

What Claude's watermark actually does

Anthropic announced the system on 14 August 2026. Future Claude models will use a version of Google DeepMind's SynthID-Text technique.

Language models generate text one piece at a time. Often there is only one sensible choice. A date must be accurate. A calculation has a correct answer. A line of code may break if one term changes.

But prose contains thousands of low-stakes choices. "Cold and grey" might work just as well as "cold and overcast". When several choices are equally suitable, Claude's watermark changes the source of randomness used to select between them. Across a long enough passage, those choices form a pattern that can be checked with Anthropic's private key.

Nothing visible is added. There are no hidden characters. Anthropic says the watermark carries no information about the user, their organisation or the chat that produced it.

The signal is strongest in long, freely written prose because Claude makes more choices. It is weaker in short, factual or exact passages. Light proofreading may leave no detectable signal because almost all the words still belong to the human author. A translation is different: Claude chooses every word in the translated version, so that text can carry the watermark.

Light editing may preserve the pattern. A complete human rewrite can remove it. Anthropic also plans to attach C2PA provenance metadata to supported image files such as PNG, JPG and SVG. That is a separate system stored in the file metadata, not the text watermark.

A detection API is coming, but Anthropic has already stated the system's most important limitation: it can indicate that Claude was likely involved, but it cannot distinguish "Claude wrote this" from "Claude heavily edited this".

That sentence should be printed above every detector result.

Why Anthropic is doing this now

The immediate reason is regulation. Anthropic signed the EU Code of Practice on Transparency of AI-Generated Content and says it is implementing watermarking to comply with the EU AI Act. It is launching the system globally because it cannot yet limit the feature reliably by region.

So this is bigger than one Claude update. Rules created in Europe are changing how AI-generated text may travel everywhere.

That is probably where the wider market is heading. Text, images, audio and video will increasingly carry machine-readable signals about where they came from and which tools touched them. In a world of cheap synthetic reviews, fake experts, cloned voices and mass-produced articles, that context can help.

But provenance is not truth. A human can write rubbish. An AI can produce an accurate paragraph. A watermarked article can be carefully researched, while an unmarked article can be entirely machine-generated by another model.

The signal tells us something about the production process. It does not tell us whether the result is original, honest, accurate or worth reading.

Is the watermark a good move?

Yes, with guardrails.

Anthropic's approach is better than adding a visible "Made by AI" label to every document Claude touches. It does not interrupt the reading experience, it does not identify the user and it acknowledges that detection is probabilistic.

It is also more defensible than many AI-writing detectors, which try to guess authorship from familiar wording and sentence patterns. Those systems have a record of turning style into suspicion. A provider-controlled statistical signal is narrower and more grounded.

The danger comes when a narrow signal is used to make a broad accusation.

A school should not fail a student because a detector found evidence consistent with Claude. An employer should not accuse someone of outsourcing their judgement. A client should not assume a commissioned article was generated from scratch. The signal cannot tell them that.

Detection should start a conversation:

It should never be the only evidence used for punishment, rejection or disciplinary action. Anthropic should also publish reliability results across languages and types of writing, support independent auditing, report uncertainty clearly and offer a meaningful appeal route wherever its detector informs a consequential decision.

Without those safeguards, honest users may carry the watermark while determined deceivers simply rewrite the output until the signal disappears.

What warrants a watermark or disclosure?

There are two separate questions here.

Anthropic can reasonably watermark model-generated text by default. A technical provenance system works best when it is applied consistently.

That does not mean every detected use deserves the public label "AI-written".

Disclosure is warranted when knowing about the AI contribution would reasonably change how the audience judges the work. That includes academic assessment, journalism, public-interest commentary, legal or professional advice, job applications, commissioned human writing, personal testimony and any situation where the writer's individual expertise or voice is part of what the audience has been promised.

It is also warranted when AI generated a substantial share of the published wording, created the analysis, introduced factual claims, produced a translation or supplied examples that remain in the final piece.

Spellcheck, punctuation fixes and light grammar correction usually do not need a warning label. We do not demand a disclosure every time software catches a typo. Local rules may still require it, particularly in education or regulated work, but the contribution is mechanical rather than authorial.

A useful test is:

Would a reasonable reader evaluate the person's competence, originality, evidence or accountability differently if they knew what the AI did?

If the answer is yes, disclose the contribution in plain English.

If AI rearranges your words, who did the work?

"AI did the work" is too vague to be useful. Writing contains several kinds of work: developing the idea, finding evidence, constructing the argument, choosing the words, editing the structure and accepting responsibility.

I would separate AI involvement into four levels.

1. Mechanical assistance

You write the ideas, structure and sentences. Claude fixes spelling, punctuation or a handful of grammatical errors.

You did the writing. The tool performed copy-editing. Calling the result "AI-written" would be misleading.

2. Structural editing

You supply the argument and the existing prose. Claude moves paragraphs, cuts repetition, suggests a clearer order or tightens awkward sections.

The human remains the principal author, but Claude has done editorial work. If the restructuring materially changes how the argument lands, "edited with AI assistance" is a fair description.

Rearranging three paragraphs is not the same as rebuilding the narrative and rewriting every transition. Degree matters.

3. Collaborative drafting

You provide the thesis, research, evidence and judgement. Claude proposes sections or sentences that you select, verify and substantially revise.

Both have contributed to the production process. The human still carries responsibility, but pretending the tool only checked spelling would be dishonest. A sensible disclosure might say:

Claude helped restructure the argument and draft selected passages. The author verified the evidence and edited the final text.

4. Predominantly AI-generated

You provide a topic or short prompt and Claude writes most of the article. You may choose a version, request revisions and approve the result.

In that case, yes, AI did the writing work. The person commissioned, directed and reviewed it, which are real contributions, but they are not the same as personally writing the article.

The honest description is not complicated: "Draft generated with Claude and reviewed by a human editor."

Does AI own the idea?

No.

Ideas are generally not protected by copyright in the first place. Copyright protects qualifying human expression: the particular wording, structure, selection or creative treatment of an idea. The details vary by jurisdiction, and the UK's treatment of computer-generated works remains contested and under review, but an AI system does not become a legal owner simply because it generated text.

Current US Copyright Office guidance says purely AI-generated material is not protected by copyright, while a person's original writing, creative selection, arrangement and modifications may be. A prompt can contain a valuable idea, but that does not necessarily make every sentence the model produces an act of human authorship.

This leaves three questions that people often muddle together:

  1. Who came up with the idea?
  2. Who produced the published expression?
  3. Which human contributions receive legal protection?

If I develop the thesis and Claude writes the first draft, the idea may be mine in the ordinary sense. Claude produced much of the wording. Claude owns neither. My claim to authorship should match what I actually contributed.

A watermark does not change any of this. Anthropic says it does not determine ownership, authorship or legal responsibility. It is simply evidence that Claude may have processed the content.

My rule for the future of content

Every meaningful AI disclosure should answer three things:

"AI-assisted" on its own is almost useless. It can mean spellcheck or a complete ghostwritten article.

"Claude reorganised my draft and suggested two paragraphs; I checked the sources and rewrote the final version" tells the reader something real.

Claude's watermark can make provenance easier to test. That is progress. But the future of trusted content will not be built by sorting everything into "human" and "AI" piles. It will be built by describing the contribution honestly and keeping human responsibility visible.

Used that way, Anthropic's move is sensible. Used as a machine-generated scarlet letter, it will punish transparency while doing little to stop deception.

Sources


Drafting disclosure: The original questions and editorial direction came from Brendan Tack. AI was used for research synthesis and drafting. Human review is required before publication.

Want to talk about your business?

Book a free Reverse Demo — we'll show you what your operation could look like with the right automations in place.

Book a Reverse Demo