All articles
Articles

The mark that only points one way

Since 2 August 2026 the EU requires AI-generated content to be machine-readable as such, and Anthropic has published how Claude will comply. The marks work in one direction only: they can suggest a machine was involved, and they can never establish that one was not.

An essay card showing a detector that returns a positive signal in one direction and nothing at all in the other

Ten days ago, on 2 August 2026, Article 50 of the EU AI Act started to apply. Anyone providing a system that generates synthetic text, image, audio or video must now make that output machine-readable as machine-made. Anthropic has signed the accompanying Code of Practice on Transparency of AI-Generated Content and published what it intends to do about it: Claude models launched in the EU from that date carry marking at launch, embedded watermarks in generated text, C2PA provenance metadata signed into generated files, and detection tooling to follow. Models that shipped before the deadline get a transition period, and the interoperable detection tooling that makes any of this checkable by third parties is not due until early 2027, so what exists today is the commitment rather than the finished machinery.

It is easy to read this as the arrival of a long-awaited line in the sand, one side machine, the other side human, at last legible. That reading does not survive contact with how the marks work. They are a one-way test. A mark that is found is weak evidence of involvement. A mark that is not found is no evidence of anything at all, and no amount of engineering will change that, because the asymmetry is structural rather than a defect of the current generation.

Two techniques, two different promises

The text watermark and the file signature are usually mentioned in the same breath and they are not the same kind of object.

A text watermark is a bias in how the model picks its next token. At each step a language model holds a probability distribution over the vocabulary and samples from it. A watermarking scheme uses a keyed pseudorandom function over the preceding few tokens to nudge that sample, so that across a long passage the choices land in a pattern that a party holding the key can recognise, while remaining statistically ordinary to everyone else. Google’s SynthID-Text, the first scheme deployed at scale and the only one with a published implementation and a Nature paper behind it, does this with a tournament between candidate tokens. Nothing is appended to the text. The mark is the word choices, which is why it survives copy and paste: there is no container to strip.

C2PA provenance metadata is the opposite construction. It is a manifest attached to a file and signed with a certificate, recording what produced the file and what was done to it. It does not touch the pixels. It is a cryptographic claim sitting alongside the content, which makes it verifiable and precise, and also trivially detachable.

Two techniques, two failure modes worth keeping apart: one degrades, the other disappears.

What the attacks actually show

Watermarking is not a young field and it has been probed properly. The most useful evaluation I know of is the ETH Zurich SRI Lab’s adversarial study of SynthID-Text, because it separates the two threats that get conflated.

Spoofing, forging the mark so that text you wrote is attributed to the model, turns out to be hard. Their strongest stealing-based attacker landed successful forgeries somewhere between 4 and 15 percent of attempts depending on configuration, and the forgeries left statistical clues that a second detector could flag. That is a real result and it matters: it means a mark that is present is not cheaply faked.

Scrubbing, removing the mark while keeping the meaning, is the opposite story. Running the watermarked text through an off-the-shelf paraphraser removed the mark in over 90 percent of attempts, without the attacker needing to know anything about the scheme. With a preliminary black-box stealing step the rate approaches 100 percent. A separate 2025 assessment reports the same direction of travel: paraphrasing, back-translation and copy-paste editing all degrade detectability sharply.

One honest caveat, and it is not a small one. These numbers are measurements of SynthID-Text, not of Anthropic’s scheme, which has not been described in technical detail yet and may well differ. What generalises is not the percentage but the shape of the problem: a watermark lives in the choice of words, and paraphrasing is precisely the operation that replaces the choice of words while preserving the meaning. Any scheme in this family inherits that exposure. The asymmetry is also worth noting on its own terms, because it is the reverse of what intuition suggests: the mark is hard to counterfeit and easy to erase, which makes it useful against the careless and useless against anyone who takes one deliberate step.

On the file side the failure is blunter. C2PA manifests are stripped by transcoding, by re-saving, by format conversion, and by most platforms on upload. A screenshot ends the chain outright. Camera-side signing has arrived on real hardware, with Leica, Sony, Nikon and Canon bodies now able to sign at capture, and several platforms read and display credentials on ingest, but the chain only holds where every link cooperates, and one uncooperative link is enough.

The asymmetry is the whole story

Put the two halves together and read the detector honestly.

Mark found. The content probably passed through the model. Anthropic’s own documentation is careful about how little this establishes: people use Claude to proofread, translate, summarise and convert files, so a mark can sit on text whose ideas, structure and argument are entirely someone else’s. The mark records contact, not authorship.

No mark found. Nothing follows. The content may predate marking, or come from a model that does not mark, or from a local open-weights model with the watermarking simply switched off in the sampling loop, which costs nothing to do. It may have been paraphrased once. It may be too short to carry a reliable signal. Its metadata may have fallen off in a screenshot. Or a person may have written it.

That last list is the load-bearing part. All those branches are indistinguishable at the detector, so the absence of a mark carries no information whatsoever about human authorship. And note who is sorted by such a test: the careless and the compliant get marked, the deliberate do not. A test that only catches people who were not trying to evade it is worth having, in the way a lock is worth having, but it is not a proof of anything about the unmarked.

This is why “AI-detected” and “human-verified” are not two settings of the same dial. Verifying a machine is a matter of finding a signal that was deliberately put there. Verifying a human means proving the absence of every possible machine, an unbounded claim that no detector can discharge. If human authorship is ever certified, it will not be by the failure of an AI detector. It will be by positive provenance built at the moment of creation, the C2PA path, signed at the camera or the keyboard, which is a different engineering project with a different and much harder adoption problem.

Provenance is not quality

There is a further step people take without noticing, and it is worth resisting: from this was machine-made to this is worse.

Provenance and quality are different claims. Provenance is a fact about a causal chain and it is verifiable. Quality is a judgement about the work and it is not. The two come apart immediately in both directions. A researcher’s careful argument, translated by Claude for a second-language audience, carries a mark. A thousand words of human-written filler carries none. If a mark became a proxy for quality, the ranking it produced would be close to arbitrary, and worse, it would be gameable by exactly one cheap operation, the paraphrase, which improves nothing about the content and removes the signal completely.

What is actually emerging as a quality signal is not machine versus human. It is declared versus concealed. Disclosure is a claim the publisher makes and stakes their credibility on, and unlike a watermark it survives paraphrasing, translation and screenshots, because it is not hidden in the artefact. It is falsifiable in the way that matters: if someone declares no AI involvement and is caught, they have lied, and that is a reputational fact with consequences. Marking is worth building because it raises the floor and gives platforms something to act on at scale. It is not the thing that will distinguish careful work from careless work.

Which brings this back to something already at the bottom of this page. Every article on this site carries a note saying how it was written: imported by an AI assistant from a personal learning journal, then reviewed by me. That note is not machine-readable, it will not survive being copied elsewhere, and no regulation required it. It is a claim I make and can be held to. The regulation arriving now is trying to automate a weaker version of the same thing, and the automation is worth having. The claim is the part that carries the weight.

Further reading

← Back to all articles
How this article is written?

This article is imported daily by an AI assistant from a personal learning journal, then reviewed by me. Shared under CC BY 4.0.

© 2026 Akciali
Legal & Privacy