As Anthropic confirms that future Claude models would embed statistical watermarks into their output, it reignited a debate that has been simmering since Google DeepMind published its SynthID-Text scheme in Nature in 2024. Does this mean that invisibly tagging AI-written text will facilitate regulation, education, and content creation?
Anthropic was one of roughly 190 signatories to the EU’s Code of Practice on Transparency of AI-Generated Content, built to help providers meet Article 50 of the EU AI Act, whose text-marking obligations became applicable on August 2, 2026. Google, OpenAI, and other labs are rolling out equivalent systems in parallel.
“These “AI-gen watermarks” are going to become excellent markers of obedience to a governance regime – while remaining vastly weaker as proof of where a given piece of AI-generated text ultimately came from.” – Ben Goertzel
How does AI text watermarking work?
At almost every point in a sentence, several next-words work equally well — “cold and overcast” versus “cold and grey.” That choice is normally settled by a random draw. A watermark biases the draw using a secret key, so the resulting sequence carries a statistical signature invisible to readers but recoverable by anyone holding the key. Anthropic frames the resulting detector output narrowly: not proof text is AI-written, and not proof it’s human-written, but an estimate of the likelihood Claude was involved.
The watermark carries no user or organisational identity. It performs poorly on short passages, where there isn’t enough accumulated bias to detect. And it goes nearly silent on constrained text — factual sentences, code, proofreading — where there’s often only one correct word and no “dice” left to load.

The case for it
It’s essentially free. Anthropic’s testing, and DeepMind’s earlier A/B tests across Gemini traffic, found no significant quality difference between watermarked and unwatermarked output. No extra tokens, no slowdown, no added cost.
It catches the laziest misuse. Bulk bot spam, unedited submissions, unmodified astroturfing — a watermark flags what would otherwise pass silently, giving platforms one honest signal among many.
It’s legally required for EU market access. Article 50(2) obliges providers to mark synthetic outputs in machine-readable format; providers operating in the EU don’t have much choice.
It reflects real international convergence on the idea that some form of machine-readable marking is now table stakes — discussed further below.
The case against
The sharper critique, laid out by Goertzel, isn’t that watermarking is technically broken. It’s that it works — just not at the thing people assume it’s for.
Removal is trivial where it matters most. A watermark applied at a closed API’s sampling layer is barely a speed bump for anyone running an open-weight model locally: swap in a standard sampler and marking never happens. DeepMind’s own documentation notes the watermark lives in the sampling procedure, not the trained weights — so whoever controls inference controls whether marking occurs at all. Even for closed models, heavy paraphrasing, translation, or a rewrite pass through a second model can degrade or erase the signal, no cryptography knowledge required.
Absence proves nothing; presence proves less than it seems. Unmarked text may simply come from a different, non-watermarking model. Marked text may be a human draft, a model merely reorganised. Neither maps cleanly onto the question people actually want answered: who is responsible for this, and how much reflects independent judgment?
There’s a genuine technical wrinkle around distillation. Research on “watermark radioactivity” (Sander et al., NeurIPS 2024) found that a smaller “student” model imitating a watermarked “teacher’s” outputs can inherit faint traces of the mark, because exact-imitation training copies stylistic quirks along with substance. This is the strongest card watermark designers hold. But follow-up research (Pan et al., 2025) has already demonstrated attacks that strip inherited watermarks while retaining distilled capability — and the underlying asymmetry favors removal: designers need every capability-preserving transformation to preserve the mark, while an attacker needs just one that doesn’t.
It risks building a compliance perimeter, not a truth detector. If marked, compliant models become detectable while unmarked or scrubbed models circulate freely, the practical effect is sorting text into “inside the governance boundary” versus “outside it” — not “AI-written” versus “human-written.” Institutions then face pressure to treat unmarked text as presumptively suspect, even when it may be entirely human-written.
It creates perverse incentives in education and publishing. A detectable watermark rewards whoever happens to use — or avoid — a compliant tool, rather than rewarding originality or rigor. Once students learn one model family is detectable and another isn’t, the system starts selecting for evasion skill over writing skill.
To be fair to watermarking’s defenders, none of them claim it’s meant to stop a determined, technically sophisticated actor. It’s framed as raising the cost of casual misuse, not sealing anything shut — and Anthropic’s own FAQ concedes a full rewrite removes it, adding that at that point “it’s arguable whether the text can any longer be described as AI-generated” at all.
Not the same as cryptographic provenance
It’s easy to conflate statistical watermarking with cryptographic content credentials — the C2PA standard Anthropic also applies to image and file outputs via signed metadata. A C2PA credential is a deliberate, checkable declaration of a signed chain of custody. A statistical text watermark is closer to an accent a speaker can’t quite suppress: suggestive and probabilistic, lost the moment the “speaker” changes through paraphrase or translation. Strip a file’s metadata and the credential vanishes while the words stay intact; rewrite watermarked prose and the words change while nothing was technically “removed.” Both offer real but bounded evidence — neither proves ultimate origin or human authorship.
The global regulatory picture
Governments are converging on the goal of marking AI content while diverging sharply on method and enforceability.
European Union: Article 50 of the AI Act requires machine-readable marking and disclosure of certain AI text. The Commission ran a seven-month, multi-stakeholder drafting process from November 2025, publishing a first Code of Practice draft in December 2025 and a second in March 2026, ahead of the obligations taking effect August 2, 2026.
China: Beijing has gone furthest. The Cyberspace Administration of China issued binding labeling measures alongside a mandatory technical standard (GB 45438-2025) in March 2025, effective September 1, 2025 — unlike the EU’s voluntary code. Rules require visible labels and embedded metadata across text, image, audio, and video, with a three-tier classification system, enforced on platforms including WeChat, Douyin, and Weibo. This builds on earlier 2023 rules targeting “deep synthesis” technology.
United States: Regulation is fragmented and state-driven. California’s AI Transparency Act (effective January 1, 2026) requires large providers to offer a free detection tool and embed provenance data, building on an earlier bill that drew support from OpenAI, Adobe, and Microsoft after amendments addressed industry objections. Colorado, Utah, and Illinois have their own transparency laws. Federally, a 2023 executive order touching AI labeling was later rescinded, leaving no unified national mandate.
Elsewhere: The Coalition for Content Provenance and Authenticity (C2PA) — backed by Adobe, Microsoft, camera makers, and now Anthropic — offers a voluntary, cross-border provenance standard that regulators increasingly treat as an interoperable baseline rather than a competing scheme.
What to actually do with this
Readers: Don’t treat a watermark’s presence or absence as a verdict on truth. It signals a compliant model was probably involved somewhere — nothing more. Evaluate claims on evidence, as you always should have.
Writers and researchers: Where disclosure matters, describe your tool use directly (“I used an LLM to reorganize a draft and check citations”) rather than letting a fragile detector speak for you.
Educators: Build assessment around demonstrated understanding — oral defense, staged drafts, in-class work — rather than detector scores that collapse the moment a student switches tools.
Enterprises: An output watermark says nothing about how data was handled upstream. Confidentiality and compliance need access controls and audit logs, not inferred provenance.
Everyone: Expect continued regulatory divergence, not convergence, in the near term — watch the EU’s enforcement posture, China’s platform-level compliance, and the growing U.S. state patchwork.
The bottom line
Statistical text watermarking is real, cheap, and useful within a narrow band: catching bulk unmodified output and supporting compliance with rules like the EU AI Act. It was never built to reliably answer whether a given passage is “AI” or “human,” nor to substitute for judging writing by its evidence, logic, and the responsibility its author takes for it. Treating a narrow, provisional signal as a verdict on authenticity is where the real risk sits — for readers and institutions alike.

Pallavi Singal is the Vice President of Content at ztudium, where she leads innovative content strategies and oversees the development of high-impact editorial initiatives. With a strong background in digital media and a passion for storytelling, Pallavi plays a pivotal role in scaling the content operations for ztudium's platforms, including Businessabc, Citiesabc, and IntelligentHQ, Wisdomia.ai, MStores, and many others. Her expertise spans content creation, SEO, and digital marketing, driving engagement and growth across multiple channels. Pallavi's work is characterised by a keen insight into emerging trends in business, technologies like AI, blockchain, metaverse and others, and society, making her a trusted voice in the industry.

