A rapid response to a new transparency measure
Anthropic’s move to embed machine-readable watermarks in text from supported Claude models has prompted an immediate counter-response from developers. Within hours of the policy becoming widely public, tools and code samples appeared claiming to remove the signal by paraphrasing text, changing word order, translating it and translating it back, or passing it through another language model.
The speed of that response is unsurprising. Text watermarking is not an unalterable label attached to a document; it is a statistical pattern spread through a model’s choices among plausible words and phrases. Changing enough of those choices can weaken or erase the pattern. Anthropic itself says that heavily edited, paraphrased, translated or mixed text may no longer retain a detectable mark.
But the more significant point is that these early “workarounds” have not yet been independently verified against Anthropic’s detector. Anthropic says it is still developing a detection API. Until such a tool is available, developers can demonstrate that a passage has been rewritten, but cannot conclusively show that it no longer triggers Claude’s watermark test.
What Claude’s watermark is designed to show
Anthropic describes its approach as a version of SynthID-Text, a method developed by Google DeepMind. Rather than inserting invisible characters or additional tokens, it alters the source of randomness used when a model selects between otherwise suitable next-word options. A detector with the relevant secret key can assess whether a sufficiently long text contains the expected statistical pattern.
That design gives the system some useful properties. Copying and pasting a Claude response should preserve the signal. Ordinary light edits may also leave enough of it intact to remain detectable. Anthropic says the mechanism does not identify a particular user, organisation or conversation, and that it does not add an explicit visible marker to the text.
It also establishes strict limits on what a positive result means. A watermark is evidence that Claude may have processed a passage; it is not a proof that Claude was the original author, that every sentence was model-written, or that a human did not substantially revise the material. Claude may have been used to translate, summarise or edit text whose central ideas came from elsewhere.
Equally, an absent mark cannot prove that a document was human-written. The model may predate the marking rollout, the sample may be too short, the text may have been substantially transformed, or a different AI system may have produced it. These are not marginal caveats: they define the appropriate use of the technology.
Why rewriting can be effective
The public removal methods described so far broadly rely on one fact: a text watermark is embedded in language choices, not stored independently of the language. Replacing words with synonyms, rearranging sentences, compressing material or generating a fresh paraphrase gradually replaces the original model’s decisions with new ones.
A full rewrite can therefore remove the evidence that a particular model made the original word choices. Anthropic acknowledges this directly, while arguing that a text whose every word has been replaced may no longer reasonably be described as the same AI-generated output.
That does not mean all editing defeats watermarking. A quick copy-edit, a few substituted adjectives or formatting changes are different from rewriting most sentences. The detector is designed to operate probabilistically across a body of text, so the effect of edits depends on the passage’s length, the degree of transformation and the amount of flexibility the original text contained.
This distinction matters for the claims circulating around watermark removers. A tool that produces readable revised prose is not automatically a tool that has demonstrated evasion. Conversely, an effective evasion method may require enough rewriting that it reduces factual precision, changes tone or takes significant processing time. The practical contest is not simply whether removal is possible, but how much useful text must change before the mark becomes unreliable.
Code is a special case
The controversy has been especially visible among users of Claude Code, but software code is not equivalent to free-form prose. Anthropic says its watermarking method depends on low-stakes alternatives where different choices would preserve meaning. Where an exact token is required for factual correctness or working software, the model has little room to encode a watermark.
As a result, Anthropic says code generally contains less watermarking than ordinary prose, although flexible elements such as comments can carry more of the signal. The European Commission’s Article 50 guidance also identifies source code among the outputs that fall outside the scope of the AI Act’s machine-readable marking obligation.
That does not necessarily remove all consequences for coding workflows. Claude Code can generate natural-language explanations, documentation, commit messages and comments alongside executable code. Anthropic’s stated policy is that generated text from supported models is marked across its products. Yet the technical and legal position is plainly more complicated for source code than for a generated essay or marketing draft.
Compliance has widened the stakes
The immediate driver is the EU AI Act. Article 50 transparency obligations have applied since August 2, 2026, requiring providers of systems that generate synthetic text, audio, image or video to make outputs machine-readable and detectable as artificially generated or manipulated, as far as technically feasible. The Commission’s related code of practice is voluntary, but is intended to guide implementation.
Anthropic has chosen to apply its marking worldwide rather than only in Europe, saying it does not yet have a durable method of limiting the system by region. Its coverage therefore extends to supported Claude models accessed through its consumer service, API, coding tools and cloud partners.
The policy response cannot rely on watermarks alone. EU guidance distinguishes provider-side machine-readable marking from separate disclosure duties for deployers in certain contexts. A hidden technical signal may assist automated checks, but it is not a substitute for visible disclosure when a rule requires people to be informed.
A signal, not a verdict
The emergence of removal attempts does not make the initiative meaningless, nor does the watermark turn AI provenance into a solved problem. Its value is more limited and potentially more realistic: to provide an additional signal in situations where the original output has travelled largely intact.
For employers, educators, publishers and platforms, that calls for restraint. A detected mark should trigger context, review and questions about how a tool was used, not an automatic finding of misconduct or authorship. A missing mark should be treated with the same caution.
Anthropic’s forthcoming detector will be the first meaningful test of the current workaround claims. It will also begin a longer contest between marking methods and transformation tools. That contest is inherent to text watermarking: language is editable, and the same flexibility that makes generative systems useful also makes their output difficult to trace once people actively rewrite it.
Sources
- Coders Say They Already Found Workarounds to Claude’s Invisible Watermarks — WIRED
- How Claude’s text watermark works — Anthropic
- How Claude marks AI-generated content — Anthropic
- Transparency obligations under Article 50 of the AI Act — European Commission
- Regulation (EU) 2024/1689 — EUR-Lex



