Claude文本水印技术拆解:4美分能洗掉98.3%,AI内容溯源这条路有多难

Anthropic日前发布官方博客,公开了 Claude 文本水印的技术细节。

once this technology was disclosed, the research community immediately ran a round of stress testing. The results are not optimistic: run it through an AI paraphrasing tool, and 98.3% of the watermark signal can be removed, at a cost of about 4 cents. For the price of a cup of coffee, you could strip watermarks from over two thousand articles.

How the watermark is added

The principle is not complicated. When an LLM generates text, every step involves choosing among multiple candidate words, and that choice process originally carries randomness. Anthropic’s approach is to replace the original random numbers with a sequence of numbers derived from a key, making word selection follow a specific statistical pattern.

The text reads exactly the same on the surface, with no change in grammar, style, or quality. Only the party holding the key can determine through statistical testing whether this passage came from Claude. Officials emphasize that the watermark has no impact on output quality.

This differs from the idea of image watermarks. Images can hide identifiers in pixel noise; text has far less room to maneuver, since every character participates in expressing meaning, leaving extremely few places to hide information.

Five hard flaws, officially admitted

Anthropic rarely laid out the complete list of watermark limitations in the blog post.

Short text basically fails. Watermark signals need text long enough to accumulate statistical evidence; short content of a few hundred characters cannot be detected.

Purely factual content works poorly. Ask the model how tall Everest is, and the answer is just a few characters; anyone’s generation looks the same, with no word-choice variance to speak of.

Proofreading scenarios are not covered. If the user asks Claude to fix only a few typos, the vast majority of the text comes from a human, diluting the watermark signal.

Translation shows an abnormally strong signal. Officials found that the watermark signal in translations is stronger than in the original text; no explanation has been given for this counterintuitive phenomenon.

Code is hard to judge. Watermarks can be embedded in code comments and variable naming, but the code logic itself does not have much freedom, leaving the judgment boundary blurry.

There is an even more fundamental problem: the watermark can only prove that Claude may have participated in generation; it cannot distinguish whether the whole piece was written by Claude or Claude only made edits. It also does not change ownership of the work or legal liability attribution; in a copyright dispute, the watermark is of no help.

What the 4-cent breaking cost means

The third-party research data is the most striking. Rewrite Claude’s output once with an off-the-shelf AI paraphrasing tool, and 98.3% of the watermark traces disappear, at a per-article processing cost of about 0.04 USD. Technically it requires no expertise at all; the barrier is nearly zero.

Anthropic has not yet released the details of its detection API, and the commercialization pace is unclear. But as long as paraphrasing attacks cost this little, the actual watermark defense can only stop those who can’t be bothered to act.

The industry is entering the era of labeling compliance

Taking a wider view, the Claude watermark is only a beginning. The transparency clauses of the EU AI Act are turning identifiable AI-generated content into a hard requirement that no one can dodge.

Several news items from the same period can be read together: OpenAI dissolved its prevention team at the end of last month, splitting safety functions into existing business units; Alibaba officially open-sourced the Qwen3.8 series; Qianwen Office launched the GLM-5.3 and DeepSeek V4 Pro flagship models. Competition at the model layer keeps intensifying, while compliance moves are ramping up in step.

For enterprise users, this has three practical implications.

First, do not treat watermarking as the main line of defense for content security. It can pass compliance checks but cannot stop deliberate evasion; important scenarios still need metadata records, audit logs, and other means.

Second, businesses dominated by short or factual content need not pay extra for watermark capability at this stage; watermarking basically does not work in such scenarios.

Third, watch the opening progress of the detection API. Once it opens, content platforms, publishers, and education will plug in, turning whether this text was written by AI into a queryable interface, and the entire content review process will have to adjust.

Final words

The technical approach of the Claude text watermark represents the industry’s current best solution, but the 4-cent breaking cost shows it is far from reliable. The value of watermarking may lie more in drawing responsibility boundaries: it is the platform’s response to regulators and an honest disclosure to users, not a universal anti-counterfeiting label.

The real answer to AI content provenance will most likely require stacking multiple mechanisms such as watermarking, key escrow, and industry-union verification. Until then, keeping moderate expectations of any single-point technology is the more sober attitude.


评论

发表回复

您的邮箱地址不会被公开。 必填项已用 * 标注

苏ICP备2025163703号-2   警徽苏公网安备32010502011527号