Software engineer Sean Goedecke has highlighted the fundamental difficulties of implementing effective AI watermarking for text, noting that these methods are often trivial to remove. Unlike image watermarking, which can utilize visual noise, text is a highly compressed medium where even subtle changes can be detected by humans or impact the quality of the output.
PLUS ULTRASecurityPolicySean Goedecke
Text AI watermarks are technically difficult and easily removed
PLUS ULTRA by Amenoyomi
There are currently two primary approaches. Google's SynthID uses a probabilistic weighting of tokens to create a mathematical pattern that is difficult for humans to perceive but easy for machines to aggregate and detect. Alternatively, some providers like OpenAI and Anthropic may use Unicode homoglyphs—replacing standard characters with visually identical but different Unicode characters—to embed patterns.
However, both methods possess significant vulnerabilities. SynthID-based watermarks can be stripped by using an LLM to paraphrase the content, while homoglyph-based marks are easily removed by replacing the special characters with their standard equivalents.
These technical limitations pose a challenge for the European Union's AI Act, which requires AI-generated content to be "detectable as artificially generated." While the Act also mentions digitally signed metadata (such as C2PA), such methods are primarily designed for file formats and are less suitable for the plain text produced by chat interfaces.
PLUS ULTRAby Amenoyomi
The fundamental difficulty of text watermarking stems from the nature of text as a highly compressed medium. In digital images, vast amounts of "noise" exist that the human eye cannot perceive, allowing data to be hidden without affecting the visual quality. Text, however, lacks this overhead; almost any change to a sentence is either noticeable to a human reader or degrades the quality and coherence of the output, making invisible embedding inherently difficult.
Current probabilistic methods, such as Google's SynthID, attempt to overcome this by embedding mathematical patterns into the selection of tokens. While this avoids obvious typos, the watermark is tied to specific vocabulary choices. Consequently, these patterns are easily destroyed through paraphrasing, as simply rewording the content while preserving the meaning breaks the mathematical relationship required for detection.
Alternative approaches using Unicode homoglyphs—replacing standard characters with visually identical ones—are easier to implement but equally fragile. Because this method relies on simple character substitution, the watermark can be completely erased by an automated process that replaces all non-standard homoglyphs with their standard character equivalents.
Other verification strategies are limited by technical constraints. Attempting to verify AI generation by measuring how closely a text matches a model's predicted tokens results in too many false positives, as humans may accidentally write in a style that mimics an AI. Additionally, digitally signed metadata like C2PA is ineffective for the plain text produced by chat interfaces, as there is no file container to hold the signature.
Sources
- テキストAIの透かし・ウォーターマークは技術的に困難な上に簡単に削除できる (GIGAZINE, 2026-09-19)
- ショーン・グーデッケ