What to Know About Digital Watermarks and Identity Verification 

As generative AI becomes more sophisticated, one of the biggest challenges facing technology companies is determining whether a piece of content was created by a human or generated by AI. Digital watermarking is emerging as one potential answer.

Recent news around Claude watermarking its AI-generated text have brought renewed attention to the technology and its potential role in identifying synthetic content. Anthropic isn’t acting alone here: Claude’s watermark is a version of Google DeepMind’s SynthID-Text method, adopted to meet transparency obligations under the EU AI Act that most major model providers have also signed up to. The basic idea is to embed  signals into AI-generated content that are effectively invisible to humans but can later be detected by a system designed to recognize them.

For identity verification, where generative AI can now be used to create or manipulate identity documents, portraits, and other evidence, that capability could become another useful signal for identifying synthetic content.

But there is an important limitation. Fraudsters don’t behave like ordinary users. Any security system built around digital watermarks therefore has to consider not only whether the watermark can identify AI-generated content under normal conditions, but whether that signal survives when an attacker is deliberately trying to destroy it.

What Is a Digital Watermark?

A digital watermark is information embedded within digital content that can later be detected to help identify its origin or authenticity.

Unlike a visible watermark placed over an image, digital watermarks can be designed so that people consuming the content don’t notice anything unusual. Instead, the content contains underlying patterns that a detection system can recognize.

For AI-generated text, for example, watermarking can influence the statistical patterns created as a model selects words during generation. The resulting text still appears normal to a human reader, but sufficiently long passages carry statistical patterns that allow a compatible system to recognize that the content was generated by AI.

What happens when the watermark is removed?
Microblink doesn't rely on watermarks alone.

The concept has potentially significant implications for AI provenance. If a system can reliably recognize content produced by an AI model, organizations have another way to distinguish synthetic content from human-created material.

However, identifying AI content in controlled conditions and detecting it during an adversarial attack are two very different challenges. The developers are candid about this. Google’s own documentation states that SynthID Text is not designed to stop motivated adversaries (Cf. Limitations in https://ai.google.dev/responsible/docs/safeguards/synthid).

The Attacker Doesn’t Have to Preserve the Watermark

Watermarks only work if they are intact to be detected. But a fraudster who knows watermarking exists has no reason to submit the original output unchanged.

Instead, the attacker could pass that text through another service, translate it into a language structurally different from the original, and then translate it back. The resulting passage might communicate essentially the same information, but the statistical properties used to identify the original AI-generated text could be disrupted.  This isn’t hypothetical, and it isn’t a criticism that the vendors dispute. Google’s own SynthID documentation lists it as a limitation: detection “can be greatly reduced when an AI-generated text is thoroughly rewritten, or translated to another language.”.

Researchers have measured how much. In one study, translating watermarked text into Chinese and back into English let roughly a third of AI-generated passages slip past the detector completely. In a separate analysis from ETH Zurich, researchers were able to scrub the watermark in more than nine out of ten attempts. Anthropic’s own guidance points the same way, noting that light editing probably won’t remove Claude’s watermark but a complete rewrite will

The attacker hasn’t necessarily defeated the watermark detector itself. Rather, they’ve altered the evidence the detector was expecting to examine. Security systems don’t operate against users who politely preserve the signals designed to catch them. They operate against adversaries whose goal is to understand those signals, manipulate them, and find ways around them. That’s why AI security requires thinking like an attacker.

What Do Digital Watermarks Mean for Identity Verification?

So, what does this all mean for identity verification? 

As AI-generated imagery improves, watermarking could provide valuable information about whether an image originated from a particular generative AI system. In identity verification, that could potentially provide another signal when assessing whether a submitted identity document or portrait is genuine.

Suppose someone creates a synthetic identity document using an AI image generator that embeds a detectable digital watermark. If the attacker uploads that original digital file directly into an identity verification workflow, watermark detection may be useful.

But, as noted above,  a sophisticated attacker will likely not preserve the original file.  They could display the synthetic document on another screen and photograph it. They could print the document and recapture it using a camera. They could potentially crop, resize, compress, edit, or otherwise transform the image before submitting it.

Each transformation creates an opportunity to weaken or destroy information embedded in the original digital asset. The document may look substantially the same to a human observer, while the underlying signal used to establish its provenance may no longer survive.

Digital Provenance Is Merely One Signal

None of this means digital watermarking isn’t valuable. Knowing that an image or piece of text was generated by AI could be an extremely useful signal. As standards and technologies for establishing digital provenance improve, these signals may become increasingly valuable for platforms attempting to understand where digital content originated.

But in fraud prevention, a good rule of thumb is that no single signal should be treated as definitive. A watermark can potentially tell you something important when it’s present. Its absence proves nothing. If an identity document contains a recognizable generative AI watermark, that information could contribute to a fraud decision. But if the watermark isn’t detected, the system still needs to determine whether the document itself is genuine.

That requires looking beyond provenance and examining the evidence contained within the identity interaction itself.

Detecting the Attack, Not Just the Tool That Created It

Generative AI has dramatically lowered the barrier to creating convincing synthetic identity documents and other fraudulent content. But attackers aren’t limited to one model, one technique, or one generation of technology.

They can switch models. They can use services from different providers, or run open-weight models locally that carry no watermark at all. They can manipulate generated content after creation. And they can move synthetic content between the digital and physical worlds specifically to eliminate signals designed to identify it.

That’s why identity verification needs multiple layers of evidence.Document authenticity, signs of digital manipulation, generative AI detection, portrait analysis, image quality, device and behavioral signals, and the context surrounding an interaction can all contribute to understanding whether an identity can be trusted.

The objective isn’t to identify which AI model an attacker happened to use. It’s to determine whether the identity evidence being presented is legitimate. Digital watermarking can strengthen that process, but it shouldn’t become a shortcut around it.

Winning the AI Security Race Means Thinking Like an Attacker

The rapid development of digital watermarking is an encouraging step toward greater transparency around AI-generated content. As generative AI becomes embedded throughout the digital world, technologies that help establish where content originated will become increasingly important.

Attackers don’t simply adopt new technology. They study the defenses built around it and look for ways to remove, manipulate, or bypass the signals those defenses depend upon.

For identity verification, that means preparing for the document after the watermark has disappeared, the AI-generated image after it has been printed and recaptured, and the synthetic content after it has passed through several transformations. Digital watermarks can help tell us where content came from. Identity security still has to determine whether we can trust what remains.

August 27, 2026

Découvrez nos solutions

L’exploration de nos solutions est à portée de clic. Essayez nos produits ou discutez avec l’un de nos experts pour approfondir notre offre.