Skip to content

Anthropic Begins Watermarking Claude Content, Including Human-Written Text It Edits

Anthropic is rolling out machine-readable watermarks across Claude, including human-written content that is edited or processed by its AI models.

Table of Contents

Anthropic is adding invisible, machine-readable markings to text, images, and other files processed by its Claude models, rolling out a new provenance system globally as the company moves to comply with transparency requirements under the European Union's AI Act.

The system will use different approaches depending on the type of content. Supported files, including PNG, JPG, and SVG images, will carry digitally signed provenance metadata using the Coalition for Content Provenance and Authenticity (C2PA) standard, while text will receive an invisible watermark embedded directly into the generated content.

But here's the catch: a watermark does not necessarily mean Claude created the underlying work.

Content can carry a Claude mark even when the model was not its original author. That includes material users ask Claude to proofread, translate, summarize, or convert, meaning human-written text that is simply copy edited by Claude could ultimately carry the same type of machine-readable signal.

Anthropic therefore cautions that a detected mark signals content was processed by Claude, and not conclusive evidence that Claude authored it. Marked content can also be modified or combined with other material after leaving Claude, further limiting what the watermark alone can establish about its provenance.

For text, the watermark is designed to remain invisible and does not alter the meaning, quality, or readability of Claude's output. Instead of attaching metadata to a document, text watermarking creates a machine-detectable statistical signal through the choices the model makes while generating its response. The approach can survive copying and pasting and some subsequent editing without requiring a visible label.

Anthropic says heavily editing, paraphrasing, translating, or mixing marked text into other writing can make the watermark undetectable. Likewise, the absence of a detectable mark does not prove that content was written by a human.

Images and other supported files will instead use C2PA, an open provenance standard that attaches cryptographically signed information about a file's history. The technology can indicate that an image or file was generated or processed by Claude and help reveal subsequent modifications. Adobe and Google are among the other major technology companies using C2PA, although provenance metadata can also be stripped through actions such as file conversion or taking a screenshot.

The markings are being introduced as part of Anthropic's commitments under the EU AI Act's Code of Practice on transparency for AI-generated content. Claude models launched in the EU on or after August 2 will support machine-readable marking from launch, while Anthropic is working to extend the system to older models.

Although prompted by European regulation, Anthropic is applying the measures globally. Because the markings operate at the model level, they will extend across Anthropic's products and services, including Claude, the Claude API and Platform, Claude Code, Claude Cowork, and Claude Tag, as well as Claude models accessed through cloud partners.

Comments

Latest