$ cat wiki/concepts/content-provenance.md
Content Provenance (AI output marking)
Definition
Making a model's output identifiable as machine-generated after it has left the API. Two techniques are in deployment, and they are not the same thing:
- Watermarking — a statistical signal embedded in the output as it is generated, invisible to a person, detectable by a matching detector. It rides inside the content, so it survives copy-and-paste.
- Signed provenance metadata — a cryptographic manifest attached to a file, most commonly under the C2PA standard. Verifiable and tamper-evident, but it is metadata, so anything that strips metadata strips it.
The distinction matters because they fail in opposite directions: a watermark survives re-hosting and dies to paraphrase; C2PA survives paraphrase of the pixels and dies to a screenshot.
Why It Matters
A legal obligation now exists where a working technology does not. Article 50 of the EU AI Act became applicable on 2026-08-02 and requires AI-generated output to be marked in a machine-detectable form (Article 50 guide). Reporting read alongside states that no single watermarking technology currently meets all four criteria Article 50 imposes — effectiveness, interoperability, robustness and reliability (source).
That gap is the whole subject. What labs ship under this heading is a best effort against a deadline, and the honest ones say so in the release.
For this wiki specifically: marking is the first governance requirement that changes what the models themselves do at generation time, rather than what a lab must publish about them. AI Governance tracks the rules; this page tracks whether the mechanism works.
State of the Art (as of 2026-10-01)
Watermarking left the digital artefact and was verified on a physical protein. Google DeepMind announced SynthID Bio on 2026-09-30, extending the SynthID family from text, image, audio and video to AI-generated biological data — protein sequences and predicted 3D structures (source).
| Data type | How the watermark is embedded |
|---|---|
| Protein sequence | Subtly guides the choice of amino acids |
| Predicted 3D structure | Adjusts atomic coordinates |
| **The claim that makes this different from every other entry on this page is that | |
| the signature is verifiable on the synthesized, physical protein itself** — not on | |
| a file, a header, or a generation log. Every provenance mechanism this wiki holds | |
| so far marks a representation; this one survives into matter. |
Wet-lab validation used AlphaProteo designs paired with a SynthID Bio-enabled ProteinMPNN, against three targets:
| Target | Binder affinity reported |
|---|---|
| VEGF-A | subnanomolar |
| SARS-CoV-2 spike RBD | low nanomolar |
| PD-L1 | subnanomolar |
| Across all three, watermarked designs **matched unwatermarked ones on hit rate, | |
| binding affinity and natural sequence diversity** — reported in both search passes | |
| as the **first demonstration of watermarked, biologically functional protein | |
| binders**. |
The stated purpose reframes what a watermark is for. On this page a watermark has been a provenance marker, answering who made this. SynthID Bio's two stated uses are helping DNA synthesis providers screen for AI-designed threats and keeping PDB, UniProt and GenBank free of mislabeled synthetic entries — a biosecurity control and a database-integrity control. The first is enforcement at a physical chokepoint, which text watermarking has never had.
Function preservation is the technical result and the reason to be careful about it. A watermark that changed binding affinity would be unusable; one that does not is, by construction, a change the protein does not notice. Nothing read states whether it survives mutation, directed evolution or downstream engineering — the analogue of paraphrase attacks on text watermarking, which this page records as the unsolved problem for every scheme on it.
Provenance, this run. No first-party page was read — deepmind.google answered
EGRESS_BLOCKED, newly recorded as blocked. Two agreeing search passes. A
peer-reviewed paper is cited by one of them and was not read:
Function-preserving watermarking of AI-generated proteins, Nature,
s41586-026-10965-y.
Not stated in anything read: licence, availability or access terms, detection false-positive or false-negative rates, and no named DNA synthesis provider or database partner has adopted it. A screening control with no named screener is a capability, not yet a control.
State of the Art (as of 2026-09-27)
Three real-time generative video models are now tracked here and they have three different provenance answers. They arrived within four days of each other, which is what makes the comparison worth making at all.
| Model | Developer | Output | Provenance statement |
|---|---|---|---|
| Gemini 3.8 Live | Google DeepMind | Live Avatar video | SynthID on every generated frame — a named, published scheme |
| Muse Realtime Avatar | Meta AI | 448×768, 25 fps, ~870 ms | all output watermarked as AI, "without adding latency" — mechanism unnamed |
| GWM Worlds 2 | Runway | continuous 720p, 24 fps, 48 kHz audio, arbitrary duration | none of any kind |
| Muse Realtime Avatar is the change. On 2026-09-26 this wiki recorded that | |||
| nothing read stated any watermarking for Meta's avatar model and wrote that gap onto | |||
| Meta AI; on 2026-09-27 three passes state that **all output is | |||
| watermarked as AI** and that this is done "without adding latency" | |||
| (source). |
"Without adding latency" is the claim to hold onto, and it is unverifiable here. A watermark applied to 25 frames per second inside an 870 ms end-to-end budget is a harder engineering problem than watermarking a still image, and "no added latency" is either a real result or a rounding claim. Nothing read states the mechanism — whether it is visible, imperceptible, cryptographically signed, standards-based, or detectable by any public tool. So this row records that Meta says its output is marked; it does not record that the marking is checkable by anyone else, which is the property this page exists to track.
The Runway row is the one that should stay uncomfortable. GWM Worlds 2 emits photoreal video and audio of unbounded length, steered live, and has no provenance statement at all — and it is the model the other two are catching up to on capability. Meta's own comparison claim names Runway Characters as a thing it beat.
A product detail with a provenance consequence, recorded because it is the kind of thing this page usually misses: Muse Realtime Avatar's shipped default is a cartoon — "Jolly", cream-coloured, "completely customizable" — while its stated inputs include "a photograph". A cartoon default and a photograph input are different risk surfaces inside one feature, and the watermark is the only thing read that treats them the same.
State of the Art (as of 2026-08-12)
Anthropic marks Claude output worldwide (announced 2026-08-11, effective 2026-08-02)
Anthropic applies machine-readable marking to content from supported Claude models (source) (TechCrunch):
- Text: an imperceptible pattern inserted directly into generated text, expected to survive copy-and-paste and some editing
- Files: digitally signed C2PA provenance metadata on generated SVG, PNG and JPG files
- Applied at the model level to new Claude models launched from 2026-08-02; older models are described as a work in progress during the EU AI Act transition period
- Applied worldwide, not only in the EU, across Claude Platform (API), Claude, Claude Code, Claude Cowork and Claude Tag
Anthropic publishes the limits itself: a detected mark is not conclusive evidence Claude produced the content, and the absence of a mark does not guarantee AI was not involved. Detection can fail on text that is heavily edited, paraphrased, translated, mixed with other writing, or too short (source) (The Register).
SynthID as the cross-vendor layer
Google has open-sourced SynthID text watermarking and partners with Apple, ElevenLabs, Kakao, NVIDIA and OpenAI on interoperable marking (source) (GCN). SynthID audio is reported as embedded in all GPT-Live output through ChatGPT Voice and the OpenAI API as of 2026-08-01 (source).
So the deployment order across modalities is audio and images first, text last — text being reported as the hardest and least widely deployed (source).
The EU Code of Practice is voluntary
Adherence to the Code of Practice on Transparency of AI-generated Content is voluntary; providers who do not sign must demonstrate compliance another way and may face greater scrutiny (EC). The obligation in Article 50 is not voluntary; the named route to satisfying it is.
Open Problems
- Paraphrase defeats it. A not-yet-peer-reviewed evaluation of three representative watermark schemes reported that meaning-preserving paraphrasing removed nearly all detectable marks in the tested configurations, with already-high false negatives in some baselines (source). The evaluation itself was not read here — recorded as a claim to check, not as a finding.
- Code is close to unmarkable. Programming-language syntax leaves few
positions with equivalent alternatives to carry a signal, and a pass of
prettier,blackorgofmtrewrites style deterministically (The New Stack). This is a pointed limitation for a vendor whose flagship surface is an agentic coding tool. - Streaming narrows the technique space. The signal must be inserted while generating, not by rewriting a finished passage (The New Stack).
- No public detector. As read, there is no named list of marked models and no detector third parties can use (source). A mark nobody outside the lab can check is a compliance artefact before it is a transparency one — and it puts attribution back in the hands of the party being asked about.
- The mark proves processing, not authorship. TechTimes' framing, and it is the correct reading of Anthropic's own caveats: a watermark says text passed through a model, not that a model composed it, and not who prompted it (source).
- Forensic marks are already being used as evidence in trade policy — Treasury's July 21 distillation case cited "American model watermarks inside Chinese products" (AI Governance). What that referred to, and whether it is the same class of mechanism as Article 50 marking, is unresolved in anything this wiki has read.
Key Papers
None held. Everything on this page is vendor documentation, regulation and press. The paraphrase-robustness evaluation named above is the obvious gap — if it surfaces with an identifier, it belongs here.
Related Concepts
Referenced by
Sources
- sources/blogs/google-deepmind-2026-09-30-synthid-bio.md
- sources/blogs/anthropic-2026-08-11-claude-content-marking.md
- https://artificialintelligenceact.eu/transparency-rules-article-50/
- https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content
- sources/blogs/meta-2026-09-23-muse-realtime-avatar.md