AI·News & analysis
Google DeepMind is putting invisible watermarks inside AI-designed proteins
Google DeepMind unveiled SynthID Bio, a watermarking system that embeds detectable signatures into AI-generated protein sequences, surviving even after they're synthesized into real proteins.

Tide
Ripple
Sci-fi
6/10
Reality
Shipping
A hidden signature baked into a protein, even after it's real.How we rate
Google DeepMind launched SynthID Bio, embedding hidden, verifiable watermarks into AI-generated protein sequences and structures that survive even after synthesis into real, physical proteins.
It targets a genuine biosecurity gap: AI-designed sequences can slip past DNA synthesis screening because they don't resemble any known threat. As AI protein design tools become more powerful and widely used, being able to trace whether a given protein came from an AI system, and which one, becomes a meaningful new layer of biosecurity infrastructure.
What to know
- Google DeepMind launched SynthID Bio, a watermarking system that embeds imperceptible signatures into AI-generated protein sequences and their predicted 3D structures.
- The watermark stays detectable even after a sequence is synthesized into a real, physical protein, not just while it exists as digital data.
- It addresses two problems: AI-generated protein sequences can slip past DNA synthesis security screening, and mislabeled AI-generated entries can corrupt scientific databases like the Protein Data Bank and UniProt.
- Testing across three real protein targets found watermarked proteins performed comparably to unwatermarked ones on binding affinity, hit rates, and sequence diversity.
- DeepMind is publishing the methods paper, open-sourcing the code and data, and releasing model weights, working with partners including Adaptyv Bio, Stanford's Hie lab, and the Arc Institute.
Google DeepMind just extended its AI watermarking technology somewhere it's never gone before: into the actual chemistry of proteins, in a way that survives long after the AI's design leaves the computer.
What exactly does SynthID Bio watermark?
SynthID Bio embeds imperceptible, verifiable signatures directly into AI-generated protein sequences and their predicted 3D structures. That's a meaningful step beyond DeepMind's existing SynthID technology, which has previously watermarked AI-generated images, text, and audio.
Why it matters: a watermark in an image or a block of text only ever exists as digital data. A protein sequence can be synthesized into an actual, physical molecule, and SynthID Bio's watermark is built specifically to remain detectable even after that synthesis happens, not just while the design lives on a computer.
What problem is this actually solving?
The catch: AI-generated protein sequences can slip past traditional DNA synthesis security screening precisely because they don't resemble any known biological threat. Screening systems built to catch dangerous sequences by comparing them against known patterns have nothing to compare a genuinely novel AI-generated sequence against, making manual review both necessary and burdensome at scale.
There's a second, quieter problem too: mislabeled AI-generated entries can end up in major scientific databases like the Protein Data Bank and UniProt, misleading future research and compromising decisions that rely on knowing whether a given protein entry is naturally occurring or AI-designed.
In real life it's like being able to tell, just by looking closely at a finished cake, exactly which recipe and which baker made it, even after it's been fully baked and frosted.
How does the watermarking process actually work?
The method adapts depending on the type of data:
- For protein sequences: SynthID Bio subtly guides which amino acids get selected during AI design, creating a detectable statistical signal without meaningfully altering the sequence's function.
- For predicted 3D structures: it fine-tunes DeepMind's AlphaFold 3 diffusion network to embed the signature directly into the predicted atomic coordinates of the structure itself.
Why it matters: watermarking both the flat sequence and the folded 3D structure means the signature can potentially be recovered whether someone is looking at raw genetic code or at a protein's actual physical shape, covering two very different ways researchers and screening systems might encounter an AI-designed protein.
Does adding a watermark make the protein worse at its actual job?
DeepMind tested watermarked proteins against three real biological targets: VEGF-A, the receptor-binding domain of the SARS-CoV-2 spike protein, and PD-L1, a protein involved in cancer immunotherapy. Across all three, watermarked proteins performed comparably to unwatermarked versions on binding affinity, hit rates, and natural sequence diversity.
Why it matters: a watermarking system that noticeably degraded a protein's real-world function would be a non-starter for actual researchers. Matching unwatermarked performance on these specific benchmarks is what makes this a genuinely usable tool rather than just a proof of concept.
Who actually uses something like this?
DeepMind positioned SynthID Bio for three groups: DNA synthesis providers conducting security screening on incoming orders, administrators of major scientific protein databases, and researchers using AI protein design tools who want to responsibly flag their own AI-generated designs as such.
Background: DeepMind is publishing the underlying methods paper, open-sourcing the code and data, and releasing model weights to the broader research community, rather than keeping the technology closed or proprietary. That approach mirrors how DeepMind has released other SynthID tools, treating broad adoption as more valuable than exclusivity for a technology meant to function as shared infrastructure.

Who did DeepMind work with on this?
The project involved Adaptyv Bio for in vitro, lab-based validation of the watermarked proteins, confirming the signal survives real biochemical synthesis rather than just working in simulation. Separately, Stanford University's Hie lab and the Arc Institute collaborated on extending the watermarking approach specifically to bacteriophages, viruses that infect bacteria, increasingly used in engineered research and therapeutic applications.
Why it matters: extending the technique beyond individual proteins to more complex engineered biological systems like bacteriophages suggests DeepMind sees this as a foundation for broader provenance-tracking across synthetic biology, not a one-off tool limited to a single type of molecule.
Does this actually solve the biosecurity problem?
Not on its own. Biosecurity policy expert Sarah Carter described SynthID Bio as "an important piece of the puzzle for tracking the provenance of biological designs," language that frames it as one layer of a larger effort rather than a complete fix.
Who's affected: that framing matters because biosecurity in synthetic biology depends on multiple overlapping systems working together, screening tools, database integrity checks, regulatory oversight, and now watermarking. A single new tool, however well-designed, doesn't replace the need for those other layers to keep improving in parallel.
Why does this trace back to AlphaFold specifically?
Background: DeepMind's AlphaFold system won its creators, Demis Hassabis and John Jumper, a share of the 2024 Nobel Prize in Chemistry for solving a roughly 50-year-old scientific challenge: accurately predicting a protein's 3D structure from its amino acid sequence. Since its release, AlphaFold has predicted the structure of more than 200 million proteins and been used by over 2 million researchers worldwide, becoming foundational infrastructure across modern biology research.
That same predictive power is exactly what makes biosecurity a real concern. The ability to accurately model how novel protein structures will fold and behave could, in principle, help someone design proteins with harmful properties just as easily as it helps someone design beneficial ones.
When DeepMind released AlphaFold 3, the team said it had consulted more than 50 experts in biosecurity, bioethics, and AI safety, ultimately concluding the system's marginal biosecurity risks were outweighed by its scientific benefits.
Why it matters: SynthID Bio is a direct continuation of that same risk conversation, not a separate initiative. It's DeepMind building a specific technical safeguard, provenance tracking through watermarking, aimed at the same category of risk that AlphaFold's own release raised, rather than leaving that concern unaddressed as the technology keeps advancing.
What it means for you
- This is infrastructure for researchers and screening systems, not something you'll interact with directly, but it shapes how safely AI protein design tools can be used as they become more powerful and widespread.
- It's a sign AI safety work is expanding beyond text, images, and audio into physical biology, a meaningfully higher-stakes domain where the gap between digital design and real-world consequence is much smaller.
- The open-source release matters for adoption. A watermarking standard only helps biosecurity broadly if DNA synthesis providers and database administrators actually implement it, which an open, freely available tool makes more realistic.
- Watch for this approach extending to other AI-designed biological systems beyond individual proteins, following the same pattern as the bacteriophage extension already underway.
The bottom line
SynthID Bio is a genuinely novel extension of AI watermarking into physical biology, solving a real, specific problem: telling whether a given protein came from an AI system after it's already been synthesized into something real.
It's not a complete biosecurity solution on its own, and DeepMind isn't claiming it is. But as AI protein design tools keep getting more capable, having a working, open-sourced way to trace a molecule's AI origins even after synthesis is the kind of unglamorous infrastructure that matters more the further this technology spreads.
Key facts
- Technology
- SynthID Bio
- What it watermarks
- AI protein sequences + 3D structures
- Survives
- Synthesis into real, physical proteins
- Tested on
- VEGF-A, SARS-CoV-2 spike RBD, PD-L1
- Release
- Open-sourced code, data, model weights
Got questions?
Quick answers, plain wordsWhat is SynthID Bio?
A watermarking technology from Google DeepMind that embeds imperceptible, verifiable signatures directly into AI-generated protein sequences and their predicted 3D structures, extending DeepMind's existing SynthID watermarking work (previously used for AI images, text, and audio) into biology.
How is this different from watermarking AI text or images?
Unlike text or images, a protein sequence can actually be synthesized into a real, physical molecule. SynthID Bio's watermark is designed to remain detectable even after that synthesis happens, not just while the design exists as digital data.
What problem does this actually solve?
Two problems: AI-generated protein sequences can bypass traditional DNA synthesis security screening because they don't resemble any known biological threat, and mislabeled AI-generated entries can quietly corrupt scientific databases like the Protein Data Bank and UniProt, misleading future research.
How does the watermarking process actually work?
For protein sequences, it subtly guides which amino acids get selected to create a detectable statistical signal. For predicted 3D structures, it fine-tunes DeepMind's AlphaFold 3 diffusion network to embed the signature directly into the predicted atomic coordinates.
Does watermarking hurt how well the designed proteins actually work?
DeepMind's testing across three real targets, including a SARS-CoV-2 spike protein region and a cancer-related protein called PD-L1, found watermarked proteins performed comparably to unwatermarked ones on binding affinity, hit rates, and natural sequence diversity.
Who is actually supposed to use this?
DNA synthesis providers conducting security screening, administrators of scientific protein databases, and researchers who use AI protein design tools and want to responsibly label their AI-generated designs.
Is this open source?
DeepMind said it's publishing the methods paper, open-sourcing the underlying code and data, and releasing model weights to the broader research community, rather than keeping the technology proprietary.
Who did DeepMind work with on this?
Adaptyv Bio for in vitro (lab-based) validation of the watermarked proteins, and Stanford University's Hie lab along with the Arc Institute on extending the watermarking approach to bacteriophages specifically.
What's a bacteriophage, and why does it matter here?
A bacteriophage is a virus that infects bacteria, increasingly used in engineered forms for research and even therapeutic purposes. Extending watermarking to bacteriophages means the same provenance-tracking approach can apply beyond individual proteins to more complex engineered biological systems.
Does this eliminate biosecurity risks from AI-designed proteins?
No single tool eliminates the risk on its own. Biosecurity policy expert Sarah Carter described SynthID Bio as 'an important piece of the puzzle for tracking the provenance of biological designs,' suggesting it's one layer of a broader effort rather than a complete solution by itself.
SourcesGoogle DeepMind
Topics and tagsGoogle, deepmind, ai, biosecurity
Related stories

A startup wants satellites to think for themselves instead of waiting for instructions from Earth
Satlyt, founded by a former Google and SpaceX product manager, raised $8 million to build software that lets AI models run directly on satellites, cutting the need to send every decision back down to ground control.

Google is quietly funding Hollywood movies to make AI look less scary
Google partnered with talent agency Range Media Partners on 100 Zeros, a fund that co-produces films and TV shows aiming to show technology, especially AI, in a more optimistic light than shows like Black Mirror.

Japan's biggest power company is building its own gas plant just to feed one AI data center
JERA, Dell Technologies, and RHAELM signed a deal to build a $15 billion, 400-megawatt AI data center powered by its own dedicated energy supply near Tokyo, aiming to start operations around 2028.
More in brief
- arXiv limits researchers to two papers a month as AI drives submissions to a recordOct 2
- California will fine robotaxi companies that block first responders for over 30 minutesOct 2
- Microsoft launches real-time transcription and new voice models for AI voice agentsOct 1
- Apple's smart home hub reportedly launches October 13, with a camera that never records videoOct 1
- Cloudflare releases Clef, open-weight AI models that make yes-or-no decisions fastOct 1
- GrayKey maker reportedly found a way around the iPhone's Inactivity RebootOct 1