Google figures out how to watermark AI-designed proteins

Researchers at Google DeepMind described a method to embed detectable watermarks into AI-designed protein sequences without breaking their function. The approach adapts Google’s SynthID watermarking concept to work with ProteinMPNN, a widely used protein design tool, allowing tagged designs to be distinguished from untagged sequences when the watermark key is known.

By AI Newsroom· Reviewed by Pranav, Founder & Editor-in-ChiefPublished about 2 hours agoUpdated about 2 hours ago0 views
Google figures out how to watermark AI-designed proteins

Why It Matters

If reliable, this technique could let trusted designers mark AI-created proteins so that downstream screening and surveillance can flag or prioritize sequences that lack such a mark, addressing a gap in biosecurity for AI-designed proteins. The method aims to make AI-generated proteins identifiable while preserving their intended biochemical activity.

Key Facts

  • Who: DeepMind team at Google
  • What: Published a research paper proposing protein watermarking (SynthIDBio) for AI-designed proteins
  • Works with: ProteinMPNN, a popular AI protein design tool developed by the Baker Lab
  • Based on: Google’s SynthID digital watermarking technology, adapted to proteins (SynthIDBio)
  • Proteins: Built from 20 amino acids; proteins of ~500 amino acids noted as comparatively small for hiding signals

Google’s DeepMind researchers propose a way to embed a covert but detectable watermark directly into protein amino-acid sequences designed by AI. The technique, called SynthIDBio, adapts the company’s SynthID digital watermarking approach to the constraints of protein chemistry. Rather than altering the protein structure after design, SynthIDBio biases the amino-acid choices made during the design process so that the resulting sequence carries a statistical signature tied to a secret key. The team implemented SynthIDBio in the second stage of the ProteinMPNN workflow, which places side chains onto a precomputed protein backbone one residue at a time. At each position the watermarking module suggests amino acids based on a key and the identities of previously chosen residues; ProteinMPNN then accepts only those suggestions that are compatible with forming a functional protein. In practice this means the watermark is inserted only where chemically and structurally tolerated. Because the approach leverages positions where multiple, chemically similar amino acids could function interchangeably (for example leucine, isoleucine, valine, or serine versus threonine), the watermark ends up randomly distributed across the sequence rather than concentrated in any single region. Detecting the watermark requires scanning the full sequence using the encoding key and measuring how often SynthIDBio’s suggested residues appear in the final protein. The paper argues the watermark can survive the tight constraints of protein design — limited amino-acid alphabet and often short sequence lengths — without abolishing the protein’s intended activity because ProteinMPNN filters out substitutions that would inactivate the design. The researchers also note the method is intended to let trusted designers label their outputs, enabling other sequences to be treated with heightened scrutiny in biosecurity screening.

Keep Reading