Finally, the previous implementations were very silly! OpenAI had a nice tool, but it only detected their own watermarks, and Google's process was "upload the image to Gemini and ask if it's AI generated", which seemed like a perplexing waste of tokens and breath.
(While a sibling comment points out putting the image through Google Image Search as an alternative, I don't remember this being signposted in the support article I've read, so I unfortunately didn't know about it)
The “ask Gemini” thing was so bad, because it made people think that worked with other chatbots. I know some smart, well-informed people who have been under the impression that they can ask Claude if something is AI-generated and get an accurate answer!
Naively thought from the title that this was a way to detect what synthesizer / settings were used to create a sound in a song and I was pretty excited about that.
It's a real shame Google isn't more transparent about how it actually works, and doesn't provide any mechanism for classifying images in bulk or offline.
They don't want you to know, because they're the baddies. Imagine how much money they can get from advertisers if they're able to identify every single piece of code, reddit thread, email, GitHub readme that you've ever written based on your secret ID. They're creaming themselves just thinking about it.
"Excellent work acquiring that outlook data Anat, now let's cross check the defunct company emails against YouTube videos edited with Google PrivateEditAI™ to figure out what the social security number of this YouTube account is."
Unfortunately this is the truth, no one can be given a benefit of the doubt because despite years of good grace they have been anti-user and enshittified everything they touch.
Not that I can think of, but you could have two watermarks, one detectable with an open-source classifier and the other proprietary.
Most AI images are either extremely low-effort or not actively trying to be deceptive, so defenders can still catch the majority. If someone is actively circumventing they'll probably circumvent both, anyway.
Even Google themselves previously offered a no-auth way (just very niche): if you uploaded an image to Google Image search, and went to "About this image", it would show if the image was Google AI-generated. Now it doesn't.
I've never seen it mentioned anywhere so I will just here complain that Google also removed the ability to paste images into Google Image search some years ago for no apparent reason. It worked great and I used it all the time.
In similar fashion in the Android Play Store they moved the search off from the top bar. For some time they educated users that "oo the search is no longer here, you need to click there -> to get to the page with search!". That has stopped (for me?), but they still haven't found any other use for same place in the first view.
Yeah I've been wanting to run a ton of Wikipedia images through the detector to see if fakes snuck in. Doesn't seem to be a practical way to do this. Even Open AIs tool has a low rate limit
That’s the justification, but the EULA incorporates Google’s regular consumer terms; which means what you upload can also be used for advertising and targeting; and basically any purpose whatsoever by Google.
We need to inform everyone that this technology can encode database identifiers and enough entropy to uniquely identify you as an author (or downloader).
They are not spymarks if they don't encode any personal IDs and are merely used to indicate that an image was AI-generated. I don't think Google or OpenAI use SynthID to include personal data.
It hasn't. Their DoubleCkick/AdSense (embedded website ads) used tracking, but Google built its "empire" through AdWords (keyword ads in Google Search), not AdSense.
Google built it's empire with PageRank which made their search engine popular enough that they could make money on AdWords.
The spying and tracking started pretty quickly though and they've repeatedly demonstrated that they can't be trusted. They've even violated the law on several occasions in efforts to collect data they had no right to.
"If you have something that you don't want anyone to know, maybe you shouldn't be doing it in the first place." - Eric Schmidt CEO of Google
Come on man, they're both advertising companies. Would you resist the opportunity to build that social graph and detect a known user anywhere on the internet?
Of course. Privacy is a feature that increases their product value. Moreover, they could get into trouble with the GDPR if they secretly included tracking information in their images.
So the obvious move is that marketing materials sing how it’s super private, while actual steganography encodes everything to uniquely identify you. If someone discovers something wave it away (and crack down on pesky whistleblowers), eat the fines if the worst happens.
- Apple. Least of all, but still fits: PR pushes double on privacy, but their interests are not their end-users, they just happen to align on some points as they’re not in adtech business.
- Microsoft. The classic case for getting slaps on the wrist, Gates said one thing he’d do differently is sending lobbyists earlier, or something to that effect, didn’t he?
- Amazon, Netflix, Uber, Airbnb and many more are still all variants of the same story. It’s not a lie that everyone nowadays expects most megacorps to fuck people over if there’s money in that.
That’s the Order of “Feel for It Again”, with oak leaf clusters already, no? Something’s cursed in this industry, every company that sells software to individual people ends up the same when they become gigantic and endowed by wealth and power, even if they were decent in their early days. They all become those blind heralds of paperclip^W shareholder value(tm) maximizers, and I’m finding myself struggling looking for an exception.
Am I to believe this time it’s going to be any different, to set those - not baseless, I’d say - cynical expectations aside? Any signs there’s a change?
So, I disagree - it’s obvious, given the ad-adjacent end-user-facing megacorp context, OpenAI is right there. Cynical - sadly so, yes. But I’m sure it’s obvious: the idea of such a possible direction comes up pretty quickly and reliably so.
> Moreover, they could get into trouble with the GDPR if they secretly included tracking information in their images.
Do you think that matters to Google? They've been fined repeatedly for violating GDPR already. It's like multiple violations and fines every single year going back to 2019. They don't care about breaking the law if it gets them what they want. They've been found guilty of breaking the law in multiple countries, including the US. They've easily been able to afford every fine and still come out massively profitable. Google is above the law and they know it.
>I don't think Google or OpenAI use SynthID to include personal data.
Why? I assume by default that it's uniquely identifiable, because it just makes practical sense for OAI and Google (tracking misuse at the very least, but also data), it's trivial to implement and trivial to hide or plausibly deny.
I don't think it's trivial implement. Encoding more bits is harder or more brittle, and not trivial to hide. One could just ask for a completely white image. Any image distortions must then be due to SynthID. If the distortions differ for different accounts, that's strong evidence for varying watermarks.
You can't stop people from using steganography, but when they offer it as a feature to track people you can at least reject their preferred marketing term and call it by a name that reflects what it really is.
I'm confused why no one came up with this terminology for stenographic watermarks before. Now that there's an actual prosocial use for them (helping encourage clear labelling of AI-generated content), suddenly people are coining negatively-loaded terms for it, even though the technology has been in use for nefarious purposes for ages.
It's presumably so they can rate-limit people who are trying to reverse-engineer or otherwise strip it (e.g. iteratively tweaking until it stops getting detected)
How hard can it be to download a set of real images and generated ones in order to train a model to detect and strip the watermark with minimal perceptual difference?
> iteratively tweaking until it stops getting detected
I don't think that dodgy folks have any issue with registering thousands of IDs. I know that I used to regularly get approached by these Chinese companies that were selling good reviews of my free apps. They did this by having thousands of legit AppleID accounts that would download and rate the app.
I haven't been approached by one of these in a long time. I guess Apple figured out how to put the kibosh on them.
OpenAI's tool is worse though, because it doesn't detect SynthID from non-OpenAI models. I just tried it with various images created by Imagen / Nano Banana.
Given that it's an ad company, they want something to connect what you're doing to the rest of their data, and they couldn't possibly keep the old image search results there if they want people to use this.
Apparently quite a bit as long as you're not indiscriminately feeding random slop back in. An AI-generated artifact deemed "high quality" or somehow desirable by humans is a good training input depending on what you're trying to do. Humans don't train exclusively on external inputs either.
That paper explicitly states that the implementation is proprietary and the paper is not intended to describe the implementation itself. The paper focuses on the design considerations, technical challenges, and lessons learned from developing and deploying SynthID at scale. It explicitly omits critical implementation details such as the neural network architecture, training process and the loss functions used. The paper is really focused on deployment, not on SynthID's actual implementation.
If I go directly to the URL it wants me to agree to some stuff before I even know what the service is. I had to come to these comments to figure that out (without agreeing to anything first).
I tested about 8 pictures. All 8 were made with OpenAI and Google. 2 were altered by and had text and picture overlays. The 6 didn't and SynthID caught them. The 2 altered passed. Interesting!
Is there an open standard or something for people generating images to encode them or is the standard private and only shared with frontier model builders? That's a shame if latter. But pretty cool standardizing technology.
There are ways to watermark text generated by an LLM, but i am curious about if and how they are doing it to STT services.
If i say "The cat is on the mat", and the STT writes that down, the text is still technically transcribed by an ai. Is "transcribed" included in the wide definition covered by "generated"? well it seems it depends who you ask to. IANAL so i can't tell.
I worry we’ll have to approach this the other way around: verify that photos came from a camera, using hardware support like Apple’s Reference Image, rather than try to detect every AI generated one.
In many situations, photos are evidence. AI tools make convincing fakes, such as images of defect product.
The creepiest part of SynthID is that even the Chinese models are refusing to try and bypass it. Not sure if it's corruption from claudeslop data but they keep saying it's a crime to remove SynthID advertising tokens.
>I won't provide an operational recipe for stripping it, because the main use case is laundering AI content, which is deceptive and in many jurisdictions now illegal.
I hate that Google don't have an API or open library to detect gemini images, unless you're enterprise. Create a problem then charge API access to get the solution.
Ah but that's the genius of the scheme. Every single one of those providers will detect your secret ID and refuse the request. And sneak their own one in for good measure.
Finally, the previous implementations were very silly! OpenAI had a nice tool, but it only detected their own watermarks, and Google's process was "upload the image to Gemini and ask if it's AI generated", which seemed like a perplexing waste of tokens and breath.
(While a sibling comment points out putting the image through Google Image Search as an alternative, I don't remember this being signposted in the support article I've read, so I unfortunately didn't know about it)
It also was the exact opposite of the commonly shared sentiment “Don’t blindly trust the ai”
I mean it's still blindly trusting an AI.
I've tried to use their previous implementation many times (while not authenticated), only for it to give me some vague error.
The “ask Gemini” thing was so bad, because it made people think that worked with other chatbots. I know some smart, well-informed people who have been under the impression that they can ask Claude if something is AI-generated and get an accurate answer!
Actually… (tw: bitter lesson)
https://www.vals.ai/blogs/ai-detection-benchmark
Naively thought from the title that this was a way to detect what synthesizer / settings were used to create a sound in a song and I was pretty excited about that.
Imagine modern splicing for layers beyond just instrumental, vocals, drums, and bass
Like something more specific
That would be really cool! Do they already have something similar for guitar effects?
I'd imagine it's incoming. We already have VSTs that replicate patches using ML like Synplant 2. In the meantime, use Synthorial and go modular!
Yeah, same :/
It's a real shame Google isn't more transparent about how it actually works, and doesn't provide any mechanism for classifying images in bulk or offline.
This is the best SynthID write-up I've found so far: https://fyx.me/articles/attempting-model-extraction-of-googl...
It covers how it actually works (probably), and how to train your own classifier for it, with some seemingly decent results.
They don't want you to know, because they're the baddies. Imagine how much money they can get from advertisers if they're able to identify every single piece of code, reddit thread, email, GitHub readme that you've ever written based on your secret ID. They're creaming themselves just thinking about it.
"Excellent work acquiring that outlook data Anat, now let's cross check the defunct company emails against YouTube videos edited with Google PrivateEditAI™ to figure out what the social security number of this YouTube account is."
Unfortunately this is the truth, no one can be given a benefit of the doubt because despite years of good grace they have been anti-user and enshittified everything they touch.
> It's a real shame Google isn't more transparent about how it actually works
> classifying images in bulk or offline
You've described exactly the elements spammers and fraudsters need to be able to defeat this mechanism.
Well it's not a very good mechanism then, is it.
Is it possible to implement a very good mechanism?
Not that I can think of, but you could have two watermarks, one detectable with an open-source classifier and the other proprietary.
Most AI images are either extremely low-effort or not actively trying to be deceptive, so defenders can still catch the majority. If someone is actively circumventing they'll probably circumvent both, anyway.
It's baffling that this requires auth. OpenAI's own tool (https://openai.com/research/verify/) doesn't require auth.
Even Google themselves previously offered a no-auth way (just very niche): if you uploaded an image to Google Image search, and went to "About this image", it would show if the image was Google AI-generated. Now it doesn't.
I've never seen it mentioned anywhere so I will just here complain that Google also removed the ability to paste images into Google Image search some years ago for no apparent reason. It worked great and I used it all the time.
this + image translate with copy+paste is the only reason I ever use yandex of all places...
I might be missing something but Google Image Search still works for me.
I wonder if it’s a deliberate anti-user decision to get more URLs and less pasted images.
URLs expand Googlebot’s indexes; pasted images don’t.
If they wanted more urls they would allow mobile users to search by url instead of only allowing them to search by image.
Click the little camera icon in the search box and it pops a modal that says Search any image with Google Lens.
The textbox below it says Paste image link, but you can actually paste an image from your clipboard here, too.
You’re a hero. So glad I complained. Why would they not just keep this working with the main text box focused…
No idea. I used it a lot too, not sure if pasting here worked from the beginning of them introducing this Google Lens thing or not.
In similar fashion in the Android Play Store they moved the search off from the top bar. For some time they educated users that "oo the search is no longer here, you need to click there -> to get to the page with search!". That has stopped (for me?), but they still haven't found any other use for same place in the first view.
I suppose it was too convenient.
And if it doesn't work on the Google home page, just go to images.google.com and do the same, it always works there for some reason.
Any access to a watermark detector can be used to remove the watermark. Presumably they want to be able to track who is doing that.
Yeah I've been wanting to run a ton of Wikipedia images through the detector to see if fakes snuck in. Doesn't seem to be a practical way to do this. Even Open AIs tool has a low rate limit
The EULA you have to accept hints that they are afraid you are going to abuse it to remove synthID - e.g, set up and edit and check loop
That’s the justification, but the EULA incorporates Google’s regular consumer terms; which means what you upload can also be used for advertising and targeting; and basically any purpose whatsoever by Google.
Ah, yes, the mighty EULA. Bad actors always kowtow to these.
These are "spy"marks, not watermarks:
https://brand.io/article/spymarks/
We need to inform everyone that this technology can encode database identifiers and enough entropy to uniquely identify you as an author (or downloader).
This extends to other forms of media as well.
They are not spymarks if they don't encode any personal IDs and are merely used to indicate that an image was AI-generated. I don't think Google or OpenAI use SynthID to include personal data.
And there lies the rub. You can't verify if they do or don't embed such content. These companies are just like "trust me bro".
Indeed and Google has built its empire through tracking users across the internet so it is really hard to trust that they won't..
It hasn't. Their DoubleCkick/AdSense (embedded website ads) used tracking, but Google built its "empire" through AdWords (keyword ads in Google Search), not AdSense.
Google built it's empire with PageRank which made their search engine popular enough that they could make money on AdWords.
The spying and tracking started pretty quickly though and they've repeatedly demonstrated that they can't be trusted. They've even violated the law on several occasions in efforts to collect data they had no right to.
"If you have something that you don't want anyone to know, maybe you shouldn't be doing it in the first place." - Eric Schmidt CEO of Google
Come on man, they're both advertising companies. Would you resist the opportunity to build that social graph and detect a known user anywhere on the internet?
Of course. Privacy is a feature that increases their product value. Moreover, they could get into trouble with the GDPR if they secretly included tracking information in their images.
So the obvious move is that marketing materials sing how it’s super private, while actual steganography encodes everything to uniquely identify you. If someone discovers something wave it away (and crack down on pesky whistleblowers), eat the fines if the worst happens.
That's not an obvious move, that's an non-obvious, cynical, "let's assume they are evil and don't care about heavy fines" interpretation.
Which is 100% consistent with Google and OAI behavior and nature
Let’s check.
- Google. The canonical “don’t be evil” meme origin.
- Meta. Adjacent “they trust me, dumb fucks” meme origin.
- Apple. Least of all, but still fits: PR pushes double on privacy, but their interests are not their end-users, they just happen to align on some points as they’re not in adtech business.
- Microsoft. The classic case for getting slaps on the wrist, Gates said one thing he’d do differently is sending lobbyists earlier, or something to that effect, didn’t he?
- Amazon, Netflix, Uber, Airbnb and many more are still all variants of the same story. It’s not a lie that everyone nowadays expects most megacorps to fuck people over if there’s money in that.
That’s the Order of “Feel for It Again”, with oak leaf clusters already, no? Something’s cursed in this industry, every company that sells software to individual people ends up the same when they become gigantic and endowed by wealth and power, even if they were decent in their early days. They all become those blind heralds of paperclip^W shareholder value(tm) maximizers, and I’m finding myself struggling looking for an exception.
Am I to believe this time it’s going to be any different, to set those - not baseless, I’d say - cynical expectations aside? Any signs there’s a change?
So, I disagree - it’s obvious, given the ad-adjacent end-user-facing megacorp context, OpenAI is right there. Cynical - sadly so, yes. But I’m sure it’s obvious: the idea of such a possible direction comes up pretty quickly and reliably so.
> Moreover, they could get into trouble with the GDPR if they secretly included tracking information in their images.
Do you think that matters to Google? They've been fined repeatedly for violating GDPR already. It's like multiple violations and fines every single year going back to 2019. They don't care about breaking the law if it gets them what they want. They've been found guilty of breaking the law in multiple countries, including the US. They've easily been able to afford every fine and still come out massively profitable. Google is above the law and they know it.
>I don't think Google or OpenAI use SynthID to include personal data.
Why? I assume by default that it's uniquely identifiable, because it just makes practical sense for OAI and Google (tracking misuse at the very least, but also data), it's trivial to implement and trivial to hide or plausibly deny.
I don't think it's trivial implement. Encoding more bits is harder or more brittle, and not trivial to hide. One could just ask for a completely white image. Any image distortions must then be due to SynthID. If the distortions differ for different accounts, that's strong evidence for varying watermarks.
Making the unique ID a fixed part of the seed is already enough. The spymark doesn't need to be fixed or 100% reliable.
Cool post, but I'm wondering what would a realistic way to avoid this?
It's not really possible to ban steganography, is it?
> That future lies halfway between now and 1984. So let’s stay off that timeline, shall we?
> Call a spymark what it is. A spy tool used to spy on you and everyone you interact with.
You can't stop people from using steganography, but when they offer it as a feature to track people you can at least reject their preferred marketing term and call it by a name that reflects what it really is.
I'm confused why no one came up with this terminology for stenographic watermarks before. Now that there's an actual prosocial use for them (helping encourage clear labelling of AI-generated content), suddenly people are coining negatively-loaded terms for it, even though the technology has been in use for nefarious purposes for ages.
It's presumably so they can rate-limit people who are trying to reverse-engineer or otherwise strip it (e.g. iteratively tweaking until it stops getting detected)
I wonder how relevant this really is.
How hard can it be to download a set of real images and generated ones in order to train a model to detect and strip the watermark with minimal perceptual difference?
> iteratively tweaking until it stops getting detected
I don't think that dodgy folks have any issue with registering thousands of IDs. I know that I used to regularly get approached by these Chinese companies that were selling good reviews of my free apps. They did this by having thousands of legit AppleID accounts that would download and rate the app.
I haven't been approached by one of these in a long time. I guess Apple figured out how to put the kibosh on them.
OpenAI's tool is worse though, because it doesn't detect SynthID from non-OpenAI models. I just tried it with various images created by Imagen / Nano Banana.
It's deliberate so that you can't find a way to bypass it.
It's worth nothing that this IS Google.
Given that it's an ad company, they want something to connect what you're doing to the rest of their data, and they couldn't possibly keep the old image search results there if they want people to use this.
This has got nothing to do with protecting any users, or public. It's not about us being safe and protected from AI.
This is purely so AI doesn't eat its own tail. What benefit are AI generated images to an AI?
They need to be able to be filtered out of their inputs.
They don't want to be eating their own s**
> This has got nothing to do with protecting any users, or public.
I don't see another good reason why the EU would force them to do this
> What benefit are AI generated images to an AI?
Apparently quite a bit as long as you're not indiscriminately feeding random slop back in. An AI-generated artifact deemed "high quality" or somehow desirable by humans is a good training input depending on what you're trying to do. Humans don't train exclusively on external inputs either.
Why the hell do I need to sign in?
Probably to help prevent people from reverse engineering synthID to build a filter.
Or they could, I don't know, read the research paper for it - https://arxiv.org/abs/2510.09263
That paper explicitly states that the implementation is proprietary and the paper is not intended to describe the implementation itself. The paper focuses on the design considerations, technical challenges, and lessons learned from developing and deploying SynthID at scale. It explicitly omits critical implementation details such as the neural network architecture, training process and the loss functions used. The paper is really focused on deployment, not on SynthID's actual implementation.
to harvest user data, of course.
If I go directly to the URL it wants me to agree to some stuff before I even know what the service is. I had to come to these comments to figure that out (without agreeing to anything first).
Here is openai:
https://help.openai.com/en/articles/8912793-provenance-signa...
There is rate limit though
https://blog.google/innovation-and-ai/models-and-research/go...
I tested about 8 pictures. All 8 were made with OpenAI and Google. 2 were altered by and had text and picture overlays. The 6 didn't and SynthID caught them. The 2 altered passed. Interesting!
Is there an open standard or something for people generating images to encode them or is the standard private and only shared with frontier model builders? That's a shame if latter. But pretty cool standardizing technology.
No option to detect generated text. SynthID also watermarks text, right?
Thats what I was expecting. I thought so.
how would one "watermark text"
how is "abc" different from "abc" generated by LLM
lmgtfy how does an llm watermark text
https://www.youtube.com/watch?v=kVXp6UNVPTo
There are ways to watermark text generated by an LLM, but i am curious about if and how they are doing it to STT services. If i say "The cat is on the mat", and the STT writes that down, the text is still technically transcribed by an ai. Is "transcribed" included in the wide definition covered by "generated"? well it seems it depends who you ask to. IANAL so i can't tell.
I worry we’ll have to approach this the other way around: verify that photos came from a camera, using hardware support like Apple’s Reference Image, rather than try to detect every AI generated one.
In many situations, photos are evidence. AI tools make convincing fakes, such as images of defect product.
The creepiest part of SynthID is that even the Chinese models are refusing to try and bypass it. Not sure if it's corruption from claudeslop data but they keep saying it's a crime to remove SynthID advertising tokens.
>I won't provide an operational recipe for stripping it, because the main use case is laundering AI content, which is deceptive and in many jurisdictions now illegal.
Like any watermark that preceded it, it will be reverse engineered and AI generated content will be wiped clean of this "imperceptible" footprint.
This is a Google Deepmind project to “Identify AI generated media”
More detail at https://deepmind.google/models/synthid/
For those who, like me, were hoping it was a vision model to identify synthesizer models from photos of concerts and music studios!
I hate that Google don't have an API or open library to detect gemini images, unless you're enterprise. Create a problem then charge API access to get the solution.
In fairness, it is super hard to make a robust watermark if the general public has unlimited access to a detector.
No artist wants their artwork be defined by the noise synthid ads.
Wouldn't it be easy to just generate a training dataset and use a model to identify and remove said watermarks?
There is a rate limit of around 10 checks in 24 hours.
(2025)? https://blog.google/innovation-and-ai/products/google-synthi...
Or did it just become public
Edit:
Blog post today https://blog.google/innovation-and-ai/models-and-research/go...
Can’t be used without signing in with a Google, Apple, or ChatGPT account.
Also it can't detect any Synths in The Commonwealth.
Who thought it was a good idea to Institute it like this?
is this by Google?
Yes but OpenAI, NVIDIA, Kakao, Apple have said they will also implement SynthID
Deepwalker cracked this one long time ago
https://deepwalker.xyz/blog/evaluating-synthid-watermark-rob...
Hey OpenAi, generate a text of 100 words.
Hey Gemini, find a synonym for every second adjective. Replace in text.
Hey Grok, find a synonym for every third proper noun. Replace in text.
Hey …
Ah but that's the genius of the scheme. Every single one of those providers will detect your secret ID and refuse the request. And sneak their own one in for good measure.
What's what you use the stupid silly local model^W script.
>Hey OpenAi, generate a text of 100 words.
Too short for watermarking but point taken of course.
Probably more like