raimove

The science behind watermark removal

“Removing a watermark” sounds like a button. In fact it is four entirely different things that happen to sit in the same file. For two of them you can show the mark is gone. For two you cannot. This page explains which is which — no prior knowledge needed.

The structure and argument of this page follow “The science behind watermark removal” by Guillaume Meyer (@guillaumemeyer), author of the watermarks-remover project that raimove is built on. The wording here is our own, written for a general reader. The evidence is the original research, listed at the bottom.

1. Invisible characters in the text

A text can contain characters you cannot see. They have no width, no colour, no shape — nothing happens on screen where they sit. To a program they are there all the same, as readable as an “A”. The best known is the zero-width space, listed in the character table as U+200B. Others are control characters meant for Arabic or Hebrew script, and characters that only modify how a neighbouring character is drawn.

To mark a text this way, you scatter such characters through it by a fixed rule: one every forty letters, say, or a particular pattern at the start of each paragraph. Anyone who knows the rule can find it again later. The Nature paper on Google's SynthID treats this as a category of its own: the mark sits in the spelling, not in what the text says.

Removal is therefore not an art here but bookkeeping. If the mark is made of characters, you delete characters. You can count them first, sort them by type, and check afterwards that none are left. That is what raimove reports: not “95 percent clean”, but “17 characters removed, here they are”.

The same invisible characters have an uncomfortable second use. Boucher and Anderson showed (Trojan Source, CVE-2021-42574) that they can be used to arrange program code so a human reads one thing and the computer runs another. Invisible content in a text is, in general, a channel through which something can travel. Removing it is hygiene — regardless of who put it there.

One group needs care: characters that look deceptively like others. The Cyrillic “а” sits next to the Latin “a” and is barely distinguishable. Replacing those genuinely changes the text, which is why that option in raimove can be switched off and is conservative by default.

The limit of this layer is easy to state: what was never written as a character will never be found this way.

2. Marks hidden in the word choice

A language model writes one word at a time. At each point several words would fit — quick, fast, rapid — and it picks one. That choice is where something can be hidden.

The basic method comes from A Watermark for Large Language Models (2023). Before each word, a secret key splits all the candidate words into two groups, call them green and red. The model gently favours the green group. A single word gives nothing away; that choice could be chance. Across a thousand words, though, a surplus of green builds up that chance practically never produces. Detection is then simple counting: too much green means “written by this AI”.

Google DeepMind turned this into a shipping product. SynthID-Text (Nature, 2024) has the candidate words compete in several rounds, rather like a tournament, and marks the text through who wins. The paper reports a quality check across nearly 20 million real Gemini responses. This is not a laboratory idea; it is in the product.

One detail matters for this page: the key belongs to the vendor. Without it nobody can check whether a mark is still present — raimove included. We can verify that characters are gone. We cannot verify what the vendor's detector is about to say.

Why rewriting helps — and why only partly

If the mark lives in the word choice, only one thing helps: different words. Moving paragraphs around achieves nothing, nor does changing headings. The text has to be reformulated sentence by sentence.

How well that works has been measured. The paper behind the rewriting tool DIPPER (2023) tested several detection programs with it. One of them, DetectGPT, spotted roughly 7 in 10 AI texts beforehand and about 5 in 100 afterwards — each measured at a setting where the program wrongly accuses at most 1 percent of human texts. Without that condition, any hit rate can be made to look good.

Repeated rewriting is starker still. A group around Sadasivan (Can AI-Generated Text be Reliably Detected?) fed the rewritten text back in to be rewritten again, five times over. The watermark's hit rate fell from 998 to 97 cases in 1000. The same paper explains why this works at all: the more AI text resembles human text, the less any detector can achieve, however well it is built.

And now the caveat that tends to go missing in marketing copy. The same researchers who invented the watermark went back to it in 2024 (On the Reliability of Watermarks). The finding: rewriting dilutes the mark, it does not erase it. Traces stay stuck in individual word sequences. Even after thorough rewriting by humans, the watermark became detectable again on average once about 800 text units were available — roughly 500 to 600 words.

Which gives an uncomfortable rule of thumb: short rewritten texts look clean, long ones accumulate traces. That is why this layer in raimove is explicitly called best-effort rather than a promise. And every pass costs something: afterwards the text sounds less like the original.

Two practical rules come straight from the research:

  • Rewrite with a model from a different family. The mark is applied while the text is being written — have Claude rewrite Claude's text and it can come back freshly marked.
  • Today's marks attach to single words, because a response has to appear word by word. Research has methods that mark whole sentences instead and are meant to survive rewriting (SemStamp and PostMark). The major vendors do not deploy them at the moment. If they do, this layer becomes considerably harder.

A third variant is left out here: marks built into the model itself during training. They do not sit in the file, they sit in the model. A program that cleans files cannot reach them.

3. The label inside the file

Every image, text or office file has two parts: the content itself and a kind of label. It records which camera took the photo, which program wrote the file, when that happened, sometimes where. None of it is visible in the document.

C2PA — announced on photos as “Content Credentials” — is such a label, digitally signed. The signature proves nobody altered the label after the fact. It does not prove the label is still there: remove it entirely and the signature goes with it.

This layer is therefore as checkable as the first. The label sits at a known place in the file. You delete it, open the file again and look whether it is gone. That is exactly what raimove does before a cleaned file is offered for download.

There is a second design, though, and it is the reason for the warning in the app. Instead of putting the label in the file, a vendor can keep it on their own server and leave only an unobtrusive mark in the picture pointing at it. Delete the metadata and the label is formally gone — but the mark in the picture lets it be reattached. A successfully removed signature therefore means: the enclosure is gone. Not: there is nothing left to find.

The pattern repeats across formats: PDF, Word, OpenDocument, PNG, JPEG, SVG, HTML, Markdown. Different packaging, same principle — what sits tidily in a field can be removed and counted. What sits in the content cannot.

4. Marks inside the picture

In images the mark sits in the pixels. Millions of them are altered very slightly, in a pattern that the matching program recognises and the eye does not notice. Three methods appear again and again in the research: StegaStamp survives even printing the image and photographing it again; Tree-Ring hides the pattern in the randomness an AI starts from when it creates the picture at all; The Stable Signature alters the image generator so every picture it produces carries the same signature.

The obvious tricks do not work. The comparison study WAVES (2024) shows that cropping, blurring and heavy compression are all survived. Something else does work: having the picture computed anew.

A group around Zhao demonstrated this (Invisible Image Watermarks Are Provably Removable Using Generative AI, 2024). You disturb the image until the mark falls apart, then let an AI reconstruct a clean picture from what is left. Because that AI does not know the mark, it does not put it back.

The catch: recompute a picture and you get a different picture. CtrlRegen (2025) softens this by guiding the process along the content and layout of the original. In the paper's measurements a StegaStamp image was previously detected practically every time and afterwards in about one case in a hundred; for Tree-Ring the rate fell from 99 to 12 percent. The stronger the setting, the more of the mark disappears — and the more of the picture.

Worth stating plainly: raimove does not offer this. This web app handles text and the labels inside files. Image recomputation is an optional add-on in the core project and needs substantial computing power. There is also no way to check the result here: whether an image mark is truly gone cannot be established without the vendor's own tool — and what we cannot measure, we do not claim.

5. Why this gets written about openly

It is fair to ask whether such methods should be described in public. Research has given the same answer for years: every attack named here comes from peer-reviewed publications, and some were published by the very people who build the watermarks. Protection whose limits nobody knows is not protection, it is an assumption.

Since Article 50 of the EU AI Act began requiring AI content to be labelled, this concerns more than specialists. Labels are becoming ordinary — and with them the question of what sits in your own files, and who decides that.

This page is written for people who own the content they want to clean: their own drafts, their own research, their own files that got stamped along the way without being asked. That provenance labelling is useful does not contradict any of it. Being pro-privacy is not the same as being against transparency.

What “removed” is allowed to mean here

DemonstrableDeleted characters, removed labels, files reopened afterwards to check. These are statements about the file itself — the app shows them rather than asserting them.
Best effortRewritten text, recomputed images. These are statements about a procedure, not about a vendor's detector. The research shows both things: that the procedures work, and that traces can remain.
Not coveredLabels kept on someone else's server; marks in audio and video; marks inside the model itself; anything that would require the vendor's secret key.

None of the cited papers claims a mark can be removed permanently and with certainty. Which leaves the one statement this page is willing to make: a removed mark does not mean no AI was ever involved in the text. A clean report is not proof of human authorship — and was never meant to be.

And in practice?

The first and third layers — invisible characters and file labels — run with no account, no key and no stored uploads: file in, report out, removed characters visible. Rewriting text needs your own access to a language model and stays what the research describes: a good attempt.

Open the cleaner →

References

  1. Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., & Goldstein, T. (2023). A Watermark for Large Language Models. ICML 2023.
  2. Dathathri, S., See, A., Ghaisas, S., Huang, P.-S., McAdam, R., Welbl, J., et al. (2024). Scalable watermarking for identifying large language model outputs (SynthID-Text). Nature 634, 818–823.
  3. Kirchenbauer, J., Geiping, J., Wen, Y., Shu, M., Saifullah, K., Kong, K., et al. (2024). On the Reliability of Watermarks for Large Language Models. ICLR 2024.
  4. Krishna, K., Song, Y., Karpinska, M., Wieting, J., & Iyyer, M. (2023). Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense (DIPPER). NeurIPS 2023.
  5. Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W., & Feizi, S. (2023). Can AI-Generated Text be Reliably Detected?
  6. Hou, A. B., Zhang, J., He, T., Wang, Y., Chuang, Y.-S., Wang, H., et al. (2024). SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation. NAACL 2024.
  7. Chang, Y., Krishna, K., Houmansadr, A., Wieting, J. F., & Iyyer, M. (2024). PostMark: A Robust Blackbox Watermark for Large Language Models. EMNLP 2024.
  8. Boucher, N., & Anderson, R. (2023). Trojan Source: Invisible Vulnerabilities. IEEE S&P; CVE-2021-42574.
  9. Coalition for Content Provenance and Authenticity — C2PA Specifications.
  10. Tancik, M., Mildenhall, B., & Ng, R. (2020). StegaStamp: Invisible Hyperlinks in Physical Photographs. CVPR 2020.
  11. Wen, Y., Kirchenbauer, J., Geiping, J., & Goldstein, T. (2023). Tree-Ring Watermarks: Fingerprints for Diffusion Images that are Invisible and Robust. NeurIPS 2023.
  12. Fernandez, P., Couairon, G., Jégou, H., Douze, M., & Furon, T. (2023). The Stable Signature: Rooting Watermarks in Latent Diffusion Models. ICCV 2023.
  13. An, B., Ding, M., Rabbani, T., Agrawal, A., Xu, Y., Deng, C., et al. (2024). WAVES: Benchmarking the Robustness of Image Watermarks.
  14. Zhao, X., Zhang, K., Su, Z., Vasan, S., Grishchenko, I., Kruegel, C., et al. (2024). Invisible Image Watermarks Are Provably Removable Using Generative AI. NeurIPS 2024.
  15. Liu, Y., Song, Y., Ci, H., Zhang, Y., Wang, H., Shou, M. Z., & Bu, Y. (2025). Image Watermarks are Removable Using Controllable Regeneration from Clean Noise (CtrlRegen). ICLR 2025.
  16. Verordnung (EU) 2024/1689 (AI Act), Artikel 50 — Transparenzpflichten für KI-generierte Inhalte.

Every link was checked before publication. The numbers in the text come from the paper named beside them and hold for that paper's setup — not as a general promise.