Lilith Lilith.

What is a watermark in AI text?

A text watermark is not a hidden character or a secret sentence. In one well-known construction, a model slightly favors tokens from a set derived from a secret key and the preceding text. A detector then tests whether those tokens occur more often than chance would predict. It is a statistical signal produced by a participating model, not a universal test for every text on the internet. Method description

For example, an editorial team may check whether an unedited draft generated by its own tool carries the expected signal. A positive result supports a provenance claim for that output. It does not identify the author, show whether a person edited the draft, or show that a text without the signal was not made with AI.

Where the signal weakens

The result depends on the scheme, sample length, and extent of editing. Research describes schemes that can retain a detectable trace after paraphrasing when enough text is available; paraphrasing and mixing the output into other writing weaken the signal. Robustness study

A separate stress test found that repeated paraphrasing can substantially reduce detection rates across several approaches, including watermarking, and described attacks that can cause human text to be labelled as machine-generated. Detector stress test A watermark should therefore not be the sole basis for a penalty or an automated decision about content origin.

What it can establish

A watermark is most useful as narrowly scoped evidence that a particular system may have generated a particular, sufficiently long, lightly edited output. Responsible use requires stating the system, the detector, and the meaning of positive and negative results. Without those conditions, a technical signal can easily become an imprecise “AI text” label.

Sources