Dominic Feron

Can a Sentence Prove Its Maker?

Claude's new text watermark may identify a statistical signature. The danger begins when institutions mistake that signal for a verdict.

Can a sentence prove that Claude wrote it?

My first answer was yes. If Anthropic changes the odds of which word Claude chooses next, those tiny choices can add up to a hidden signature. A detector with the right key can test for the pattern. No odd spaces, secret characters, or visible stamp are needed.

That sounds stronger than it is.

The detector is running a statistical test. It can say that a passage contains more of the favored choices than chance should produce. It can set a very low false-positive rate under test conditions. It cannot watch the sentence being born. The jump from “this pattern is unlikely by chance” to “this person used Claude to cheat” is made by a school, employer, publisher, or court.

The jump contains most of the danger.

Suppose a student asks Claude for a draft, rewrites half of it, translates one paragraph, and adds quotations from two human authors. How much of the final essay did Claude write? The watermark score will move as words are replaced. Even a perfect reading of the signal cannot tell you who supplied the argument, who checked the facts, or who is responsible for the result.

My first answer fails because authorship is not a property stored inside a sentence. It is a history.

A watermark can record part of that history. It cannot preserve the chain once the text leaves the generator. Another model can paraphrase it. A person can edit it. An open model can produce the same kind of prose without adding Anthropic’s mark at all. The more rewriting allowed, the less complete the record becomes.

Making the mark tougher creates another problem. Researchers have shown that attackers can learn enough about some watermark schemes to scrub the signal. They can also spoof it. In a piggyback attack, marked text is altered or extended with malicious content while enough of the signature survives. Robustness then turns against the provider: the stronger the surviving mark, the more convincing the false attribution may look.

This is where the copyright trap becomes interesting.

Researchers recently used a jailbreak and a costly extraction method to recover about 96 percent of Harry Potter and the Philosopher’s Stone from Claude 3.7 Sonnet. That result is not a copyright judgment. A statistical mark would not prove infringement on its own either. Courts care about the protected expression, similarity, access, defenses, and who did what.

But a reliable Claude signature could become one piece of attribution evidence. The same feature sold as a way to identify harmless synthetic prose might help connect a copied passage to Anthropic’s system. A watermark does not mark ownership. It may mark involvement.

That is not necessarily bad for Anthropic. Providers also need ways to investigate abuse, measure synthetic content, and meet the EU’s new transparency rules. The law itself uses sensible language: marking should be effective and reliable as far as technically feasible. It does not demand magic.

Institutions might.

Andrej Karpathy’s advice to schools was blunt: assume that work done outside the classroom may have used AI, because detectors can be defeated. I would narrow that claim. A provider watermark can work well when the model, key, text length, and editing conditions are known. Inside a controlled system, it may be useful for audits and aggregate measurement.

What it cannot do is identify all AI text, from every model, after arbitrary editing, with no false accusations. The open world gives an attacker both erasers and counterfeit stamps. It also gives ordinary writers mixed tools and messy workflows that do not fit a binary label.

So will AI-text identification become a virus-style arms race? For hostile users in an open channel, probably. Detection will improve. Evasion and spoofing will follow. One hundred percent identification is imaginable only in a closed process where generation is logged, outputs are signed, and the chain of custody remains intact.

Outside that process, the best available answer has a condition attached: trust the watermark as evidence about a signal, never as a verdict about a person.