Nikhil Deshpande

provenance

The Em Dash With a Standards Body

·LinkedIn post

Two sheets of handmade paper under raking light; one carries a faint embossed mark, the other a single tiny blue fibre, with a magnifying glass resting on it.

Anthropic has published how Claude will mark AI-generated content. Watermarks woven into the text, C2PA metadata on files, detection tools to follow.

I have been trying to work out what this actually gives us, and I keep arriving at the em dash.

For about two years now, the em dash has been a folk test for AI writing. Someone posts a paragraph with a few of them and the comments fill up with accusations. It is a terrible test. Plenty of people have used em dashes their whole lives. Plenty of machine text has none. But the test spread anyway, and the result is that writers have started avoiding a piece of punctuation to escape suspicion. The signal changed behaviour without producing any truth.

Watermarking is the same test with a standards body behind it.

I am not guessing at the failure modes. They are in Anthropic’s own document, stated plainly, which I do respect. A mark can appear on text whose words and ideas are entirely human, because people use these tools to proofread, translate and summarise. And a mark can be absent from text the machine wrote, if it was heavily edited, paraphrased, translated, kept short, converted between formats, or screenshotted.

So it is unreliable in both directions. Which would be fine if everyone treated it as weak evidence. Nobody will. Official infrastructure gets trusted far past its accuracy, and the absence of a mark will very quickly be read as proof that a human wrote something, which it is not.

Now follow the incentive.

If marks vanish under paraphrasing, format conversion and screenshots, then anyone determined to pass machine text off as their own has a clear and easy path. Run it through another model. Retype it. Screenshot it. Convert it.

Meanwhile the person who used Claude to translate her own Marathi story into English, so she could send a sample to a publisher who does not read Marathi, is carrying a mark on work that is entirely hers.

The system flags the honest and clears the sophisticated. That is not a small design flaw. That is the whole thing running backwards.

I want to be fair about what marking is genuinely for. Provenance metadata on files is real and useful. Keeping synthetic text out of the next round of training data matters. Platforms filtering at scale will find it valuable. It is decent plumbing.

It just does not answer the question anyone is actually asking, which is not “did a machine touch this” but “who is responsible for what I am reading”.

Last Diwali we published SrujanMitra, an AI-generated Diwali Ank in Marathi. Under every piece we printed the name of the model that wrote it. Gemini, Claude, ChatGPT, Grok, Deepseek. The poetry was poor and we published it anyway, with the model named.

No watermark. No detection tool. No standard. A line of text under each piece.

That told readers more than any embedded signal will, and it cost nothing but the willingness to say it.

So, in that spirit.

I used Claude to write this post. The argument is mine, the em dash comparison is mine, SrujanMitra is my work and the Marathi translation example comes from people I know. The sentences were drafted with the machine and then pulled apart and put back by me. Some are more mine than others.

This post will almost certainly carry a watermark.

Which is the case I have just spent this whole post describing.