Claude is becoming quite ambitious with watermarking, and I can sense the issues from far away.
Anthropic aims to facilitate the identification of AI-generated text, and in theory, I have little to critique. The company is testing an invisible watermark that can be integrated directly into the output created by Claude.
This seems like a logical approach. AI-generated text is prevalent, and knowing the source could be beneficial. Additionally, Anthropic is not merely hiding a marker in the document; its strategy alters the way Claude selects words to create a detectable statistical pattern.
However, there is one aspect that concerns me. Anthropic is examining how durable this watermark can be, even after modifications to the text.
That’s where I sense potential issues.
Nadeem Sarwar / Digital Trends
Did Claude alter my writing, or did it actually write it?
Consider translation for a moment. If someone crafts an entire essay in Spanish and asks Claude to translate it to English, the concepts, research, and arguments belong to them. Claude's only role is the translation.
Still, the resulting text might bear Claude’s watermark.
The same concern exists for proofreading. If a person writes something and then asks Claude to correct grammatical errors, what occurs if they want to shorten a paragraph, adjust its tone, refine dictated text, or simplify a complex sentence?
These scenarios are no longer outliers. More individuals are utilizing tools like ChatGPT, Gemini, and Claude for routine tasks that do not involve the creation of original content. A watermark can indicate Claude’s participation in a text, but it doesn’t clarify whether Claude authored it. Anthropic emphasizes that the watermark signifies Claude’s involvement, not the identity of the original creator.
Now, envision having to clarify that distinction to a professor after their detection software flagged your essay.
Unsplash
We are aware of the challenges associated with AI detection
I wouldn’t be nearly as concerned if the track record of AI detection were particularly successful. Unfortunately, it is not.
MIT Sloan’s recommendations are clear regarding current AI detectors. They indicate high error rates that can lead instructors to mistakenly accuse students of wrongdoing.
We’ve witnessed the consequences. There are students who have had to defend themselves against claims that they submitted AI-generated work when they insisted on writing it themselves. In one documented case by The Guardian, a student’s essay was incorrectly marked as completely AI-generated, despite the student stating they had only utilized approved spelling and grammar correction tools. Their appeal was ultimately upheld.
To clarify, Claude’s watermark is fundamentally different. Traditional AI detectors analyze writing and estimate whether an AI could have produced it. Anthropic, on the other hand, intentionally embeds a detectable signal in Claude’s outputs. In theory, this should enhance the reliability of the system significantly. However, reliability is not the sole concern; interpretation is as well.
Nadeem Sarwar / Digital Trends
We are employing AI to prove we didn’t use AI
Things have reached a somewhat absurd level.
Students anxious about AI detection are turning to so-called AI humanizers, which specifically rewrite text to reduce the chances of triggering detection systems. Some students even utilize these tools on their own work due to fears of false positives. Naturally, companies developing detection systems are now trying to devise ways to spot these humanizers.
Consider that again.
A human writes something, worries that it might be flagged as AI-generated, modifies it with another AI to appear more human, and then has yet another system determine if the AI adequately humanized it.
It resembles a technological ouroboros.
Creating a watermark in Claude that can withstand editing and translation is technically remarkable. Prior studies have demonstrated that translation can compromise certain text-watermarking techniques; addressing this vulnerability would be a significant advancement.
Nonetheless, I question whether enhancing the watermark’s resilience truly resolves the more pressing issue.
Rachit Agarwal / Digital Trends
A watermark requires context
There are valid reasons to watermark AI-generated content. It can assist in identifying mass-distributed misinformation, undisclosed synthetic text, or AI-generated material that may later be included in training datasets.
The issue lies in the fact that AI assistants are now utilized for far more than just creating content from scratch. People employ them to translate text, proofread documents, summarize research findings, assist with coding, enhance accessibility, or simply refine an email before sending it. In those situations, detecting AI involvement does not inherently indicate who actually produced the work.
All of these interactions involve AI to varying extents. If Claude composes an essay from inception, knowing that is valuable. However, if Claude translates an essay that someone spent three weeks researching and crafting, the awareness of Claude’s involvement is significantly less informative.
The watermark can effectively respond to the question, “Did Claude influence this?” My concern is what may happen when people begin to treat the answer as evidence of “Did Claude author this?”
Anthropic can develop the most advanced watermark possible. However, if users do not comprehend that distinction, I fear we will encounter significant issues.
Other articles
Claude is becoming quite ambitious with watermarking, and I can sense the issues from far away.
Anthropic is working on the challenge of recognizing AI-generated text, but a continuous watermark might lead to another issue if AI support is confused with AI authorship.
